The Public Health Metric That Cities Need to Borrow from Vaccines
Hatched by Emil Funk Vangsgaard
Aug 27, 2026
10 min read
2 views
86%
What if the most important question in public health is not how much illness exists, but whether an intervention can actually make the body, the community, or the health system respond better?
That distinction sounds subtle. It is not. In Manchester, disease related to overweight and obesity was estimated to cost £185.1 million in 2015. A number that large creates urgency, but it does not tell us which actions deserve investment, which outcomes matter first, or how much confidence we should have in a proposed solution.
Vaccine science offers a useful contrast. It does not judge a vaccine only by whether it causes antibodies to appear. It asks whether those antibodies perform a relevant function: can they help immune cells identify and destroy a pathogen? It then compares that functional response with a standard, accounts for variation, and sets a threshold for deciding whether the new intervention is good enough.
The deeper lesson is this: public health improves when it measures function rather than intention, activity, or appearance. Cities confronting obesity need something like an OPA GMR for policy: a disciplined way to ask whether an intervention produces a meaningful protective effect compared with what already exists.
The difference between a signal and a solution
A public health program can generate impressive signals without solving the problem it was designed to address. Thousands of people may attend workshops. A school may distribute healthier menus. A city may open a new walking route. These are useful activities, but activity is not the same as impact.
The distinction resembles the difference between antibody concentration and immune protection. A laboratory can detect antibodies after vaccination, yet the practical question is whether those antibodies can support opsonophagocytic activity, the process by which pathogens are marked for destruction by immune cells. The functional assay moves the conversation from presence to performance.
Obesity policy faces the same measurement problem. A program may increase knowledge about nutrition while leaving food prices, neighborhood design, working hours, stress, and access to exercise unchanged. It may produce short term weight loss among highly motivated participants while doing little for the wider population. It may even improve an average while increasing inequality, because the people best positioned to participate benefit most.
The first mental model, then, is the translation test:
A result matters only when it translates from a measurable signal into the function the intervention is supposed to improve.
For a vaccine, the translation is from an immune response to the ability to help clear a pathogen. For an obesity strategy, it might be from a new service to sustained changes in eating patterns, physical activity, metabolic health, or the financial burden of disease. Each link should be made explicit.
This does not mean that intermediate measures are useless. Antibody titers are important, just as program attendance, food purchasing patterns, and self reported activity can be important. But they are steps in a causal chain, not the final destination. The danger begins when a convenient measurement is mistaken for the outcome that citizens actually need.
Why comparison matters more than impressive numbers
The geometric mean ratio used in vaccine research contains a second lesson: an intervention has meaning only in relation to a comparator.
Suppose a new vaccine produces an average functional antibody response of 120 units. That may sound powerful. But powerful compared with what? If the established vaccine produces 100 units, the ratio is 1.2. If it produces 300, the new result looks very different. The ratio makes the comparison visible, while the geometric mean helps prevent a handful of unusually high or low results from distorting the picture.
Policy debates often lack this discipline. A city may report that a program helped 10,000 residents, but the relevant comparison could be no program, an existing service, a cheaper intervention, or a redesigned version of the same service. Without a comparator, even a positive result is difficult to interpret.
Imagine two approaches to reducing obesity related illness. The first funds intensive coaching for a small group of residents. The second changes the placement and pricing of food in public facilities, improves access to active transport, and supports healthier meals in schools. The first may produce a larger average weight change among participants. The second may produce a smaller change per participant but reach tens of thousands of people and reduce exposure to unhealthy defaults every day.
Which is better? There is no responsible answer without defining the unit of comparison. We need to examine effect per participant, total population effect, cost per improvement, durability, and distribution across income groups. A single average cannot carry all of that information.
The second mental model is the counterfactual ladder. For every proposed intervention, ask five questions:
- What happens with no new action?
- What happens if current services continue unchanged?
- What happens with a lower cost alternative?
- What happens when the intervention is offered at realistic scale?
- Who benefits, and who is left out, under each scenario?
This ladder prevents a common error: comparing a carefully supported pilot with an imaginary baseline instead of with the imperfect system that would actually operate in its place.
The threshold is not the same as perfection
Vaccine trials also show why public decisions need thresholds rather than fantasies of certainty. In some comparisons, noninferiority is declared when the lower bound of the 95 percent confidence interval for the geometric mean ratio remains above 0.5. The exact threshold belongs to a technical context, but the general principle is widely useful: an intervention may be accepted not because it is superior in every respect, but because evidence indicates that it is not unacceptably worse than the standard.
This is especially important in obesity policy, where interventions are rarely perfect. A new measure may not produce dramatic weight loss. It may still be worthwhile if it reduces complications, improves blood pressure, makes healthier choices easier, or reaches people who have never benefited from clinical services. Conversely, a program may show a statistically significant average improvement and still be too expensive, too short lived, or too unequal to justify broad adoption.
A good threshold should therefore answer three different questions:
Is the effect real? The observed improvement should be distinguishable from random fluctuation and measurement error.
Is the effect meaningful? The size of the improvement should matter to health, quality of life, or public spending, not merely to a spreadsheet.
Is the effect acceptable at scale? The intervention should remain useful when deployed across ordinary neighborhoods, ordinary staff, and ordinary budgets.
These questions are often collapsed into one word: success. That word is too blunt. A project can be scientifically credible but practically trivial. It can be clinically meaningful but financially unsustainable. It can be effective on average but unfair in distribution.
A more mature approach uses a portfolio of thresholds. For example, a city might require a program to show a modest improvement in metabolic risk, maintain participation after one year, cost less than a specified amount per quality adjusted life year, and avoid widening health inequalities. The purpose is not to make innovation impossible. It is to make the tradeoffs visible before enthusiasm hardens into policy.
From a £185.1 million burden to a decision system
The £185.1 million estimate for Manchester in 2015 is valuable because it converts an abstract health problem into an economic one. Yet a burden estimate is a starting point, not a strategy. It tells us that inaction is costly, but not whether a proposed action will reduce that cost.
To turn a burden into a decision system, policymakers need to connect four layers:
1. The burden
What costs are being generated, and where? These may include medical treatment, lost productivity, social care, reduced quality of life, and pressure on families. Different costs fall on different institutions, so a program that saves the health service money may require an upfront investment from schools, transport, or local government.
2. The mechanism
How is the intervention expected to work? A food pricing measure, a clinical referral program, a redesigned street, and a school meal reform operate through different pathways. If the mechanism is unclear, evaluation becomes a search for good news rather than a test of a theory.
3. The functional endpoint
What observable change would demonstrate that the mechanism is working? This may involve body weight, waist circumference, blood pressure, diabetes incidence, physical activity, diet quality, or reduced use of health services. The endpoint should be close enough to the mechanism to detect change, but important enough to matter.
4. The decision threshold
How much improvement is enough to expand, adapt, or stop the intervention? This threshold should include uncertainty, cost, durability, and equity. It should be specified before results arrive whenever possible, because people are remarkably skilled at redefining success after seeing the data.
This framework exposes why large public health problems cannot be solved by a single metric. The £185.1 million figure describes the size of the fire. A functional evaluation tells us whether a particular hose is reaching the flames.
There is also a lesson about time. The immune response measured after vaccination can be an early indicator of future protection, but it is not identical to every long term outcome. Similarly, a short term change in diet or activity may be a leading indicator, while disease incidence and public expenditure are lagging indicators. Evaluation must track both. If leaders demand only immediate changes in distant outcomes, useful interventions may be abandoned too early. If they accept only early signals, ineffective programs may survive indefinitely.
Designing policies that behave like good experiments
The practical implication is not that every city must run a laboratory trial for every decision. It is that policy should be designed to learn.
Consider a city introducing healthier procurement standards across public buildings. Rather than announcing success after implementation, officials could define a sequence of tests. First, are healthier options actually available and affordable? Second, do purchasing patterns change? Third, do people consume the new options repeatedly? Fourth, are there measurable effects on health indicators? Fifth, are the benefits concentrated among residents who were already healthier?
Each step corresponds to a different point in the causal chain. Failure at an early step suggests an implementation problem. Success at early steps but not later ones suggests that the assumed mechanism may be weak. Success across the chain provides a stronger basis for expansion.
This approach also encourages adaptive scale. Start with a clearly bounded intervention. Establish a comparator. Measure functional outcomes. Expand when the evidence crosses the predefined threshold. Modify when the intervention works for some groups but not others. Stop when the costs exceed the plausible benefit.
The same logic can improve personal decision making. Someone trying to improve health need not rely on a single dramatic goal. They can define a functional target, such as being able to walk for thirty minutes without discomfort, and track the behaviors that support it. Weight may remain relevant, but it becomes one indicator among several rather than the sole verdict on progress.
Key Takeaways
- Measure function, not just activity. Attendance, awareness, and antibody levels can be useful signals, but ask whether they translate into the health function that matters.
- Always name the comparator. A result has no practical meaning until it is compared with current practice, no intervention, or a realistic alternative.
- Set thresholds before judging success. Define what counts as a meaningful, affordable, durable, and equitable improvement before results create pressure to move the goalposts.
- Track the causal chain. Measure implementation, immediate behavior, intermediate health indicators, and long term outcomes. Each reveals a different kind of failure or success.
- Treat equity as part of effectiveness. An intervention that improves an average while leaving disadvantaged communities behind is not fully successful.
The central challenge in obesity policy is not a shortage of concern. The scale of the estimated burden makes that clear. The challenge is converting concern into a sequence of testable decisions: what should change, through which mechanism, for whom, at what cost, and with what degree of confidence?
Vaccine science offers a powerful habit of mind. It asks whether a response performs the job required of it, how it compares with a standard, and whether the uncertainty leaves enough room to trust the conclusion. Public health policy should ask the same questions.
A city does not become healthier because it launches more initiatives. It becomes healthier when its initiatives reliably produce protection in the places where people live, work, eat, and travel. The real measure of progress is not the size of the program, or even the size of the budget. It is the distance between an intervention's visible signal and its actual function, and how rigorously we close that distance.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣