When You Cannot Count Everything, Count What Changes: The Hidden Logic of Realist Evaluation
Hatched by Anemarie Gasser
Jun 06, 2026
10 min read
1 views
82%
The Most Useful Question Is Not “Did It Work?”
A public health program can be crowded, expensive, and technically sophisticated, and still leave one brutal question unanswered: what actually changed because of it? Not what was delivered, not what was intended, not what was measured in a spreadsheet, but what shifted in the real world, among real people, under real constraints.
That question matters because social change rarely behaves like a clean laboratory experiment. Policies meet politics. Programs meet incentives. Communities reinterpret interventions. And the most important effects are often the least obedient to predefined indicators. A clinic opens, but attendance rises only after transport vouchers are added. A health campaign launches, but the real breakthrough comes when local leaders start repeating its message in their own language. A nutrition program fails in one district and succeeds in another, not because the core idea changed, but because the surrounding conditions did.
This is where a deeper tension emerges: we want evidence that is rigorous, but we also need evidence that is alive to complexity. If we only count what was planned, we miss what was actually transformed. If we only collect stories, we risk mistaking vividness for validity. The challenge is not choosing between numbers and narratives, but learning how to detect causal change in messy systems.
That is the shared intellectual territory between realist synthesis and outcome harvesting. One asks, in effect, what mechanisms are activated for whom, in what contexts, and with what results? The other asks, what outcomes actually occurred, and how can we trace back to the contribution of the intervention or network that helped produce them? Put together, they point to a more mature idea of evaluation: not as a verdict, but as a process of disciplined interpretation.
The Old Fantasy of Evaluation: Control the World, Then Measure It
Many evaluation systems are built on an implicit fantasy. First, define success in advance. Then create indicators. Then collect data. Then decide whether the program worked.
This sounds reasonable, but it quietly assumes that change behaves like a machine: linear, visible, and predictable. In reality, public health and social change behave more like weather systems. Small shifts in one place can have outsized effects elsewhere. Similar interventions can produce different outcomes depending on trust, timing, leadership, resource constraints, or local meaning. A strategy that looks weak on paper may become powerful when it interacts with the right social mechanism.
The problem is not just technical, it is philosophical. When evaluation is obsessed with pre-set metrics, it often mistakes what was easy to count for what was important to know. The result is a kind of epistemic blindness. We end up with confidence about outputs and uncertainty about outcomes. We know how many trainings were held, but not whether any relationships changed. We know how many pamphlets were distributed, but not whether the community’s understanding shifted. We know how many facilities were built, but not whether anyone trusted them enough to use them.
This is why the question “Did it work?” is too blunt. It hides at least three more useful questions:
- What changed?
- What helped make that change possible?
- Under what conditions did the change happen?
Those are not just evaluation questions. They are a theory of how social reality works.
In complex systems, the main job of evaluation is not to pronounce judgment, but to reconstruct causality under uncertainty.
Realist Thinking: The World as Context Plus Mechanism
The realist lens begins with a deceptively simple insight: interventions do not work by themselves. They work through mechanisms, and those mechanisms are triggered only in certain contexts. This is why the same program can produce opposite results in different settings.
Imagine a school nutrition initiative. In one area, free meals increase attendance because children are hungry and parents value the support. In another, the same initiative barely moves attendance because the main barrier is not hunger but seasonal labor demands. The intervention is identical, but the mechanism differs, or never activates. The useful question, then, is not whether the intervention has a universal effect. It is whether a particular configuration of context can activate a mechanism that leads to a desired outcome.
This way of thinking is powerful because it rescues evaluation from crude yes or no judgments. It also explains why some of the best interventions seem modest in one place and transformative in another. A peer support group may be ordinary where social trust is low, but catalytic where communities already have informal mutual aid. A community health worker may be ineffective when overloaded and unsupported, yet highly effective when embedded in a responsive local network.
A useful mental model here is the spark and fuel model:
- Context is the fuel: existing norms, resources, politics, trust, infrastructure.
- Mechanism is the spark: the reasoning, motivation, or response triggered by the intervention.
- Outcome is the fire: the visible change produced when spark and fuel meet.
Without the fuel, the spark does nothing. Without the spark, the fuel sits inert. Evaluation becomes the art of identifying which fuels are available, which sparks the intervention can generate, and which fires are likely to follow.
This matters because it changes the role of evidence. Instead of asking for universal proof that a program works everywhere, realist evaluation asks for transferable insight about how it works, where it works, and why. That is a more demanding standard, not a softer one. It requires the evaluator to understand the machinery of change, not just its surface effects.
Outcome Harvesting: When the Most Important Results Were Never in the Plan
Realist thinking helps explain why change happens differently across settings. Outcome harvesting tackles a second blind spot: sometimes the most important outcomes were never specified at the start.
This is not a minor issue. In many initiatives, the original plan names a set of expected outputs, but the actual significance of the work emerges later, often in unexpected places. A policy dialogue might not immediately pass legislation, but it may shift how a ministry talks about a problem. A local advocacy campaign might not change national law, but it may inspire neighboring groups to adopt a new strategy. A training program might not produce the exact behavioral change predicted, but it might build a network that later becomes crucial in a crisis.
Outcome harvesting starts with the outcome itself, then asks how it came to be. That can mean tracing observable changes in behavior, relationships, policy, discourse, or institutional practice, then gathering evidence from multiple actors to understand contribution. This approach is especially useful when the terrain is adaptive, political, or hard to predefine. In such settings, the key signal is not whether a target was hit, but whether meaningful change occurred and whether the intervention plausibly contributed to it.
This is a profound shift in attention. Instead of forcing reality to match the logframe, the evaluator learns to follow reality as it unfolds. That does not mean abandoning rigor. It means embracing a different rigor, one rooted in careful documentation, triangulation, and explanation.
A practical analogy helps here. Think of a gardener who plants seeds in a wild, shifting plot of land. A rigid evaluator would return only to check whether the exact flowers listed on the packet appeared in the exact spots on schedule. An outcome harvester would notice something more interesting: which plants took root, which ones spread, which patches changed soil conditions, and which unexpected species transformed the ecosystem. The point is not to admire surprise for its own sake. The point is that in complex environments, unexpected outcomes are often where the most valuable learning lives.
The Real Breakthrough: Evidence as a Conversation With Reality
The deepest connection between these two approaches is this: both reject the idea that evidence should merely confirm a prior plan. Instead, evidence should help us converse with reality.
That phrase matters. A conversation is not passive recording, and it is not domination. It involves listening, revising, clarifying, and asking better questions in response to what is said. Evaluation in complex public systems should work the same way. We bring a theory, but reality answers back. We bring indicators, but outcomes appear in other forms. We bring an intervention, but the system responds in ways we did not anticipate.
This is why the most useful evaluation systems are not only measurement systems. They are learning systems. They help organizations discover which assumptions were right, which were wrong, and which were incomplete. They do this by linking two disciplines that are often separated:
- Explanation, which asks why change happened.
- Discovery, which asks what change happened in the first place.
Realist synthesis strengthens explanation. Outcome harvesting strengthens discovery. Together they produce something more powerful than either alone: disciplined inference under complexity.
Think about a maternal health program. A conventional evaluation might report increased prenatal visits, but miss that the real breakthrough was a shift in trust between nurses and young mothers. A realist lens would ask what mechanism made that trust possible, perhaps respectful communication, same day appointments, or support from community health workers. Outcome harvesting would notice the unplanned yet crucial result that local women began advising each other to seek care earlier, creating a social norm that no indicator had anticipated.
The full picture is richer than any single metric. More importantly, it is more actionable. If the real mechanism is trust, then future investment should build trust. If the real outcome is peer diffusion, then strategy should support community diffusion, not just direct service delivery.
A Better Way to Evaluate Change: Four Questions
When these approaches are combined, they suggest a practical framework for evaluation and learning. Before asking whether an initiative succeeded, ask these four questions:
1. What changed that matters?
This is the outcome harvesting question. Do not begin by asking whether the original targets were met. Begin by identifying meaningful changes in behavior, relationships, rules, language, access, or institutional practice.
2. What mechanism likely produced that change?
This is the realist question. Did the intervention change incentives, knowledge, trust, identity, legitimacy, coordination, or capacity? A program does not work because it exists. It works because it triggers some human or institutional response.
3. In what context did it happen?
Context is not background noise. It is part of the causal story. Political support, local norms, staffing levels, historical trust, seasonality, and infrastructure can all determine whether a mechanism activates.
4. What evidence would make the story credible?
Complexity should never become an excuse for vagueness. Good evaluation still demands triangulation, comparison, and careful reasoning. The question is not whether we can eliminate uncertainty. It is whether we can reduce it enough to make good decisions.
Credible evaluation in complex systems is not about perfect measurement. It is about making the most defensible causal story from incomplete evidence.
This framework is useful precisely because it avoids two common traps. The first trap is reductionism, pretending that all important change can be pre-specified and counted. The second trap is romanticism, pretending that any interesting story counts as evidence. Good evaluation lives between them.
Key Takeaways
- Stop treating pre-set indicators as the whole truth. They are useful, but they often capture only the easiest layer of change.
- Look for mechanisms, not just outputs. Ask what shifted in motivation, trust, incentives, coordination, or understanding.
- Treat context as causal, not decorative. The same intervention can fail or succeed depending on local conditions.
- Track unexpected outcomes deliberately. Some of the most important effects emerge outside the original plan.
- Use evidence to learn, not only to judge. The best evaluations refine strategy, not just produce a scorecard.
The Real Lesson: Systems Change Us Before We Can Fully Measure Them
The most profound insight from combining these approaches is that change is often visible in the system before it is visible in the metrics. A community begins talking differently. A frontline worker behaves differently. A policy debate shifts framing. An institution starts responding faster. By the time the numbers catch up, the deeper transformation may already be underway.
That is why evaluation should not be treated as a ceremonial afterthought. It is part of the intervention itself. What we choose to notice shapes what we choose to reinforce. What we choose to measure shapes what we fund. What we choose to explain shapes what we believe is possible next.
If there is a single reframing worth keeping, it is this: the goal is not to prove that change happened exactly as planned. The goal is to understand how real change became possible. That is a more humble ambition, but also a more intelligent one. It respects complexity without surrendering rigor.
In a world where public health problems, social programs, and policy efforts unfold in moving terrain, the best evidence is not the evidence that closes the conversation. It is the evidence that opens a better one. And once you start asking not just what was done, but what changed, why it changed, and in what conditions it could change again, you stop evaluating programs as isolated events. You start learning how transformation actually works.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣