Why Good Evaluations Fail Without a Theory of Change
Hatched by Anemarie Gasser
Apr 19, 2026
9 min read
6 views
62%
The uncomfortable truth: evidence rarely speaks for itself
A program can work, and still look like it failed. It can fail, and still look like it worked. That is the uncomfortable center of evaluation: numbers do not interpret themselves, and outcomes do not arrive wearing name tags.
This is why so many decisions built on evidence feel oddly fragile. People ask, “Did it work?” as if the answer were a simple reading from a thermometer. But evaluation is not weather reporting. It is closer to detective work, where the same clue can support multiple stories, and the quality of the conclusion depends on whether you can trace the path from action to outcome.
The deeper question is not whether something changed. It is: what changed, through what pathway, under what conditions, and what else could plausibly explain it?
That is where the real tension lives. The first instinct is to hunt for a result and declare victory or failure. The harder, more useful task is to build an interpretation that survives contact with reality. Without that, evaluation becomes a ritual of certainty performed over uncertain evidence.
Results are not explanations: the missing middle matters
Imagine a city launches a job training initiative. Six months later, employment among participants rises. That sounds like success. But if you stop there, you have only a pattern, not a reason.
Maybe the training improved skills. Maybe the labor market improved. Maybe participants were already more motivated than average. Maybe the rise would have happened anyway. Maybe the program helped some people while leaving others behind. Each of these possibilities leads to a different interpretation, and a different decision.
This is the central weakness of many evaluations: they jump from observed change to implied causation without examining the middle steps. The missing middle is where meaning lives. It is the chain of events, decisions, behaviors, and contexts that turn an intervention into an outcome, or block it entirely.
A useful way to think about this is as a causal pathway. A program does not create impact by magic. It creates impact, if it does, through a sequence: inputs lead to activities, activities change intermediate behaviors or capacities, those changes influence outcomes, and context shapes every link. If any link weakens, the final result may be diluted, delayed, or distorted.
The most misleading evaluation is not the one with no data. It is the one with data but no pathway.
This is why a theory of change is not decorative paperwork. It is the map that tells you what the result means. Without it, a positive number can be accidental and a negative number can be premature.
The real job of evaluation is not judgment, but interpretation
Most people think evaluation is about verdicts: success or failure, effective or ineffective, worth scaling or worth cutting. But the more sophisticated role of evaluation is interpretation under uncertainty.
That distinction matters because real programs live in messy systems. They are not laboratory experiments with clean boundaries. They are implemented by humans, received by humans, and shaped by politics, incentives, timing, and local conditions. In that environment, the question is rarely “Did it work?” in the absolute sense. It is more often:
- Did it work for whom?
- Under what circumstances?
- Through which mechanisms?
- Compared with what alternative?
- What should we do differently next time?
This is where contribution analysis becomes powerful. Instead of pretending you can always isolate a single cause, you ask whether the available evidence supports a plausible causal contribution. In plain language: can you make a credible case that the intervention mattered, even if it was not the only thing that mattered?
That shift is profound. It moves evaluation away from courtroom logic, where one cause must win and all others lose, and toward historical logic, where complex events are explained by converging causes. A child’s growth, for example, cannot be attributed to one dinner, one night of sleep, or one doctor visit. Yet each may have contributed. Similarly, social change rarely belongs to a single actor, but that does not mean no actor contributed meaningfully.
The task is not to eliminate ambiguity. The task is to discipline ambiguity so that it can inform action.
How to tell whether a change is meaningful
If interpretation is the real challenge, then the evaluator needs tools for sorting signal from noise. One useful mental model is to ask whether a result passes four tests: plausibility, specificity, robustness, and coherence.
1. Plausibility
Does the outcome make sense given the intervention and the context? If a low-cost text-message reminder program reduces missed appointments, that is plausible. If it suddenly halves hospital mortality, skepticism is warranted unless the pathway is very strong.
2. Specificity
Did the pattern change in a way that matches the intended mechanism? A reading program should improve literacy, not necessarily every academic metric at once. If everything improves equally, something broader may be going on. If the expected indicators do not move at all, the causal story weakens.
3. Robustness
Do multiple forms of evidence point in the same direction? Quantitative trends, participant accounts, implementation data, and comparison groups may all tell parts of the story. When independent clues converge, confidence increases. When they conflict, the evaluation becomes more interesting, not less.
4. Coherence
Do the observed effects fit together as a system? A job training program that boosts enrollment but not completion, or completion but not employment, may be revealing a broken pathway rather than a failed idea. Evaluations often miss this because they look for one headline metric instead of a linked sequence.
Think of this like diagnosing a car problem. If the engine will not start, the useful question is not “Is the car bad?” It is whether the battery, fuel line, starter, or ignition system is failing. Evaluation works the same way. Outcomes are symptoms. Pathways are diagnostics.
Good interpretation does not ask whether the result is impressive. It asks whether the result is intelligible.
That shift saves organizations from two common errors: overclaiming success when a lucky trend appears, and overreacting to failure when the pathway was partially working but blocked at a later stage.
A better framework: from outcome worship to pathway thinking
Many institutions are organized around final numbers because final numbers are easy to report. But final numbers are often the least informative place to learn. The deeper insight is that impact is usually a chain of small truths, not a single big one.
This leads to a more useful framework for any evaluation or decision process:
Step 1: Define the intended causal chain
Write down the story of change in plain language. What is supposed to happen first, second, and third? What assumption must hold at each step?
Step 2: Identify failure points in the pathway
Do not ask only whether the intervention failed overall. Ask where it may have broken. Was the target population reached? Did they engage? Did behavior change? Did the environment support the change?
Step 3: Collect evidence at each link
Do not rely on the final outcome alone. Gather indicators for implementation quality, intermediate behaviors, participant experience, and contextual shifts. A good evaluation resembles a chain of clues, not a single photograph.
Step 4: Compare competing explanations
Any result can be explained in multiple ways. A strong interpretation actively tests alternatives. What else changed at the same time? Who benefited? Who did not? What would we expect if the program were not responsible?
Step 5: Decide with calibrated confidence
The goal is not perfect certainty. The goal is a decision that reflects the weight of evidence and its limits. Sometimes that means scaling. Sometimes it means redesigning. Sometimes it means waiting for more information.
This framework is valuable because it changes the emotional posture of evaluation. Instead of treating uncertainty as a threat, it treats uncertainty as part of the object of study. That is a far more mature stance, and it produces better decisions.
Why this matters beyond evaluation
The logic here applies far outside the world of programs and metrics. In business, leaders often mistake revenue growth for product health without asking whether growth is driven by brand, pricing, market timing, or unsustainable acquisition costs. In healthcare, a policy may improve one measure while worsening another because the causal chain was never fully understood. In education, a test score bump may hide a narrowing curriculum or a burden shift onto teachers.
The same pattern repeats everywhere: people crave simple verdicts, but reality arrives as an interconnected story.
This is why the most valuable thinkers in complex environments are not the ones who sound most certain. They are the ones who can say, with discipline, “Here is the evidence, here is the pathway, here are the plausible alternatives, and here is how confident we should be.” That is not weakness. It is intellectual honesty with operational value.
It also changes how organizations learn. If leaders reward only final outcomes, teams will hide uncertainty and optimize for appearances. If leaders reward clear causal reasoning, teams will surface weak links earlier, learn faster, and adapt more intelligently. In that sense, evaluation is not just a measurement practice. It is a culture-shaping practice.
Key Takeaways
- Do not confuse outcomes with explanations. A result is not evidence of causality until you can trace the pathway that produced it.
- Build a causal chain before you judge performance. Ask what must happen first, second, and third for the intended effect to appear.
- Test interpretations against alternatives. A credible evaluation compares competing explanations instead of accepting the first plausible story.
- Look for convergence, not just a headline number. Strong conclusions usually come from multiple forms of evidence pointing in the same direction.
- Treat uncertainty as information. Uncertainty often reveals where the pathway is weak, blocked, or context dependent.
The deeper lesson: good decisions depend on better stories, not just better stats
The biggest mistake in evaluation is believing that better measurement automatically produces better judgment. It does not. Better judgment comes from better causal stories grounded in evidence. Data matter enormously, but data only become wisdom when they are organized around a pathway that explains how change happens.
That is the real synthesis here: evaluation is not the art of announcing whether something worked. It is the discipline of figuring out what kind of change occurred, why it occurred, and how sure we should be that the intervention contributed to it.
Once you see that, the goal of evaluation changes. You stop asking for a single answer and start asking for a defensible interpretation. And that is a much more useful standard, because it reflects the way the world actually works: not as a clean verdict, but as a web of causes, conditions, and consequences.
In the end, the most powerful question is not “Did it work?” It is:
What story of change can survive the evidence?
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣