Why Good Evaluations Fail Without a Theory of Change
Hatched by Anemarie Gasser
Jun 24, 2026
10 min read
2 views
87%
The Strange Problem with Measuring Success
Most organizations do not fail because they cannot measure anything. They fail because they measure the wrong thing at the wrong time, then mistake the measurement for understanding.
That is the hidden tension at the heart of evaluation work: a result can be real and still be misleading. A program may report higher attendance, more workshops delivered, or more beneficiaries reached, yet none of those numbers tell you whether the thing that actually matters is changing. If you only track outputs, you can become extremely precise about activity while remaining vague about impact.
This is why evaluation often disappoints the people who need it most. Leaders want a verdict, but what they really need is a map. A verdict tells you whether something happened. A map tells you how change is supposed to happen, where it might stall, and what evidence would count as progress along the way.
Evaluation without a theory of change is like checking the weather after already deciding whether to carry an umbrella. It gives you data, but not judgment.
The deeper question is not, “Did this initiative work?” The deeper question is, “What chain of change would have to be true for this initiative to work, and how would we know at each step?” Once you ask that question, evaluation stops being a postmortem and becomes a design tool.
A Theory of Change Is Not a Diagram, It Is a Bet
A theory of change is often presented as a neat flowchart: inputs lead to activities, activities to outputs, outputs to outcomes, and outcomes to impact. But that tidy visual can be deceptive. The real purpose of the theory is not to decorate a report. It is to make an explicit claim about causality in a world that is messy, nonlinear, and full of competing explanations.
In other words, a theory of change is a bet on the logic of transformation.
If a nonprofit trains teachers, the real theory is not merely “training improves teaching.” It may be something more specific: if teachers receive practical coaching, have time to reflect, and work in schools that support experimentation, then their classroom practices will change, which will improve student engagement, which may improve learning outcomes. Each link in that chain is a hypothesis, not a fact.
That is why weak theories of change are so dangerous. They often skip the difficult middle layer, where most actual change occurs. They assume that once a service is delivered, the outcome will follow. But human systems do not work like machines. They respond to incentives, trust, timing, institutional friction, and local context. The missing middle is where good intentions either become results or disappear into routine.
A strong theory of change forces three questions:
- What has to happen first?
- What might prevent it from happening?
- What observable evidence would show the chain is moving?
This matters because evaluation is not only about proving success. It is about discovering whether the chain itself is plausible. If one link is weak, the whole system deserves rethinking.
The Core Mistake: Confusing Activity with Progress
One of the most common failures in evaluation is the activity trap. Organizations count what is easy to count because it feels objective and visible. They report the number of sessions delivered, surveys completed, reports published, or people contacted. These are not useless metrics. But they are only meaningful if they are connected to a believable theory of change.
A simple analogy makes the problem obvious. Imagine trying to improve public health by counting how many bandages are distributed. The number may rise, but that tells you almost nothing about wound healing, infection rates, or long-term recovery. Distribution is an activity. Healing is the outcome. They may be related, but they are not the same thing.
This confusion becomes especially costly in complex social programs. A youth employment initiative may celebrate the number of participants enrolled. But enrollment is only valuable if participants stay engaged, gain relevant skills, find opportunities, and experience better economic trajectories. Without the middle steps, the headline metric can become a comfort blanket that hides failure.
The real discipline of evaluation is to resist the seduction of easy counts. Instead, it asks: What evidence would actually indicate movement along the change pathway? Sometimes that means measuring knowledge gain, behavioral change, network formation, decision quality, or institutional adoption. Sometimes it means admitting that impact is too distant to observe directly, so you need intermediate indicators that are more informative than outputs and more honest than impact claims.
This shift changes the role of data. Data is no longer a scoreboard. It becomes a conversation with reality.
Evaluation as Navigation Through Uncertainty
The best evaluation processes do not pretend the world is linear. They acknowledge that interventions operate inside living systems. People adapt. Context shifts. One change triggers another. A program that works in one district may fail in another for reasons that have little to do with the core intervention itself.
That is why a theory of change should be treated as a navigation instrument, not a certificate of certainty.
Think about sailing. A sailor does not draw a single perfect line from departure to destination and trust that the sea will cooperate. The sailor uses a map, a compass, the wind, and frequent course corrections. Evaluation works the same way. The theory of change supplies direction, but measurement tells you whether you are drifting, advancing, or facing a hidden current.
This is where outcome evaluation becomes powerful. It is not merely the end stage of a project, when someone asks whether the goal was achieved. Done well, it becomes a way of testing whether the causal assumptions embedded in the design were accurate. Instead of asking only, “Did we reach the final destination?” it asks, “What signs show that we are traveling in the right direction?”
That difference matters because many initiatives are too complex for simple before-and-after judgment. Suppose a municipal program aims to reduce homelessness. Did shelter usage increase because outreach improved, because more people are willing to seek help, because rents fell, or because economic conditions worsened? A good theory of change helps disentangle these possibilities. It tells evaluators what to look for, what sequences to expect, and what deviations would require a revised explanation.
A theory of change is not valuable because it predicts the future perfectly. It is valuable because it names the assumptions that must survive contact with reality.
The Best Evaluations Ask Better Causal Questions
Once you see a theory of change as a causal bet, evaluation becomes less about proving and more about interrogating. That shift is profound.
Instead of asking whether the program produced a final result, you ask whether the mechanism is working. Did the intervention change beliefs, access, incentives, skills, relationships, or institutional behavior? Which mechanism mattered most? Which one failed? Did the outcome occur because of the program, or alongside it for other reasons?
This is especially important when multiple forces operate at once. A school reform might coincide with demographic shifts, policy changes, or improved funding. A health campaign might coincide with seasonal trends. A poverty reduction program may benefit from broader economic growth. If you only examine the outcome, you risk mistaking correlation for causation.
A rigorous theory of change helps evaluation distinguish between three different questions:
- Did change happen?
- How did it happen?
- Would it have happened anyway?
The first question is descriptive. The second is explanatory. The third is causal. Too many evaluations stop at the first and pretend they have answered the third.
This is why outcome evaluation should be iterative rather than ceremonial. Early evidence is not just an early verdict. It is a diagnostic signal. If the expected intermediate outcomes are not appearing, the response should not be to defend the original plan more loudly. The response should be to examine the chain. Maybe the intervention is too weak. Maybe it is reaching the wrong people. Maybe the context blocks adoption. Maybe the theory itself is wrong.
That humility is not a weakness. It is what makes learning possible.
A Useful Mental Model: The Theory of Change as a Bridge, Not a Promise
The most useful way to understand the relationship between theory and evaluation is to imagine a bridge.
On one side is intention. On the other side is impact. In between are the spans, supports, and load-bearing assumptions that carry people across. A bridge is not evaluated by its beauty alone. It must be tested at every critical joint. Are the foundations stable? Is the weight distributed properly? Does the structure hold in different weather conditions?
A theory of change is the engineering plan for that bridge. Evaluation is the stress test.
This metaphor clarifies a mistake many organizations make. They treat a theory of change like a promise, something to be defended at all costs. But a theory of change is not sacred. It is provisional. Its job is to expose what must be true, so that reality can confirm, refine, or reject it.
That is why the best evaluations are not embarrassed when a theory fails. Failure is not merely bad news. It is information. It tells you which support beam cracked, which assumption was too optimistic, or which context variable mattered more than expected. In a sophisticated system, failure is how learning becomes precise.
Consider a job training program. The theory may assume that unemployed participants lack only technical skills. But evaluation may reveal that the real barriers are childcare, transportation, unreliable schedules, or weak employer demand. In that case, the original bridge was missing crucial supports. The program was not necessarily useless. It was incomplete.
This is the practical brilliance of pairing theory with outcome evaluation. Theory prevents measurement from becoming shallow. Evaluation prevents theory from becoming self-referential.
What Makes a Strong Change Model in Practice
A good theory of change is not the most complicated one. It is the most testable one.
That means it should be specific enough to guide decisions, but flexible enough to adapt when evidence changes. It should identify the key assumptions, not bury them in jargon. It should make room for context, because human systems are not interchangeable. And it should define success in ways that reveal movement, not just completion.
Here are the qualities that make a theory of change genuinely useful:
- Clarity: It explains the logic in plain language.
- Specificity: It names the actors, behaviors, and conditions that matter.
- Testability: It identifies evidence that could confirm or challenge the theory.
- Sequencing: It recognizes that change happens in stages.
- Context awareness: It acknowledges that local conditions shape outcomes.
The most underrated element is sequencing. Many plans fail because they assume all good things can happen at once. But change is usually cumulative. Trust comes before participation. Participation comes before habit. Habit comes before sustained outcome. When a theory of change respects sequence, evaluation becomes far more intelligent.
This also changes how organizations set indicators. Instead of choosing only lagging indicators, they can build a ladder of evidence. At the bottom are activity measures. In the middle are signs of uptake, behavior, or institutional response. At the top are long-term outcomes. The point is not to flood the dashboard with metrics. The point is to ensure that every metric answers a different and necessary question about the pathway to change.
Key Takeaways
- Stop asking only whether something worked. Ask what causal chain would need to be true for it to work, and test each link.
- Do not confuse activity with progress. High output can coexist with zero meaningful change.
- Use evaluation as a diagnostic tool. Look for intermediate outcomes, not just final results.
- Treat theory as provisional. If evidence contradicts your assumptions, revise the model instead of defending it.
- Build a ladder of evidence. Measure outputs, uptake, behavior change, and long-term outcomes as distinct stages.
The Real Purpose of Evaluation
At its best, evaluation is not an audit of past performance. It is a disciplined way of learning how change actually happens in the world.
That is a much larger ambition than counting deliverables or publishing endline reports. It asks organizations to become more honest about uncertainty, more precise about causality, and more humble about their own assumptions. It also gives them something more valuable than a verdict: the ability to improve the next attempt.
The deepest insight is this: impact is not a number at the end of a pipeline. It is the visible trace of a causal story that survived testing.
When you start from that premise, everything changes. Indicators become clues. Outcomes become checkpoints. Failures become data. And theory becomes useful not because it simplifies reality, but because it helps you navigate reality without lying to yourself.
That is the real partnership between theory of change and outcome evaluation. One tells you what should happen. The other tells you whether the world agrees.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣