Why Good Evaluation Starts by Admitting You Cannot See Causality Directly
Hatched by Anemarie Gasser
Jun 19, 2026
10 min read
2 views
92%
The uncomfortable truth behind every dashboard
What if the most dangerous thing in evaluation is not bad data, but a story that feels too complete?
Organizations love tidy answers. A program launched, numbers moved, success was declared. Yet the movement of a metric is not the same thing as understanding why it moved. A rise in attendance may reflect better outreach, a warmer season, or a competing event being canceled. A drop in readmissions may signal better care, or it may simply mean sicker patients stopped showing up. The world is full of correlations that look like explanations until you ask the hardest question: what would have happened otherwise?
That question sits at the center of every serious attempt to evaluate an intervention. But there is another question hiding beside it, equally important and often neglected: through what chain of events did the change happen? Not just whether a policy worked, but how, for whom, and under what conditions. This is where causal inference and theory based evaluation meet. One gives us discipline about effects. The other gives us structure for meaning.
Together they point to a deeper idea: you cannot evaluate change responsibly unless you treat causality as both a measurement problem and a theory problem.
The seduction of outcomes without mechanism
It is easy to become impressed by outcomes. A literacy intervention raises test scores. A workforce program increases earnings. A public health campaign reduces smoking. These are useful findings, but they can become intellectually lazy if we stop there.
Imagine two cities launch similar anti-litter campaigns. City A sees cleaner streets within three months. City B sees almost no change. A superficial evaluator might conclude the campaign is effective in A and ineffective in B. But suppose A paired posters with more trash bins and better enforcement, while B relied only on slogans. The real lesson is not merely that one campaign worked and the other failed. It is that the mechanism mattered. The same label covered two different causal systems.
This is the central limitation of outcome fetishism: it treats programs as if they were pills. But most interventions are not pills. They are systems of incentives, messages, relationships, institutions, and timing. Their effects depend on how people receive them, how implementers interpret them, and what else is happening in the environment.
A result is not an explanation. It is only the beginning of one.
Causal inference helps guard against false certainty by forcing us to ask what comparison would make the effect credible. Theory based evaluation adds another guardrail by insisting that the causal path itself be articulated. When used together, they transform evaluation from a scorekeeping exercise into a disciplined inquiry into how change is produced.
Two blind spots: the missing counterfactual and the missing mechanism
Think of evaluation as trying to answer two questions at once.
- Did the intervention make a difference?
- How did that difference come about?
The first question needs a counterfactual. Without some estimate of what would have happened in the intervention’s absence, it is almost impossible to attribute change confidently. That is why before after comparisons are so fragile. The world changes on its own, and interventions rarely arrive in a vacuum.
The second question needs a theory of change. Without a plausible account of how activities are supposed to lead to outcomes, evaluation becomes an archaeology of numbers without a map. You may discover that something changed, but not whether the change came from the program, the context, or a coincidence.
These two blind spots are distinct. A study can be causal in a narrow statistical sense and still be conceptually empty. For example, a randomized trial may show that an SMS reminder increases medication adherence. Great. But if the trial tells us nothing about whether the effect came from reduced forgetfulness, social accountability, or easier scheduling, then scaling the intervention becomes guesswork. Conversely, a richly described theory of change can feel convincing and still be wrong if it is not tested against evidence that rules out alternative explanations.
The deepest mistake is to treat these as competing approaches. They are complementary disciplines of doubt.
One disciplines the question, “Did it work?” The other disciplines the question, “Why should we believe that?”
A practical mental model: evaluation as a chain, not a snapshot
A useful way to combine these ideas is to imagine every intervention as a causal chain with five links:
- Inputs: money, staff, tools, time
- Activities: what the program does
- Mediators: the immediate changes in behavior, perception, access, or capability
- Outcomes: the desired end states
- Context: the conditions that can strengthen, weaken, or redirect the chain
This model matters because many evaluations focus only on the last link. But if the chain breaks, the program may fail for reasons unrelated to the basic idea. Maybe the intervention is sound, but the activity was delivered inconsistently. Maybe the activity was delivered well, but the mediator never changed. Maybe the mediator changed, but the outcome was blocked by context.
For example, consider a job training program. A superficial evaluation might ask whether participants found jobs. A causal evaluation asks whether the training increased the probability of employment compared with similar people who did not receive it. A theory based evaluation asks where the chain could fail: Did participants acquire the right skills? Did employers recognize those skills? Did the local labor market have openings? Were there transportation barriers? Were childcare constraints ignored?
This is not mere elaboration. It is the difference between knowing that a bridge collapsed and knowing which beam failed under what load.
The most useful evaluations therefore do two things at once. They estimate an effect, and they map the pathway. The effect tells us whether the intervention matters. The pathway tells us whether the effect is robust, portable, and worth trusting.
The real tension: rigor versus relevance is a false choice
A common misunderstanding is that causal rigor and theory driven nuance pull in opposite directions. In practice, the opposite is often true. Rigor without theory can be precise but shallow. Theory without rigorous comparison can be insightful but untrustworthy.
This tension shows up everywhere.
A hospital introduces a discharge follow up call to reduce readmissions. The numbers improve. Did the calls work because patients felt cared for, because they were reminded to take medications, or because staff identified complications early? If the next hospital copies the practice but sees no benefit, was the intervention weak, or was the context different? Without theory, the statistic cannot travel. Without causal comparison, the theory cannot be tested.
Or take a school district that adopts a new reading curriculum. Test scores rise after implementation. A naive interpretation is that the curriculum succeeded. But perhaps the improvement was caused by teacher training, smaller class sizes, or a cohort effect. A more serious evaluator asks whether the curriculum changed classroom practice, whether teachers adopted the new method, and whether the learning gains were strongest where implementation fidelity was highest. Now the evaluation becomes useful for decision making, not just reporting.
This is where the synthesis becomes powerful: causal inference protects us from mistaking coincidence for impact, while theory based evaluation protects us from mistaking impact for understanding.
That distinction matters because institutions do not just need evidence. They need actionable evidence. Actionable evidence tells leaders what to continue, what to adjust, what to scale, and what to abandon. You cannot get that from outcomes alone, and you cannot get it from theory alone. You need both the effect and the architecture of the effect.
A better question than “Did it work?”
The most mature evaluators eventually ask a harder question than success or failure. They ask:
What exactly worked, for whom, under what conditions, and through which pathway?
This phrasing changes everything. It stops us from treating programs as monoliths. It also respects the fact that interventions are rarely uniformly effective. A mentorship program may help first generation students more than others. A payment incentive may improve one clinic’s behavior but not another’s. A policy may succeed in urban areas and stall in rural ones. Those are not nuisance findings. They are the substance of reality.
This is where theory becomes indispensable. A theory of change is not just a diagram for funders. Done well, it is a map of assumptions. It says which steps must occur, which conditions must hold, and where failure is likely. That map lets evaluators design better comparisons and interpret results with greater precision.
At the same time, causal inference keeps theory honest. It asks whether the observed pattern is actually different from what would have happened anyway. In other words, it converts a hopeful narrative into a testable claim.
The best evaluation is neither a verdict nor a story. It is a disciplined explanation.
That phrase matters because organizations often behave as if evidence exists to issue a final judgment. But mature evaluation is less like a court sentence and more like a diagnostic process. The goal is not simply to pronounce the intervention good or bad. The goal is to understand the system well enough to make better choices next time.
How to think like an evaluator, even if you are not one
You do not need to be a statistician or a program analyst to use this synthesis. You only need to start asking your questions in the right order.
First, identify the counterfactual instinct. Whenever someone claims success, ask: compared to what? If the comparison is vague, the claim is weak. If the comparison is explicit, you can begin to assess credibility.
Second, identify the mechanism instinct. Whenever someone claims an intervention worked, ask: through what intermediate change? If the answer is “people changed behavior,” press further. What behavior changed, and why would that lead to the final outcome?
Third, identify the context instinct. Ask where the intervention is likely to travel and where it is not. A policy that depends on stable staffing, high trust, or strong administrative capacity may fail in settings that lack those features. Knowing that in advance is not pessimism. It is intelligence.
Fourth, identify the fidelity instinct. Did the intervention happen as designed? Many failed programs were not truly tested because implementation drifted from the original plan. An evaluation that ignores fidelity risks blaming the idea for the execution.
Fifth, identify the adaptation instinct. If the intervention succeeded, what can be safely modified without breaking the causal chain? This is where theory becomes strategic. It helps distinguish the essential ingredients from the incidental ones.
A simple test: if you cannot explain how the intervention should work in one paragraph, you probably do not yet understand what you are trying to evaluate.
Key Takeaways
- Do not confuse outcomes with causality. A changed metric is not proof that a program caused the change.
- Do not confuse theory with evidence. A plausible mechanism is not enough unless it is tested against a credible comparison.
- Evaluate both the effect and the pathway. Ask not only whether something worked, but how the change happened.
- Treat context as part of the causal story. Interventions often depend on implementation quality, incentives, and local conditions.
- Use evaluation to learn, not just to judge. The most valuable evidence helps you decide what to scale, adapt, or abandon.
The deeper lesson: causality is not hidden, it is constructed
The world does not hand us causality in a neat package. We construct our understanding of it by combining comparison, theory, and disciplined skepticism. That is why the best evaluations are not just technically correct. They are intellectually humble.
This humility is not weakness. It is a strength. It acknowledges that programs operate in living systems where causes interact, effects vary, and success can be real for reasons that are not immediately visible. It also acknowledges that a beautiful theory without a hard comparison can mislead, while a sharp estimate without a mechanism can remain strangely useless.
So the next time a result lands on your desk, resist the urge to ask only whether it is positive or negative. Ask instead what chain of events produced it, what alternative explanations remain, and whether the evidence is strong enough to rule them out.
That shift in posture changes evaluation from a retrospective report into a way of thinking.
And once you start thinking that way, you notice something profound: the real subject of evaluation is not the program alone, but the relationship between action, context, and change. That is where insight lives. Not in the number by itself, and not in the narrative by itself, but in the disciplined meeting of both.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣