The Hidden Contract Between Causality and Evaluation
Hatched by Anemarie Gasser
Apr 28, 2026
10 min read
3 views
91%
The question beneath every good policy
Why do so many smart programs fail to prove their value, even when the people running them are convinced they work? The usual answer is that the data were messy, the measurements were weak, or the evaluation came too late. But the deeper problem is more philosophical than technical: we often confuse seeing change with understanding what caused change.
That confusion matters everywhere. A school launches a tutoring initiative and test scores rise. A city expands a housing subsidy and homelessness edges down. A company rolls out a new onboarding process and employee retention improves. In each case, the temptation is immediate: the intervention worked. Yet the world is noisy. Trends shift, people self select, other policies intervene, and outcomes drift for reasons no one planned. If we cannot separate what happened from what made it happen, we are not really evaluating. We are narrating.
The most powerful insight connecting causal thinking and theory based evaluation is this: good judgment requires both a counterfactual and a theory. Counterfactual thinking asks, “What would have happened otherwise?” Theory based thinking asks, “Through what pathway was this supposed to work?” One without the other is brittle. Together, they give us a way to think clearly about action in a complex world.
Why results alone are an illusion
Results are seductive because they look objective. They arrive in numbers, charts, and dashboards. But outcomes, by themselves, are only the end of a story whose beginning is usually hidden. A program may succeed because of the intervention, despite the intervention, or for reasons unrelated to it. If we do not ask the causal question, we can end up rewarding coincidence and punishing genuine effort.
Consider a job training program. Participants get jobs after the program ends. That sounds like success. But maybe those participants were already unusually motivated. Maybe the local labor market improved. Maybe only the most prepared applicants enrolled. Without a comparison to what would have happened otherwise, the apparent effect may be an illusion. This is the basic discipline of causal inference: not just asking what changed, but what changed because of us.
Yet even a clean causal estimate can be empty if it is detached from the logic of the intervention. Suppose the same job training program increases employment, but we do not know whether the key ingredient was resume coaching, credentialing, peer support, or simple participation requirements. If we only know the program works, we cannot improve it intelligently, scale it safely, or adapt it to a new setting. A number without a mechanism is a verdict without a diagnosis.
This is where theory enters. A theory of change is not bureaucratic decoration. It is a map of how inputs are supposed to become outputs, and outputs become outcomes. It makes hidden assumptions explicit. It says: if we do X, then Y should happen, because Z. That “because” is where evaluation becomes intellectually serious.
A result tells you that something happened. A theory tells you what to look for next.
The real unit of evaluation is not the program, but the claim
One of the biggest mistakes in evaluation is to treat the program as the thing being tested. In practice, what we are really testing is a claim: if we intervene in a certain way, under these conditions, we expect a certain outcome. The claim has structure. It includes assumptions about people, timing, incentives, context, and measurement. If those assumptions are wrong, the evaluation may be technically correct and practically misleading.
Think of a vaccination campaign. If cases fall, one might say the campaign succeeded. But the real claim is more specific: vaccination caused a reduction in disease transmission through immune protection, at a rate that outweighed competing forces. That claim could fail in different ways. Distribution might miss the highest risk groups. The vaccine might be less effective against a new variant. Public trust might shape uptake. Each failure points to a different part of the causal chain.
A theory based evaluation works best when it treats the intervention like a machine with visible and invisible gears. The gears are not just the final outcome, but the intermediate steps: awareness, access, participation, fidelity, behavior change, and sustained effects. If one gear slips, the machine may appear broken even if the design is sound. If the machine works, we learn which gears mattered most.
This is why evaluation should not be a post hoc scorecard. It should be a disciplined investigation of a causal story. Did the intervention reach the right people? Did they respond as intended? Did the mechanism activate? Did external conditions support or block the pathway? These questions turn evaluation from an accounting exercise into a learning system.
A useful mental model is to imagine any intervention as a chain of four links:
- Exposure: Did the intended people encounter the intervention?
- Engagement: Did they actually participate or use it?
- Mechanism: Did the intervention change beliefs, incentives, knowledge, or behavior?
- Outcome: Did those changes produce the desired result?
A negative result at the end of the chain does not tell you where the breakdown occurred. A theory based approach does. That is why it is so much more useful than a simple yes or no conclusion.
Causality tells us whether something worked. Theory tells us why it worked, or why it did not.
The deepest connection between causal inference and theory based evaluation is that each corrects the blind spot of the other. Causal inference is strong on attribution. It helps estimate the effect of an intervention rather than just its correlation with outcomes. But attribution alone can become sterile if it does not illuminate process. Theory based evaluation is strong on process. It reveals how change is supposed to happen. But process alone can become wishful thinking if it is never tested against actual causation.
This pairing matters because many real world programs operate in environments where randomized trials are impossible, undesirable, or incomplete. In those settings, theory becomes the scaffold that keeps evaluation from collapsing into guesswork. But theory is not a substitute for evidence. It is the guide for where to look, what to measure, and how to interpret what you see.
Imagine a public health initiative designed to reduce smoking among teenagers. A simple outcome evaluation would ask whether smoking rates declined. A causal evaluation would ask whether the initiative caused the decline. A theory based evaluation would go further: did the campaign change peer norms, reduce access, increase perceived risk, or shift family conversations? If smoking did not fall, which link broke? If it did fall, which pathway deserves credit?
This is not just academic elegance. It has practical consequences. If the program worked because peer norms shifted, then future investments should strengthen social influence, not just deliver more pamphlets. If the program failed because teenagers ignored the message, then the problem is not the outcome metric. It is the mechanism of communication.
The best evaluation does not merely answer whether a policy worked. It identifies the exact point where the world changed.
That is an extraordinary advantage. It allows leaders to decide whether to scale, revise, segment, or abandon an intervention. It transforms evaluation from retrospective judgment into design intelligence.
A framework for thinking clearly about interventions
When people evaluate programs, they often jump too quickly to the final outcome. A better approach is to ask four questions in sequence. Together, they create a more honest and actionable view of reality.
1. What was the causal claim?
Before evaluating a program, state the claim in plain language. Not “the initiative improved performance,” but “if we provide this support to this population in this context, then outcome X should improve relative to what would otherwise have happened.” This forces clarity about the target, the comparison, and the expected direction of change.
2. What was the theory of change?
List the mechanism. What has to happen first, second, and third for the outcome to occur? What assumptions must hold? What conditions might block the pathway? The more specific the theory, the easier it becomes to test.
3. Where might the chain break?
Look for bottlenecks at each stage: reach, engagement, mechanism, and outcome. A program can fail because of weak delivery, low uptake, poor design, hostile context, or insufficient time. The point is not to blame. The point is to diagnose.
4. What would count as evidence at each stage?
Do not wait for the final outcome alone. Measure intermediate indicators that map onto the theory. If a new literacy program is supposed to improve reading by increasing practice time, measure practice time, not just test scores. If a nutrition initiative is supposed to work by changing grocery purchases, measure purchasing behavior, not just body weight.
This framework is powerful because it prevents two common errors. First, it prevents overclaiming when outcomes improve for unrelated reasons. Second, it prevents underlearning when outcomes do not improve, but the mechanism may still be sound and merely underpowered, mistimed, or poorly implemented.
A good evaluation therefore behaves less like a judge and more like a physician. A judge asks, “Did it succeed?” A physician asks, “What is happening inside the system, and where is the fault line?” That shift in posture changes everything.
The most valuable question is not “Did it work?” but “What is the smallest true story?”
Complex interventions often generate complex outcomes. It is a mistake to demand a single simple answer when the reality is multi layered. A program may help one subgroup and not another. It may work in one neighborhood, with one facilitator, during one policy climate, and fail elsewhere. A rigid binary conclusion can hide the most important lesson: effects are conditional.
That is another place where causal and theory based thinking meet. Causal inference teaches us to respect comparability and variation. Theory based evaluation teaches us to expect that different contexts activate different pathways. Together they encourage a humbler, more precise kind of knowledge.
A subsidy, for example, might not reduce hardship uniformly. For families with stable transportation, the benefit may be substantial. For families without transportation, the same subsidy may be useless because access is the real barrier. The outcome average could look modest, but the theory reveals a crucial design flaw. Perhaps the right fix is not a larger subsidy, but bundled support that solves access first.
This is why evaluation is inseparable from design. If you do not know how a program works, you cannot know how to improve it. If you do not know whether a result is causal, you cannot know whether to trust it. The real goal is not merely accountability. It is learning under uncertainty.
That phrase matters because many institutions still treat evaluation as a scorekeeping exercise. They want a definitive pass or fail. But in complex systems, the more useful habit is to seek the smallest true story that can still guide action. Sometimes that story is: the intervention had no effect. Sometimes it is: it worked only where implementation quality was high. Sometimes it is: the mechanism was right, but the timing was wrong. Each is more valuable than a vague declaration of success or failure.
Key Takeaways
- Always state the causal claim in advance. If you cannot say what comparison would convince you, you are not evaluating causally.
- Map the theory of change into measurable steps. Break the intervention into exposure, engagement, mechanism, and outcome.
- Do not treat final outcomes as self explanatory. Ask where in the chain the effect appeared, disappeared, or changed direction.
- Use intermediate indicators to diagnose mechanism. They are not substitutes for outcomes, but they are often the fastest way to learn.
- Look for conditional effects, not just averages. Programs often work in specific contexts, for specific people, through specific pathways.
Conclusion: evaluation is a way of respecting reality
The deepest lesson at the intersection of causality and theory is not methodological. It is moral. Reality is more intricate than our slogans, and people living through interventions deserve more than applause or blame based on thin evidence. They deserve judgments that are earned.
To respect reality is to admit two things at once: that outcomes can mislead, and that theories can deceive if they are never tested. The best evaluation practice refuses both naivety and cynicism. It asks what changed, why it changed, and what would have happened otherwise. That combination is rare, but it is the difference between measuring activity and understanding impact.
In the end, the purpose of evaluation is not to produce a verdict. It is to produce better actions in a world that never stops changing. When we unite causal inference with theory based evaluation, we gain something far more valuable than confidence. We gain the ability to learn, carefully and honestly, from the real world itself.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣