The Most Important Question in Evaluation Is Not What Happened, but What Would Have Happened Anyway
Hatched by Anemarie Gasser
Jul 10, 2026
9 min read
1 views
61%
The hidden trap in every outcome story
A program improved test scores. A campaign increased donations. A new workflow reduced churn. These sentences feel complete because they describe a change after an intervention. But they conceal the most important part of the story: change compared to what?
That missing comparison is the difference between sounding informed and actually knowing something. It is also the difference between an outcome that merely appears in the same timeline as an action and an outcome that can credibly be linked to it. The core challenge of evaluation is not measuring results. It is resisting the human urge to confuse sequence with causation.
This is why so many organizations drown in metrics while remaining unsure whether they have learned anything. They count what happened, track what improved, and present tidy dashboards. Yet without a disciplined way to ask what would have happened anyway, those numbers can become elegant decoration.
The hardest question in evaluation is not whether something changed. It is whether the change would have occurred without you.
That question sounds simple. In practice, it overturns almost everything people do when they try to prove impact.
Why outcome measurement so often lies politely
Most people think evaluation begins with collecting data. In reality, it begins with confronting a philosophical problem: counterfactual thinking. To evaluate an intervention, you need to imagine a world in which the intervention did not happen, then estimate what would have occurred in that world.
That sounds abstract until you look at a concrete example. Suppose a nonprofit launches a tutoring program and student grades rise afterward. Is the program responsible? Maybe. Or maybe the school also changed grading policy, a new cohort arrived, attendance improved, or the exams simply became easier. If grades rise during flu season after a tutoring launch, the timing alone tells you almost nothing.
This is where weak evaluation often goes wrong. It treats any post intervention improvement as evidence of impact. But outcomes move for many reasons at once: seasonality, regression to the mean, selection effects, economic shifts, policy changes, and plain chance. If you do not account for those forces, you may end up rewarding the wrong thing and abandoning the right thing.
A useful mental model is to think of every observed outcome as having three layers:
- The intervention effect, if any.
- Background drift, the natural movement of the system over time.
- Noise, the random fluctuations that make patterns look meaningful before they are.
Most unstructured evaluations collapse all three into one number and call it success.
The seduction of the dashboard
Dashboards are persuasive because they convert complexity into confidence. A rising line feels like progress. A falling line feels like failure. But a line is not a cause. It is only a trace of events, shaped by everything happening around the intervention.
Consider a sales team that adopts a new outreach script. Conversions rise by 15 percent the next month. The instinct is to credit the script. But what if the rise coincided with the company’s strongest prospecting quarter, a major industry event, and an unusually responsive market? The script may have helped, but the raw increase does not tell you how much.
This is why outcome evaluation cannot be reduced to outcome reporting. Reporting says, “Here is what changed.” Evaluation asks, “What changed because of us, and how do we know?” The second question is much harder, but it is the only one that prevents you from being misled by your own success story.
The real unit of analysis is the comparison you failed to make
The best evaluations are not defined by having more data. They are defined by better comparisons.
When people say they want to measure impact, they often imagine more measurement. In fact, the crucial move is design. You need a structure that approximates the missing world, the one where the intervention was absent. Sometimes that means a randomized comparison group. Sometimes it means a carefully chosen before and after window. Sometimes it means leveraging natural experiments, matched controls, interrupted time series, or other quasi experimental tools.
The specific method matters less than the principle behind it: do not compare the intervention group only to itself after the intervention. That is not a causal question. It is just a before and after story with a flattering costume.
A useful analogy is medicine. If a fever breaks after a patient takes a drug, you might feel encouraged. But fevers also break on their own. The relevant question is not whether the patient got better after the pill. It is whether they got better more than similar patients who did not get the pill, or more than they would have without it. Without that comparison, the pill gets credit for the body’s own recovery.
A practical framework: the three questions of credible evaluation
Whenever someone presents an outcome claim, ask three questions:
- Compared to what?
- What else changed at the same time?
- How plausible is the alternative explanation?
These questions sound obvious, but they are rarely answered with discipline. They are what turn raw outcomes into evidence.
The first question forces a counterfactual. The second forces context. The third forces humility. Together, they prevent a common error: over attributing causality to the intervention simply because it is the most visible thing in the story.
Evaluation is not the art of proving that something worked. It is the discipline of ruling out the easier explanations first.
This matters because organizations often make high stakes decisions from low quality causal stories. They fund programs because participants improved, cut initiatives because improvement was absent, and scale efforts based on anecdotes that ignore comparison. In each case, they mistake a narrative for an estimate.
Why “step by step” beats “big reveal” in outcome evaluation
A strong evaluation is rarely a single dramatic insight. It is a sequence of smaller, more honest steps.
That step by step mindset changes the entire practice of measurement. Instead of asking, “Did it work?” in one leap, you ask a chain of more defensible questions:
- Did the intervention reach the intended people?
- Did behavior change in the short term?
- Did the intermediate outcomes move in the expected direction?
- Could those changes plausibly be due to the intervention rather than another force?
- Did the final outcomes shift enough to matter?
This matters because programs fail in different ways. Some fail at delivery. Some succeed at delivery but fail to alter behavior. Some alter behavior but not outcomes. Some improve outcomes, but not because they caused the improvement.
A stepwise approach does two things. First, it helps locate the mechanism. Second, it reduces false certainty. If a tutoring program improves attendance but not grades, that may still be a meaningful signal. It tells you where the chain breaks. If a marketing campaign increases site visits but not conversions, the problem may not be awareness. It may be the landing page, the offer, or the audience mix.
In this sense, evaluation is not only a verdict. It is a diagnostic.
Outcomes are not the same as effects
This distinction is easy to miss and crucial to preserve. An outcome is what happened. An effect is the change attributable to an action, relative to what would have happened otherwise.
A school may see attendance rise after introducing breakfast. That is an outcome. But the effect is the portion of that rise that exceeds what would have happened without breakfast, after accounting for weather, calendar shifts, or other policy changes. When you confuse outcomes with effects, you end up attributing every improvement to intervention and every decline to failure.
That confusion also creates bad incentives. Teams learn to choose interventions that are easy to measure rather than those that are truly valuable. They optimize for visible change, not causal change. In the long run, that is a recipe for theater.
A better way to think about impact: the causal audit
The most useful synthesis of evaluation thinking is to treat every claim of impact as requiring a causal audit.
A causal audit is not a statistician’s luxury. It is a practical habit. Before celebrating or rejecting a result, audit the claim using four layers:
1. Temporal layer
Did the change occur after the intervention? If not, stop. If yes, continue, but do not confuse order with causality.
2. Comparison layer
What is the best available comparator? A control group, baseline period, matched population, or benchmark trend. No comparison means no causal estimate, only a description.
3. Mechanism layer
Through what pathway should the intervention work? If a pathway cannot be explained, measured, or at least plausibly traced, the claim is fragile.
4. Rival explanation layer
What else could explain the change? Incentives, selection, seasonality, policy shifts, measurement artifacts, and external shocks should all be considered before taking credit.
This audit is powerful because it is portable. You can apply it to public policy, product analytics, education, philanthropy, health care, and even personal decision making.
For example, imagine you start waking up earlier after buying a new alarm clock. Did the clock change your behavior, or did the new month, better sleep hygiene, and a more stable schedule do most of the work? The causal audit does not require perfection. It requires that you ask the right questions before you tell yourself a flattering story.
Why organizations resist causal thinking
Causal evaluation is uncomfortable because it threatens identity. If a program did not cause the improvement, then a team cannot simply claim victory. If a cherished initiative did not produce the promised effect, then prestige is at risk.
This is why many organizations prefer outcome theater to outcome truth. Theater is faster, simpler, and emotionally rewarding. Truth is slower and often less glamorous. But truth compounds, because it allows better decisions over time.
The deepest benefit of causal thinking is not that it produces cleaner reports. It is that it protects learning from self deception. An organization that can distinguish signal from coincidence becomes adaptive. One that cannot becomes superstitious.
Key Takeaways
- Never treat a post intervention improvement as proof of impact. Improvement is not the same as causation.
- Always ask “compared to what?” A comparison, even a rough one, is the heart of credible evaluation.
- Separate outcomes from effects. Outcomes describe what happened, effects estimate what changed because of the intervention.
- Use stepwise evaluation. Check reach, behavior change, intermediate outcomes, and final outcomes instead of jumping straight to conclusions.
- Run a causal audit before making decisions. Test temporal order, comparator quality, mechanism plausibility, and rival explanations.
The deeper lesson: humility is part of measurement
The irony of evaluation is that the more seriously you take it, the less certain you become. That is not a flaw. It is a feature.
Good evaluation does not make the world simpler. It makes your judgments more honest. It replaces the comforting illusion of direct visibility with a more rigorous form of seeing, one that admits uncertainty while still extracting useful conclusions.
That is why the central question is never just what happened. It is what would have happened anyway. Once you internalize that, every metric becomes more interesting, every success claim more testable, and every failure more informative.
In the end, the goal is not to stop telling stories. It is to tell better ones, stories that survive comparison with the world that might have been. That is the standard that turns measurement into understanding, and understanding into wiser action.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣