Why Good Evaluations Fail Without a Theory of How Change Happens

Anemarie Gasser

Hatched by Anemarie Gasser

Apr 30, 2026

9 min read

38%

0

The hidden problem in measuring success

Most evaluations fail for a strange reason: they ask whether something worked before they fully understand how it was supposed to work. That sounds obvious until you notice how often organizations, governments, and nonprofits skip the crucial middle step. They gather outcomes, compare before and after, and then pronounce judgment, while the actual chain of events remains a black box.

That black box matters more than most people think. A program can improve a metric for the wrong reason, fail on the surface while still producing valuable intermediate change, or appear successful only because of outside forces. If you do not know the mechanism, you do not really know what you have learned. You have a score, not an explanation.

This is the deeper tension connecting evaluation and process tracing: measurement tells you whether change happened, but explanation tells you whether the change was caused by the thing you changed. The first is useful. The second is indispensable.

A result without a mechanism is a headline. A result with a mechanism is a lesson.

The difference between those two determines whether an evaluation becomes a report that sits on a shelf or a tool that improves future decisions.


Outcome is the destination, process is the map

Imagine you run a city program designed to reduce unemployment. Six months later, joblessness drops. Celebration follows. But what if the decline came from a booming local industry, a seasonal hiring cycle, or a change in federal policy? The outcome is real, but the attribution is uncertain. You know the destination changed, but you still do not know whether your program moved the vehicle.

Now imagine the reverse. The unemployment rate barely changes, but the program successfully helps participants build resumes, complete training, and get interviews. The final number is disappointing, yet the internal machinery is working. Maybe the market is saturated. Maybe the time horizon is too short. Maybe the intervention is one step away from producing visible effects.

This is why process tracing is so powerful. It treats causality not as a leap from input to output, but as a sequence of observable links. Instead of asking only, “Did it work?” it asks, “What had to happen in between for it to work?” That shift changes everything.

A useful analogy is a detective story. If you arrive at a scene and see that the window is broken, you have evidence. But you do not yet know whether the break happened before the theft, after the theft, accidentally, or as part of a different event entirely. Good evaluation works the same way. It does not stop at the broken window. It reconstructs the sequence.

In practice, that sequence can include things like awareness, participation, behavior change, institutional response, and downstream effects. Each link is a chance to test whether the causal story is real. If one link is missing, the theory weakens. If several links appear in the right order, confidence grows.

The beauty of this approach is that it respects complexity without surrendering to it. Many people treat complexity as an excuse for vagueness. Process tracing does the opposite. It says complexity is precisely why you need a disciplined way to follow the trail.


The real question is not “Did it work?” but “What would have to be true for it to work?”

This is the conceptual pivot that transforms evaluation from bookkeeping into reasoning. Outcome evaluation is often framed as a verdict, but a better frame is hypothesis testing. Every intervention implies a theory of change, even if that theory is informal or invisible. The real job of evaluation is to pressure test that theory against evidence.

Suppose a mentoring program claims to improve student outcomes. The simple outcome question is whether grades or graduation rates rise. The process question is more revealing: Did students show up regularly? Did they build trust with mentors? Did the mentors influence study habits, planning, or confidence? Did those shifts precede better performance?

Now the evaluation becomes more than a scorecard. It becomes a search for the causal pathway.

This matters because causal pathways are where design lives. If the intervention fails, the pathway tells you where. Maybe recruitment was weak. Maybe the dosage was too low. Maybe the mechanism only helps a subset of participants. Without that information, leaders tend to do one of two things: abandon the program too quickly or scale it too confidently. Both mistakes are expensive.

There is also a deeper intellectual benefit. Process tracing disciplines our tendency to confuse correlation with causation. A positive outcome may coexist with a weak mechanism, and a disappointing outcome may coexist with a strong one. The causal story helps us decide whether to refine, repeat, or retire an intervention.

Good evaluation is not just an accounting exercise. It is a mechanism audit.

That phrase captures a crucial shift. An accounting exercise records what happened. A mechanism audit asks whether the moving parts actually moved in the intended sequence.


A practical framework: three layers of confidence

One way to unify outcome evaluation and process tracing is to think in terms of three layers of confidence.

1. Outcome confidence

This is the simplest layer: did the metric move? If the target is employment, graduation, disease incidence, or voter turnout, you need to know whether the needle shifted. This is where many evaluations stop, but it should be the beginning.

2. Mechanism confidence

Did the expected intermediate steps occur? Did participants receive the intervention, understand it, engage with it, and change behavior in a way that the theory predicted? This layer reveals whether the program was implemented in a way that could plausibly cause the outcome.

3. Attribution confidence

If both the outcome and the mechanism line up, how much confidence do we have that the intervention, rather than other forces, caused the change? Here the evaluator looks for alternative explanations, rival pathways, and contextual factors.

These layers matter because they separate different kinds of uncertainty. A program might have high outcome confidence and low mechanism confidence, which means something changed but we do not know why. It might have high mechanism confidence and low outcome confidence, which means the intervention likely worked internally but was blocked externally. Or it might have moderate confidence in both, enough to justify iteration rather than grand conclusions.

This framework is especially useful because it prevents a common error: treating all evidence as if it answers the same question. It does not. A line on a chart, a participant interview, and a policy timeline each answer different parts of the causal puzzle. Together, they create a more trustworthy picture than any one could alone.

Consider a public health campaign encouraging vaccination. Outcome data might show rising vaccination rates. Process tracing might reveal that uptake increased only after schools, employers, and local clinics aligned their messaging. That means the campaign was not a solo cause. It was a catalyst inside a broader system. The lesson is not “the campaign worked” in isolation. The lesson is “the campaign worked because a specific sequence of social and institutional conditions activated its mechanism.”

That is a much more useful conclusion for future policy design.


Why mechanism beats myth

Organizations love simple success stories. They are easy to tell, easy to fund, and easy to repeat. But simple stories often become myths when they hide the real causal structure.

A myth says: we launched a program and the outcome improved, therefore the program worked. A mechanism says: we launched a program, specific groups changed behavior, those changes triggered intermediate effects, and under these conditions the final outcome improved. The first story flatters intuition. The second builds knowledge.

This difference matters because myths scale badly. A program that works in one setting may fail elsewhere if its mechanism depends on trust, timing, infrastructure, or incentives that do not travel. Mechanism-based evaluation helps distinguish between a portable intervention and a locally dependent one.

Think of a seed. The seed is not the whole story. Soil, water, sunlight, and temperature all matter. If you only measure whether a plant grew, you may miss the fact that the same seed fails in a different environment. If you trace the process, you learn what conditions are necessary for growth. That is knowledge you can reuse.

In that sense, process tracing is not merely a research method. It is a way of avoiding false universals. It asks: what exactly made the effect possible here, and what would need to be present elsewhere for the same effect to happen again?

This makes evaluation far more strategic. Instead of asking whether a program is good in the abstract, you begin asking under what conditions it is effective, for whom, and through which sequence of change. Those are the questions that support smarter scaling and better adaptation.


Key Takeaways

  1. Do not treat outcomes as explanations. A change in the metric is evidence, but it is not yet a causal story.

  2. Map the mechanism before you celebrate the result. Ask what intermediate steps must occur for the intervention to work, then look for evidence of those steps.

  3. Separate outcome confidence from attribution confidence. A good result can have a weak causal story, and a weak result can still contain a strong mechanism.

  4. Use rival explanations as a strength, not a threat. Alternative causes help you test whether the intended pathway is genuinely necessary.

  5. Evaluate for reuse, not just for judgment. The best evaluation produces insight that can improve future design, not just a yes or no verdict.


From verdict to learning

The deepest value of combining outcome evaluation with process tracing is that it changes the purpose of evaluation itself. The goal is not merely to determine whether something should be praised or punished. The goal is to learn how change actually happens in a messy world.

That shift has a moral dimension as well as an analytical one. When we stop at outcomes, we risk rewarding luck and punishing effort that has not yet had time to bear fruit. When we trace processes, we become more fair, more precise, and more capable of improvement. We stop mistaking the visible end of a story for the whole story.

The broader lesson is this: causality is not a shortcut from action to result. It is a path. If you want to understand whether your interventions matter, you have to walk that path, step by step, checking each link as you go.

So the next time a dashboard tells you something worked, do not stop there. Ask what had to happen for that number to move. Ask which part of the chain was strong, which was fragile, and which forces you did not account for. The most valuable evaluations are not the ones that deliver the cleanest verdict. They are the ones that reveal the shape of reality well enough to help you change it.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣