When Programs Stop Pretending They Caused Success: The Hidden Logic of Realistic Evaluation

Anemarie Gasser

Hatched by Anemarie Gasser

Apr 23, 2026

9 min read

74%

0

The uncomfortable question behind every successful program

What if the most important question in evaluation is not Did it work?, but What exactly worked, for whom, in what circumstances, and why did it matter here instead of somewhere else?

That shift sounds subtle. It is not. It changes evaluation from a courtroom verdict into a learning system. Instead of treating an intervention like a machine with a single output, it treats it like a set of interacting mechanisms moving through a living context. The result is less certainty, but far more usefulness.

This matters because many programs do produce real benefits, yet those benefits rarely arrive in a clean, repeatable, one size fits all way. The same training, policy, or service can thrive in one place and stall in another. A program can look successful on paper and still fail to travel, because what actually drove the change was not the logo on the initiative, but the hidden combination of relationships, incentives, timing, trust, and local constraints.

The deeper tension, then, is this: we want simple judgments, but the world rewards causal humility.


Why the usual success story is too flat

Traditional evaluation often asks a deceptively narrow question. If outcomes improved, we conclude the intervention worked. If they did not, we conclude it failed. That logic is attractive because it is crisp, comparable, and easy to communicate. But it collapses the complexity of real change into a single yes or no.

Imagine a school launches a reading program. Test scores rise. Did the program cause the rise? Maybe. But maybe the rise came from strong teacher leadership, a new principal, extra parent engagement, or simply the fact that the school also began targeting students who were already on the cusp of improvement. If we stop at the headline result, we learn almost nothing that helps the next school.

This is where causal pathways become essential. A result is not just an endpoint. It is the visible trace of a chain. Resources trigger activities, activities alter behaviors, behaviors interact with local conditions, and those conditions either support or block the intended outcome. The chain matters because change is rarely produced by one factor alone. It is produced by a pattern.

A program does not succeed because it exists. It succeeds when its mechanisms are activated by the right context.

That single sentence reframes the whole game. The practical question is no longer whether a program is universally effective, but which components are doing the work, and under what conditions they can be expected to do it again.


The missing middle: how change actually travels

The most useful way to think about evaluation is as a search for the missing middle between action and outcome. That middle is where mechanisms live.

A mechanism is not just a step in a linear plan. It is the process by which people respond to an intervention. For example, a cash transfer may work partly because it increases purchasing power, but also because it reduces anxiety, improves decision making, and creates room for families to invest in school attendance. A mentorship program may work not because advice is inherently magical, but because it increases identity, accountability, and belonging. Those are very different causal channels, and they may operate differently in different settings.

This is why a useful evaluation cannot merely count outputs. Counting workshops, visits, or participants tells us what happened, not how change happened. If a job program places 200 people, the number alone does not tell us whether placement was driven by new skills, employer relationships, motivation, or local labor shortages. Without the pathway, we have a story with the most important pages torn out.

A practical way to think about this is to treat every intervention like a chain with four links:

  1. Context: the environment before the intervention begins.
  2. Mechanism: what the intervention changes in perception, behavior, or incentives.
  3. Action: the observable activities that follow.
  4. Outcome: the final result that matters.

If results differ across settings, the first instinct should not be to say the intervention is broken. The first instinct should be to ask which link broke, or whether the chain was never assembled in the same way.

This is the core insight that makes evaluation more useful in the real world: causality is not a single force, it is a sequence of conditions.


Contribution is not a weaker truth, it is the right kind of truth

Many people resist this approach because they think it gives up on rigor. If we cannot prove that an intervention alone caused the outcome, are we left with guesswork? Not at all. We are left with a more honest and often more actionable claim: the intervention contributed to the outcome through identifiable pathways, alongside other influences.

That distinction matters. In complex settings, especially public policy, development, health, education, and organizational change, singular attribution is often a fantasy. Real life is crowded. Multiple actors are moving at once. People adapt. Systems respond. A policy can be necessary without being sufficient. It can be catalytic without being fully controlling.

Contribution analysis gives language to that reality. Instead of demanding impossible purity, it asks whether the evidence supports a plausible causal chain, whether alternative explanations have been considered, and whether the observed changes fit the theory of change. This is not a downgrade from causality. It is a smarter version of causality for complex systems.

Think of it like weather forecasting. A meteorologist does not need absolute certainty to make a useful prediction. She looks at pressure systems, humidity, wind patterns, and historical behavior. She develops probability, not prophecy. Evaluation should aspire to the same discipline. It should aim for causal confidence, not causal fantasy.

In complex systems, the question is rarely “What caused this alone?” The better question is “What combination of factors made this outcome likely?”

This is especially valuable when a project faces scrutiny from multiple sides. Funders want accountability. Implementers want learning. Communities want relevance. Contribution thinking can satisfy all three, because it avoids both naive triumphalism and lazy skepticism.


The real payoff: evaluation becomes a design tool, not just a report card

The biggest mistake is to treat evaluation as something that happens after the fact, when the program is over and the score is in. If causal pathways are the focus, evaluation becomes part of design from the start.

That means every intervention should begin with a sharp theory of change, not as a decorative diagram, but as a working hypothesis. What exact mechanism is supposed to produce change? Which assumptions must hold? What context is required? What would count as evidence that the mechanism actually fired?

For example, suppose a nonprofit wants to increase vaccine uptake. A shallow theory of change might say: run awareness campaigns, then uptake rises. A deeper one asks whether the barrier is awareness, trust, access, misinformation, transport costs, or previous negative experiences with institutions. Each barrier implies a different mechanism. If the real blocker is mistrust, more information alone may do little. If the blocker is access, then better information without easier appointments is almost useless.

This is why an evaluation that traces contribution is often more useful than one that only measures final outcomes. It tells you what to change when the outcome disappoints. Did the mechanism fail? Did the context shift? Did the intended chain never activate? That diagnostic power is what turns evaluation into a design asset.

A useful mental model here is the difference between a thermostat and a detective. A thermostat tells you whether the room is too hot or too cold. A detective asks how the temperature got there, who changed the setting, whether the window is open, and whether the furnace is broken. Most organizations need detective style evaluation, because their problems are not stable enough for thermostat style thinking.

The best programs are not those that claim universal causation. They are those that learn how their causation behaves.


A practical framework for thinking in pathways

If you want to use this approach without drowning in complexity, try asking five questions after any intervention or policy effort:

  1. What was expected to change first? Identify the earliest visible shift, not just the final target.

  2. What mechanism would explain that shift? Name the human or system response that should connect the intervention to action.

  3. What context conditions were necessary? Ask what had to already be true for the mechanism to work.

  4. What else could have produced the same outcome? Consider alternative explanations early, not as an afterthought.

  5. What would we expect to see if the theory were right? Translate the causal story into observable clues.

This framework has a powerful side effect. It forces teams to stop speaking in abstractions. Instead of saying “the program improved engagement,” they have to specify whether engagement rose because meetings became more convenient, managers became more responsive, participants trusted the process, or incentives changed. The story gets harder to tell, but easier to use.

That is the tradeoff worth making.


Key Takeaways

  • Stop asking only whether something worked. Ask what mechanism produced the change, and whether that mechanism is likely to hold in other settings.
  • Treat context as causal, not decorative. Conditions such as trust, timing, power, incentives, and institutional capacity are part of the explanation, not background noise.
  • Prefer contribution over false certainty. In complex environments, proving sole causation is often impossible and less useful than building a strong case for contribution.
  • Use evaluation as design intelligence. A good causal story tells you how to improve the intervention, not just how to judge it.
  • Look for the first link in the chain. Early signs of change often reveal whether the mechanism is alive long before final outcomes appear.

The deeper lesson: outcomes are the surface, pathways are the truth

The temptation in evaluation is always to climb straight to the result. Did the poverty rate drop? Did the students improve? Did the adoption increase? Those questions matter, but they are the last questions, not the first ones. If we do not understand the pathway, the outcome becomes a fact without a lesson.

The more mature stance is to accept that real-world change is always partly collective, partly contextual, and partly contingent. That does not make causality weaker. It makes it truer. It means we stop pretending that programs are isolated levers and start seeing them as interventions inside living systems.

That shift is more than methodological. It is philosophical. It says that the world is not a scoreboard waiting for a final verdict. It is a network of relationships in motion, and the job of evaluation is to learn how change actually moves through it.

Once you see that, you stop asking, “Did the program cause success?” and start asking a better question: What had to happen for success to become possible here?

That is the question worth keeping.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣