When You Cannot Prove a Program Worked, Start by Tracing What Changed

Anemarie Gasser

Hatched by Anemarie Gasser

Jul 30, 2026

10 min read

73%

0

The strange problem with impact

What do you do when the thing you care about most, whether a policy, program, or intervention, does not leave behind a clean fingerprint?

That is the uncomfortable reality of evaluation. Real-world change is messy. People adapt, institutions respond, context shifts, and outcomes emerge from a tangle of influences that rarely cooperate with tidy before and after comparisons. Yet many evaluations still behave as if the world were a laboratory, where every effect can be isolated, measured, and neatly attributed.

The deeper question is not simply whether something worked. It is this: how do we learn responsibly when causality is partial, uncertain, and distributed across many actors and conditions?

That question changes everything. It moves evaluation away from a courtroom model, where the job is to prove or disprove guilt, and toward a detective model, where the job is to assemble enough clues to understand what happened, why it happened, and what to do next. In that shift, theory based evaluation and outcome harvesting are not competing methods. They are two answers to the same problem: how to make sense of change without pretending change is simpler than it is.

The most useful evaluation is not the one that claims perfect certainty. It is the one that can explain uncertainty well enough to improve action.


Why measurement alone is not enough

A spreadsheet can tell you whether a target was reached. It cannot tell you whether the target was the wrong one, whether the path to it created hidden harms, or whether the same approach would fail in a different neighborhood, ministry, or year. This is the first trap of conventional evaluation: it confuses counting outcomes with understanding change.

Imagine a literacy program that reports a rise in test scores. That number matters. But it does not explain whether the rise came from better teaching, more motivated students, a new exam format, increased parental support, or a temporary intervention effect. It also does not tell you what would happen if the program scaled, shifted regions, or faced budget cuts. Data without theory can become a scoreboard detached from reality.

This is why evaluation needs more than metrics. It needs a change logic, a story about how activities are supposed to lead to outcomes. A theory based approach asks evaluators to surface that story, challenge it, and revise it as evidence accumulates. It treats a program not as a black box but as a set of assumptions, links, and dependencies.

The power of that approach lies in humility. Instead of asking, “Did we cause the result?” it asks, “What chain of events seems to connect our actions to the changes we observe?” That is a more honest question, and often a more useful one.

A good theory of change works like a map. It does not eliminate the terrain, but it helps you notice where the road bends, where the bridges are missing, and where you may have mistaken a mountain for a hill. Without such a map, evaluators risk mistaking correlation for mechanism, or mechanism for guarantee.


The hidden value of following outcomes backward

If theory based evaluation begins with an explanation of how change should happen, outcome harvesting begins somewhere else: with the change itself. It looks for what changed, then works backward to ask who contributed, how, and under what conditions.

That reversal matters more than it first appears.

Many important changes do not announce themselves in advance. A coalition unexpectedly shifts a local policy. A community adopts a new practice after informal peer exchange. An advocacy effort changes how a ministry talks about an issue before any formal rule changes. These are outcomes, but they often emerge in environments too fluid for rigid indicators to capture in advance.

Outcome harvesting is well suited to such settings because it starts with the evidence of change and then reconstructs the contribution story. It treats outcomes not as the final checkpoint of a planned route, but as traces left by evolving social processes. If theory based evaluation asks, “Did our assumptions hold?” outcome harvesting asks, “What happened that we did not fully predict, and how can we verify its significance?”

This is a profound shift in posture. It acknowledges that important change is often discovered, not merely measured. In complex systems, people, networks, and institutions often adapt faster than plans can keep up. Trying to force every outcome into a prewritten framework can blind evaluators to the most important developments precisely because those developments were not expected.

Think of a gardener. A theory based approach is like planting with a design in mind: where the beds are, what will be grown, how water should flow. Outcome harvesting is like walking the garden later and noticing which plants actually took root, which self-seeded, which were helped by bees, wind, and weather, and which gaps matter more than the original plan. One approach is not more “scientific” than the other. Each is seeing a different part of reality.


The real tension: prediction versus discovery

The deepest tension connecting these approaches is not methodological. It is epistemological. It is the tension between prediction and discovery.

Prediction says: if we design the right intervention, specify the right chain of logic, and track the right indicators, we can anticipate change well enough to manage it. Discovery says: the world is adaptive, surprising, and often path dependent, so we must remain open to outcomes we did not foresee.

Most failures in evaluation happen when one of these dominates the other.

If prediction dominates completely, evaluation becomes brittle. It can only recognize what it already expected. This is dangerous in policy and development work, where the most meaningful effects are often indirect, delayed, or contested. A program can appear ineffective on paper while quietly reshaping norms, networks, or institutional routines that only later produce visible results.

If discovery dominates completely, evaluation can become impressionistic. It may collect interesting stories but struggle to distinguish signal from noise. Not every reported change matters. Not every contribution story is credible. Without a theory of how change should work, you risk mistaking anecdote for evidence.

The synthesis is to treat theory and harvesting as complementary lenses. Theory gives you direction. Harvesting gives you sensitivity. Theory tells you what kind of change would count. Harvesting tells you what change actually occurred. Theory helps explain plausibility. Harvesting helps reveal surprise.

A useful analogy is aviation. A flight plan is indispensable, but a pilot also needs instruments, weather reports, and the ability to respond when turbulence appears. A theory of change is the flight plan. Outcome harvesting is the instrument panel that shows you what the world is doing, not just what you hoped it would do. Good evaluation requires both.

The goal is not to choose between a plan and reality. The goal is to build a learning system where each corrects the other.


A practical framework: from intended logic to observed change

The most powerful way to combine these perspectives is to use a two stage learning cycle.

Stage 1: State the intended logic clearly

Before launching a program, articulate the assumptions that must be true for success.

Ask:

  1. What changes first, second, and third?
  2. Which actors need to shift behavior, incentives, or relationships?
  3. What context must remain stable, and what context can vary?
  4. What would count as early evidence that the pathway is working?

This is not bureaucratic paperwork. It is a discipline of honesty. If a job training initiative assumes employers will hire participants after a short course, say so. If a public health campaign assumes trust in local messengers, name that assumption. The clearer the logic, the easier it becomes to learn from reality.

Stage 2: Harvest outcomes as they emerge

Then, instead of waiting only for endline indicators, continuously scan for meaningful change.

Ask:

  1. What changed in behavior, policy, relationships, or discourse?
  2. Who noticed the change first, and how was it recognized?
  3. What evidence can verify it?
  4. What contribution did the intervention plausibly make?
  5. What else may have influenced the result?

This creates a feedback loop. The theory tells you what to look for. The observed outcomes tell you whether the theory is holding, failing, or only partly true. Over time, the program becomes less of a static plan and more of a learning organism.

Here is a simple example. A civic organization launches a campaign to increase youth participation in local meetings. The original theory predicts attendance will rise if transport barriers are reduced and meetings are made more accessible. Three months later, attendance does rise, but not mainly because of transport. The surprising outcome is that young participants began bringing peers after a few respected local figures publicly endorsed their role.

A narrow evaluation might say the program succeeded because attendance increased. A better evaluation asks a deeper question: what actually changed in the social environment that made participation feel legitimate? That answer is more transferable than the raw number. It helps the next city understand that access was necessary, but social permission was the real catalyst.

This is the kind of insight that only appears when planned theory and emergent outcomes are read together.


What evaluation becomes when it is taken seriously

When these approaches are combined well, evaluation stops being a postmortem and becomes a form of institutional intelligence.

That phrase matters. Intelligence is not just information. It is information organized for judgment. It helps organizations distinguish noise from signal, success from luck, and durable change from fleeting fluctuation. A good evaluation system does not simply report what happened. It helps leaders decide what to do next.

This has three important consequences.

First, it changes how organizations treat uncertainty. Instead of fearing it, they can work with it. Uncertainty becomes something to map, not eliminate. That mindset produces better learning environments because it reduces the pressure to manufacture certainty where none exists.

Second, it changes how success is understood. Success is not only achieving a target. It is also refining the causal story behind that target. A program that misses a numerical goal but discovers a powerful pathway to influence may be more valuable than one that hits a target for the wrong reasons.

Third, it changes what counts as evidence. Numbers remain important, but so do verified narratives, stakeholder accounts, contextual shifts, and unexpected patterns. The point is not to abandon rigor. It is to widen rigor so it can handle real complexity.

A useful mental model here is the difference between a photograph and a documentary. A photograph captures a moment. A documentary shows motion, context, and sequence. Many evaluations act like photographs. The best ones behave more like documentaries, because change itself is a story unfolding over time.


Key Takeaways

  1. Do not confuse outcomes with explanation. A positive result is not enough. Ask how the result came about and what conditions made it possible.

  2. Write down your theory of change in plain language. If you cannot explain the causal logic clearly, you cannot test or improve it.

  3. Scan for unexpected outcomes, not just planned indicators. The most important changes are often those nobody predicted at the start.

  4. Use both direction and discovery. Theory helps you know what should matter. Outcome harvesting helps you notice what actually mattered.

  5. Treat evaluation as a learning loop. Revisit assumptions, verify outcomes, and update your understanding continuously rather than waiting for a final verdict.


The better question to ask

The biggest mistake in evaluation is to ask a question that is too small. Did it work? Was it effective? Did we meet the target? These are not bad questions, but they are incomplete.

The better question is: what changed, through what pathway, and what does that teach us about influencing future change?

That question is more demanding because it refuses the comfort of simple causality. But it is also more useful because it respects the actual shape of the world. Change is rarely linear. It is negotiated, delayed, distributed, and occasionally surprising. The best evaluation methods do not deny that fact. They turn it into an advantage.

In the end, theory based evaluation and outcome harvesting are not just tools for programs. They are tools for thinking. They teach a larger lesson: when the world is complex, the goal is not to predict everything in advance. The goal is to build a disciplined way of noticing, interpreting, and learning from change as it unfolds.

That is a more mature idea of evidence. And in many settings, it is the only kind worth trusting.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣