When You Cannot Prove Causality, Start by Counting What Changed

Anemarie Gasser

Hatched by Anemarie Gasser

Jul 08, 2026

10 min read

88%

0

The Evaluation Problem We Keep Mistaking for an Evidence Problem

Most organizations do not fail because they lack data. They fail because they ask the wrong question of the data they already have.

The habitual question is: Did our program cause the outcome? That sounds rigorous, but it often smuggles in an impossible standard. In the real world, especially in complex social systems, change rarely arrives in a neat, controlled package. Policies interact with institutions, people improvise, incentives shift, and context bends every plan. By the time a result appears, many forces have touched it.

That is where a deeper tension emerges: if causality is messy, how can judgment still be disciplined? The answer is not to lower the bar. It is to change the unit of analysis. Instead of forcing evaluation to behave like a laboratory experiment when it is really operating in a living system, the better question becomes: What changed, for whom, in what direction, and what plausible contribution did the intervention make?

That shift is more than a technical adjustment. It is a philosophy of learning. It replaces the fantasy of total attribution with a more realistic and often more useful discipline: tracing change, testing assumptions, and building a credible theory of how impact happens.


From Attribution to Contribution: A Better Question for Complex Change

Traditional evaluation often behaves as if the world were a machine. Pull a lever, measure the output, isolate the variable, repeat. This approach works well when the system is stable and the mechanism is simple. If a light bulb goes on when you flip the switch, attribution is easy.

But social change is not a light bulb. It is more like weather. Many conditions must align, and no single force fully explains the storm. A literacy program may increase reading scores, but only partly because of the curriculum. Attendance, teacher morale, parental support, local politics, and even seasonality may matter. In such settings, insisting on pure attribution can produce either false certainty or paralyzing doubt.

This is why contribution is a more mature standard than attribution. Contribution asks whether a program helped make change happen within a broader causal landscape. It does not abandon rigor. It demands a stronger kind of rigor, one that is explicit about assumptions, pathways, and context.

The question is not whether one actor owns the outcome. The question is whether the story of change is coherent, evidenced, and plausible.

That reframing matters because it changes how organizations behave. If they believe the only valuable evidence is proof of direct causation, they will avoid ambitious work, overclaim success, or reduce themselves to narrow metrics that miss the point. If they accept contribution as the goal, they can pursue a more honest practice: map their theory of change, watch for signs along the pathway, and revise their understanding as reality speaks back.


The Hidden Power of a Theory of Change: Not Prediction, but Discipline

A theory of change is often misunderstood as a presentation slide, a logic model, or a fundraising artifact. Its real value is sharper and more demanding. It is a discipline of explanation.

At its best, a theory of change forces an organization to state the chain of reasoning connecting inputs to outcomes. Why should this activity produce this intermediate shift? Why should that shift matter? Under what conditions would the chain break? Once those questions are made visible, evaluation becomes less about checking boxes and more about interrogating assumptions.

This is where many organizations make a crucial mistake. They treat evaluation as an audit of finished results instead of a conversation with reality about mechanism. But mechanism matters because outcomes do not appear by magic. If a job training program improves earnings, is it because participants gained skills, because employers changed their hiring practices, because the program created new networks, or because the local labor market was unusually strong? A theory of change gives you a map of possible answers before the numbers arrive.

The elegance of this approach is that it turns uncertainty into inquiry. Instead of asking, “Did it work?” in the abstract, you can ask:

  1. Which assumptions held?
  2. Which links in the causal chain were visible?
  3. Where did the intended pathway fail or get rerouted?
  4. What did the context amplify or suppress?

This is not a softer form of evaluation. It is often a harder one, because it refuses to let a single headline number stand in for understanding.

Think of a community health campaign. If vaccination rates rise, the program may be credited. But theory-based evaluation asks a more interesting question: did awareness increase, did trust improve, did access become easier, did misinformation decline, did local leaders endorse the message, and did those changes accumulate into uptake? Each answer reveals something different about why the result happened and whether it is likely to persist.

In that sense, theory is not decoration around data. Theory is what makes data legible.


Outcome Harvesting: When the Outcome Is the Evidence

The second half of the puzzle is a method that starts from a different premise: sometimes you do not know in advance exactly what will change, and waiting for preselected indicators may cause you to miss the most important effects.

That is the strength of outcome harvesting. Rather than beginning with a fixed outcome list and checking whether it happened, you look for changes first, then work backward to understand their significance and contribution. This is especially powerful in environments where change is emergent, political, or nonlinear. In advocacy, systems change, institution building, or adaptive programming, the most important outcomes may not be the ones you predicted.

Imagine trying to evaluate a coalition working to improve public procurement transparency. The most meaningful change might not be a final law passing, but a ministry beginning to publish bidding data, a journalist using that data to surface patterns, or a watchdog network forming relationships inside government. A rigid framework may miss those steps because they are not the end result. Outcome harvesting treats them as evidence of movement.

The method is deceptively simple in spirit and demanding in practice: gather descriptions of changes, verify them, and determine how the intervention contributed. But its deeper significance is epistemological. It says that in complex systems, the world does not always announce its important outcomes in advance.

That insight corrects a common failure of measurement culture. Organizations often design metrics before they understand the phenomenon they hope to influence. Then they optimize for what they can count instead of what actually changed. Outcome harvesting reverses the sequence. It begins with observed change, then asks what it means.

In stable systems, measurement can begin with the indicator. In living systems, measurement must begin with the outcome.

This is why outcome harvesting feels liberating to practitioners who work in volatile contexts. It validates unexpected change. A new partnership, a policy reinterpretation, a shift in discourse, a local imitation effect, all of these can matter even if they were never listed in the original workplan.

But there is a danger here too. Without a theory, harvesting can become an indiscriminate pile of anecdotes. Change alone is not understanding. The challenge is to distinguish signal from noise, and that is where theory-based evaluation enters as the necessary counterpart.


The Real Breakthrough: A Closed Loop Between Theory and Discovery

The deepest insight arises when these two approaches are combined. One provides structure, the other provides discovery. One starts with expected pathways, the other starts with observed change. Together, they create a learning loop that is much more powerful than either approach alone.

Here is the mental model:

  • Theory-based evaluation tells you what kinds of change should matter and why.
  • Outcome harvesting tells you what change actually occurred, including the unexpected.
  • Together, they let you test whether the theory was right, incomplete, or wrong.

This combination matters because most failures in evaluation are category errors. People either demand too much precision from a theory that is only a hypothesis, or they collect too much change data without a model to interpret it. The first mistake creates false confidence. The second creates descriptive chaos.

A useful analogy is navigation. A map without landmarks is hard to use, but landmarks without a map tell you little about where you are going. Theory is the map. Harvesting is the landmarks appearing on the road. If you rely only on the map, you may miss detours, shortcuts, and unexpected terrain. If you rely only on landmarks, you may know you have moved without knowing whether you are closer to your destination.

This closed loop also changes power dynamics inside organizations. Evaluation stops being a tribunal at the end of the project and becomes a practice of collective sensemaking during the work. Program staff, beneficiaries, partners, and funders can all participate in asking what changed and why. That matters because the people closest to the work often see weak signals first, while senior leaders often control the official story.

The most sophisticated organizations do not choose between planned outcomes and emergent outcomes. They build systems that can hold both. They begin with a theory strong enough to guide action, but humble enough to be revised by reality.


A Practical Framework: Three Questions That Beat the Vanity Metric

If you want a simple way to apply this synthesis, use the Three Questions Framework.

1. What did we think would change?

This is the theory side. State the pathway plainly, including the assumptions. Avoid vague language like “improve capacity” unless you can specify what capacity means in behavior, relationships, or decision making.

2. What actually changed?

This is the harvesting side. Look broadly, not just at the final target. Capture intended and unintended changes, small and large, direct and indirect.

3. Why is that change plausibly connected to us?

This is the contribution test. You are not trying to prove monopoly causation. You are asking whether your work appears to have played a meaningful role in a broader causal pattern.

This framework is useful because it forces a disciplined conversation between expectation and observation. It prevents two common traps: claiming success on the basis of activity, and dismissing real change because it was not pre-registered in a dashboard.

Consider a nonprofit that trains rural entrepreneurs. If it only tracks number of workshops delivered, it may confuse motion with impact. If it only tracks revenue at year end, it may miss the interim changes that make revenue possible, such as confidence, supplier relationships, or improved record keeping. The three questions force attention to the full chain, from hypothesis to evidence.

And once you start asking these questions regularly, something subtle happens. The organization becomes more honest. It stops pretending that uncertainty is failure. It starts treating uncertainty as information.


Key Takeaways

  • Stop asking for total attribution when the system is complex. Aim for credible contribution instead of impossible certainty.
  • Make your theory explicit before you measure results. If you cannot state the causal pathway, you cannot learn from success or failure.
  • Look for outcomes broadly, not just the ones you predicted. Important change often appears first as a shift in relationships, behavior, or norms.
  • Use evidence to test and revise assumptions, not just to report performance. The point of evaluation is learning, not only accountability.
  • Combine structure with discovery. A strong theory gives direction, and outcome harvesting keeps you honest about what is actually happening.

The Deeper Reframe: Evaluation Is Not a Verdict, It Is a Conversation With Reality

The most important shift here is philosophical. Evaluation is often treated as a verdict rendered after the fact. Did we succeed or fail? Did the intervention work or not? That framing is seductive because it feels definitive. But in complex change, definitive answers are frequently the least useful ones.

A better frame is that evaluation is an ongoing conversation with reality. The theory tells you what you hoped would happen. The harvested outcomes tell you what did happen. The gap between them is not just a scorecard. It is the place where learning lives.

That is why the most valuable evaluations do not simply pronounce judgment. They reveal mechanisms, expose blind spots, and sharpen the next iteration of action. They ask not only whether something changed, but how change travels through a system and what kinds of evidence can responsibly speak to that travel.

If you remember only one idea, let it be this: the point of evaluation is not to prove that your story was right. It is to discover the truest story of change you can tell.

Once you adopt that standard, data becomes more than reporting material. It becomes a way of seeing. And when an organization learns to see change as it actually unfolds, not as it wishes it were unfolding, it gains something far more valuable than a clean metric. It gains judgment.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣