When You Cannot Predict the Outcome, You Must Prove the Trail

Anemarie Gasser

Hatched by Anemarie Gasser

May 09, 2026

9 min read

72%

0

The hidden problem with evaluation: success is often visible only after the fact

What do you do when the most important results of a project are not the ones you planned, the ones you can measure neatly, or the ones that fit inside a dashboard? This is the quiet crisis in evaluation: the world keeps generating outcomes that are real, consequential, and messy, while our evidence systems often insist on tidy inputs, predefined indicators, and linear causality.

That tension reveals a deeper question: is evaluation supposed to confirm a plan, or help us discover what actually changed? If the answer is the first, then transparency and reproducibility are the main virtues. If the answer is the second, then the evaluator needs something more like a field naturalist than a laboratory technician, someone who can notice unexpected traces, reconstruct chains of change, and remain honest about uncertainty.

The most useful way to think about this is not as a choice between rigor and flexibility. It is a choice between two logics of truth. One logic says, “I knew what I was looking for, and I can prove I found it.” The other says, “I did not know in advance what mattered, but I can still show how I learned what mattered.” The first is the logic of reproducibility. The second is the logic of outcome harvesting. The real challenge in modern evaluation is to make them work together.

The trap of measuring only what was already imagined

Traditional evaluation practices reward foresight. Before a project begins, we specify indicators, targets, and methods. This makes sense in stable environments where the path from intervention to outcome is fairly predictable. But many real-world efforts do not behave that way. Policy change, community organizing, institutional reform, behavioral shifts, and systems interventions often produce effects that are indirect, delayed, or entirely unanticipated.

Consider a public health campaign intended to improve vaccination rates. A conventional logic might track uptake, attendance, and awareness. Useful, yes. But what if the campaign also changes local trust in health workers, creates informal peer educators, or triggers a backlash that later reshapes political alliances? Those changes may matter more than the original indicators, yet they can remain invisible if the evaluation framework only sees what it planned to see.

This is where the deeper weakness of many evidence systems appears. They can be highly precise about the wrong question. They can reproduce a method faithfully while missing the phenomenon it was meant to understand.

A perfectly reproducible measurement of the wrong outcome is still a failure of understanding.

The problem is not reproducibility itself. The problem is that reproducibility is often treated as if it were synonymous with relevance. But when change is emergent, the evaluation task is not just to verify a predefined signal. It is to detect signals that only become legible after they occur.

Outcome harvesting as the discipline of seeing backward without cheating

Outcome harvesting starts from a simple but radical premise: if you cannot always predict change, you can still document it rigorously after it appears. Instead of beginning with only expected outcomes, the evaluator looks for evidence of what changed, who changed, and why that change plausibly matters. The process is retrospective, but not casual. It is structured reconstruction, not storytelling by memory alone.

That distinction matters. A weak version of retrospective evaluation becomes impressionistic. People remember what they liked, cherry-pick dramatic anecdotes, and inflate causality. A strong version does the opposite. It asks for traces, corroboration, sequence, and plausibility. It treats change like an archaeological site. You cannot watch the event happen, but you can still examine the layers left behind.

Imagine trying to understand how a new idea spreads through a neighborhood. A survey might tell you how many people say they heard about it. Outcome harvesting would ask a richer set of questions: Who changed behavior first? What conversations preceded the shift? Which institutions adapted? What evidence exists in emails, meeting notes, policy drafts, or public statements? The point is not to force change into a prewritten theory. The point is to let the evidence reveal what kind of change actually occurred.

This makes outcome harvesting especially valuable for complex interventions where the most meaningful effects are often unplanned outcomes. In such settings, causality is less like pulling a lever and more like nudging a landscape. The terrain responds in multiple directions. Some effects are immediate, others are delayed, and some appear in places the intervention never directly touched.

Transparency is not the enemy of flexibility, it is what makes flexibility trustworthy

At first glance, transparency and outcome harvesting seem to belong to different cultures. Transparency evokes protocols, documentation, data sharing, audit trails, and reproducibility checks. Outcome harvesting evokes adaptability, qualitative judgment, and emergent discovery. But these are not opposites. They are complementary answers to the same problem: how do we know that our interpretation of change is not just convenient fiction?

The key insight is that flexibility without transparency becomes improvisation, while transparency without flexibility becomes blindness.

A transparent evaluation process does not require that every outcome be known in advance. It requires that the process of finding outcomes be visible, inspectable, and defensible. If an evaluator discovers a surprising shift in institutional behavior, transparency means documenting how that outcome was identified, what evidence supported it, what alternative explanations were considered, and why the conclusion is credible. Transparency is the chain of reasoning. Reproducibility is the ability of others to examine or rerun that chain.

This matters because the most common critique of qualitative or adaptive evaluation is not that it is inherently subjective. It is that it can become opaque. When the path from observation to conclusion is hidden, trust erodes. But when the path is visible, adaptive evaluation can be as rigorous as any pre-registered design, sometimes more so, because it is honest about uncertainty instead of pretending uncertainty does not exist.

Here is the deeper synthesis: reproducibility should not mean reproducing only the result. It should mean reproducing the reasoning conditions under which the result was reached. That is a much higher standard, and a much more useful one.

A better mental model: evaluation as signal detection, not just hypothesis testing

A helpful way to connect these approaches is to think of evaluation as a signal detection problem. In a noisy environment, the question is not only whether a signal exists. It is also whether your instruments are designed to hear the right frequencies.

Predefined indicators are like a radio tuned to one station. They work well when the transmission is stable and known. Outcome harvesting is like scanning the airwaves for unexpected broadcasts. It does not abandon rigor. It broadens the search field. Transparency and reproducibility then act as calibration tools, ensuring that what you detected was not static, wishful thinking, or selective attention.

This model clarifies why some evaluations fail even when they are methodologically sophisticated. They confuse precision with attunement. A precise instrument can still miss the phenomenon if the phenomenon is changing shape. In complex programs, the most important outputs may not be outputs at all in the conventional sense. They may be shifts in norms, relationships, decision rules, or institutional memory.

To put it concretely, imagine evaluating a project aimed at improving collaboration among local actors in a climate adaptation network. A narrow evaluation might count meetings held or documents produced. A signal detection approach would also ask: Did actors begin sharing data earlier than before? Did a formerly skeptical agency start referencing the project in its planning memos? Did a new norm of joint problem solving emerge? These are outcomes, even if they do not fit a spreadsheet neatly.

The evaluator’s task, then, is not merely to measure a destination. It is to detect a route that was not fully visible at departure.

The practical synthesis: from rigid indicators to disciplined curiosity

The most powerful combination of these ideas is a workflow that begins with openness and ends with accountability. That means allowing room for unanticipated outcomes while still insisting on documented evidence and clear reasoning. It is a form of disciplined curiosity.

Disciplined curiosity asks three questions throughout an evaluation:

  1. What did we expect to happen?
  2. What actually changed that we did not expect?
  3. What evidence shows that the change is real, meaningful, and plausibly linked to the intervention?

This sequence preserves rigor without prematurely narrowing the field of inquiry. It also changes how teams work. Instead of defending indicators as if they were sacred, they can treat them as hypotheses about where change might appear. When something unexpected happens, that is not a nuisance. It is a lead.

A concrete example helps. Suppose an education initiative aims to improve classroom learning through teacher training. The planned indicators might include attendance, test scores, and lesson completion. But during implementation, evaluators notice that teachers begin forming peer networks, sharing lesson plans informally, and pushing for changes in school scheduling. These are not side notes. They may be the mechanism through which learning improves. Outcome harvesting captures these emergent changes, while transparency ensures that the case for their importance is traceable and credible.

This approach is especially valuable when working with systems that adapt in response to being observed. In such settings, the evaluation itself becomes part of the system. If the process is too rigid, it can distort behavior. If it is too loose, it loses trust. The answer is not to choose one pole. It is to design for structured openness: predefined commitments about evidence standards, paired with permission to discover outcomes not anticipated at the outset.

Key Takeaways

  • Do not confuse a predefined indicator with the full reality of change. Important outcomes often emerge after the project begins and may never appear in the original logic model.
  • Treat transparency as a method for making flexibility credible. Document how outcomes were identified, what evidence supports them, and what alternatives were considered.
  • Think in terms of signal detection, not only hypothesis testing. In complex settings, evaluation should scan for unexpected but meaningful changes, not just confirm known ones.
  • Use outcome harvesting to discover, then use reproducibility to defend. The first helps you find what matters, the second helps others trust it.
  • Aim for disciplined curiosity. Be open to surprise, but never casual about evidence.

The real lesson: evidence should help us learn what we could not have planned for

The deepest connection between transparent reproducibility and outcome harvesting is not technical. It is philosophical. Both are responses to the same humility: the world changes in ways our plans do not fully anticipate. One response is to make our methods so clear that others can inspect and trust them. The other is to make our inquiries so open that they can detect change we did not know to look for.

That combination is more than a methodological compromise. It is a more mature theory of knowledge. It says that good evaluation is neither blind certainty nor formless interpretation. It is a practice of seeing clearly in conditions of uncertainty.

Perhaps the most useful shift is this: stop asking whether an evaluation was only rigorous or only flexible. Start asking whether it was rigorous enough to support discovery, and flexible enough to notice what discovery required.

In complex worlds, the point of evidence is not to prove that reality obeyed our plan. The point is to reveal how reality actually moved.

Once you see that, evaluation changes from a courtroom into a mapmaking exercise. The question is no longer just whether the route matched the itinerary. It is whether we learned enough from the terrain to navigate the next journey better than the last one.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣
When You Cannot Predict the Outcome, You Must Prove the Trail | Glasp