Why Good Evaluation Starts Before Anything Happens

Anemarie Gasser

Hatched by Anemarie Gasser

Jul 29, 2026

9 min read

58%

0

The hidden mistake in most evaluation

Most people think evaluation begins at the end, after a program, policy, or project has already done its work. That instinct is understandable, but it is also the source of one of the most expensive errors in decision making: we try to judge outcomes without first understanding the causal story that produced them.

A project can look successful for reasons that have little to do with the intervention itself. A school program may improve test scores because of a shift in student demographics. A job training initiative may appear ineffective because its participants faced barriers the program was never designed to solve. A public health campaign may seem powerful when, in fact, it rode a preexisting trend. The numbers are real, but the interpretation is fragile.

That is why the deeper question is not simply, “Did it work?” The deeper question is: What had to be true for it to work, and how do we know?

This changes evaluation from a scorekeeping exercise into a discipline of explanation. Instead of treating outcomes as the final answer, we begin to see them as clues.


Outcomes are only the surface of the story

A result is not the same thing as a cause. That sounds obvious, yet many decisions are built as if the two were interchangeable. Outcome evaluation often asks whether change happened. Theory based evaluation asks why change happened, or failed to happen, by tracing the chain from inputs to activities to outputs to intermediate effects to final outcomes.

That chain matters because real world programs rarely behave like light switches. They are more like domino setups inside a room full of moving air. You can arrange the pieces carefully, but the environment still matters. Participants bring their own motivations, institutions shape implementation, and external conditions can accelerate, distort, or erase effects.

A useful mental model is this: outcomes are footprints, theory is the path. If you only measure footprints, you know something moved. You do not know whether it walked, ran, stumbled, or was carried. A strong theory of change gives you the path so you can tell whether the footprints make sense.

This is why theory based evaluation is not a luxury reserved for sophisticated organizations. It is a practical necessity whenever:

  • the intervention is complex,
  • the context is variable,
  • results may emerge slowly,
  • or multiple forces could plausibly explain the same outcome.

In other words, almost everywhere that human action meets human systems.


The real challenge is not measurement, it is causality

Many evaluation efforts are hampered by an unspoken assumption: if you collect enough data, the truth will reveal itself. But data does not interpret itself. You can have a mountain of numbers and still not know whether your intervention mattered, for whom, and under what conditions.

The deeper work is causal reasoning. That means building and testing a plausible explanation of change before you claim success or failure. It means asking questions like:

  • What mechanism is supposed to produce the outcome?
  • What conditions must be present for that mechanism to activate?
  • What evidence would confirm or weaken that story?
  • What alternative explanations could fit the same result?

Consider a simple analogy. Suppose a restaurant says its new menu increased revenue. Was it the menu, the seasonal tourist surge, a nearby event, or a viral social media post? Revenue alone cannot answer that. But if you know the theory, for example, that shorter menus reduce decision fatigue and speed up orders, you can look for intermediate signals: faster table turnover, fewer abandoned orders, higher satisfaction among first time diners.

That is the power of a theory based approach. It does not replace outcome measurement. It gives outcomes meaning.

Without a theory of change, an outcome is just a number attached to an ending.

There is a second advantage as well: theory helps distinguish between two very different kinds of failure. One is implementation failure, when the idea is sound but execution is weak. The other is theory failure, when execution was fine but the logic of change was flawed. Those are not the same problem, and they require different fixes. Confusing them is how organizations keep repeating the same mistakes with new branding.


A better model: evaluation as a chain of trust

The most useful way to think about evaluation is not as a verdict, but as a chain of trust.

At the far end of the chain is the final outcome, the thing everyone cares about. But trust in that outcome depends on several links being intact:

  1. Clarity of purpose: What change are we actually trying to create?
  2. Theory of change: By what mechanism should the intervention create that change?
  3. Operational fidelity: Was the intervention actually delivered as intended?
  4. Intermediate evidence: Do we see the expected stepping stones along the way?
  5. Outcome evidence: Did the final result appear, and is it credibly linked to the intervention?
  6. Contextual interpretation: What else was happening that might explain the result?

If any one of these links is weak, the final judgment becomes less trustworthy. A program can win on paper because it was evaluated only at the endpoint, while the actual causal chain was broken in the middle. Conversely, a program can look weak in a blunt outcome metric but reveal strong intermediate movement that suggests long term value.

This is why sophisticated evaluation is not about proving everything succeeded. It is about pinpointing where the causal chain held and where it broke.

Imagine trying to diagnose why a city’s bike lane policy did not reduce car traffic. If you only observe the final traffic count, the policy may look like a failure. But a theory based analysis might reveal that bike lanes increased ridership, yet commuters still lacked safe intersections, workplace showers, or reliable transit connections. The policy worked partly, but the system around it did not.

That diagnosis is much more useful than a yes or no score. It tells leaders what to fix.


What theory based evaluation changes in practice

The best thing about theory based evaluation is that it changes how teams behave long before the final report is written. It forces clarity at the start, not post hoc rationalization at the end.

First, it makes assumptions visible. Every intervention rests on beliefs, often hidden, about human behavior and system dynamics. For example, a mentoring program may assume that guidance alone changes student outcomes. But perhaps students also need transportation, time, emotional safety, and family support. Naming those assumptions early prevents later disappointment from being misread as randomness.

Second, it turns evaluation into a learning tool. If a program is designed around a chain of change, then each stage can be monitored. Leaders do not need to wait months or years to know whether they are on track. They can use leading indicators to see whether the mechanism is functioning. That allows faster correction.

Third, it improves comparability across contexts. The same intervention may work in one setting and fail in another because the causal conditions differ. A theory based approach does not treat that as an embarrassment. It treats it as data. The question becomes not “Does it work?” but “Under what conditions does it work, and what needs to be adapted?”

That is a more honest question, and a more useful one.

A concrete example helps. Suppose a nonprofit launches a reading program for early learners. An outcome only evaluation asks whether reading scores improved. A theory based evaluation asks a richer sequence:

  • Were children attending consistently?
  • Did they receive enough practice with feedback?
  • Did teachers implement the program as intended?
  • Did children show better phonemic awareness before the final scores changed?
  • Were there confounding changes, such as new curriculum adoption or parental involvement campaigns?

If the final scores rise, the program can explain why. If they do not, the organization can see whether the issue was attendance, implementation, mechanism, or environment. That is the difference between learning from data and merely collecting it.


The deepest insight: good evaluation is a design discipline

Here is the most important reframing: evaluation is not something you do after design. Evaluation is part of design itself.

When you build a theory of change, you are already making choices about what matters, what can be measured, what must happen first, and which assumptions are worth testing. In that sense, a strong evaluation is an act of intellectual architecture. It does not merely inspect a program. It shapes the program’s logic.

This explains why some organizations become trapped in endless measurement. They measure outcomes because they never clarified the causal model. Then, when the result is ambiguous, they add more indicators. The dashboard grows, but the understanding does not. More metrics cannot compensate for a weak theory.

A better approach is to start with a disciplined question: What would we need to observe if our explanation were true? That question naturally produces better indicators. It also reduces vanity metrics, because every measure must earn its place by illuminating a step in the causal chain.

Think of evaluation like testing a bridge. You do not only look at whether the bridge is standing after traffic passes over it. You also inspect the supports, the stress points, the load distribution, and the materials. A bridge that looks fine from a distance may still hide structural weakness. In the same way, a program that produces a desirable outcome may still be fragile, unrepeatable, or dependent on conditions that will not hold next time.

This is why the most valuable evaluation questions are not always the simplest ones. Sometimes the better question is not whether a policy reduced homelessness, but whether it reduced specific pathways into homelessness, such as eviction, income shock, or family disruption. That reveals leverage. It tells policymakers where interventions can actually bite.


Key Takeaways

  • Start with a causal story, not a score. Before judging impact, define the mechanism by which change should happen.
  • Measure the chain, not just the endpoint. Track intermediate steps that show whether the intervention is working as intended.
  • Separate implementation failure from theory failure. If results are weak, determine whether the problem was delivery or design.
  • Use evaluation to learn conditions, not just averages. Ask when, where, and for whom the intervention works.
  • Treat assumptions as evidence targets. Every hidden belief in your design should become something you can test.

The final reframe

The biggest misconception about evaluation is that it exists to pass judgment. In reality, its highest purpose is to make change intelligible.

When you understand the causal chain, outcomes stop looking like mysterious verdicts handed down by the world. They become legible consequences of design, context, behavior, and timing. That shift is profound, because it turns disappointment into diagnosis and success into something you can reproduce.

So the next time you ask whether something worked, pause before looking at the final number. Ask instead: What had to be true for this outcome to happen, and did we actually build those conditions?

That question does more than evaluate performance. It teaches you how change really works.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣