The Hidden Discipline Behind Useful Evaluation: Why Good Programs Need Better Stories of Causation

Anemarie Gasser

Hatched by Anemarie Gasser

Jul 06, 2026

10 min read

87%

0

What if the main problem is not measuring impact, but understanding what kind of world your program is acting in?

Most evaluation debates sound like a fight over instruments. Do we need randomized trials, logic models, qualitative interviews, dashboards, or cost benefit analysis? That is the wrong starting point. The deeper question is more unsettling: what exactly counts as evidence when a program is trying to change something messy, human, and context dependent?

A vaccination campaign, a job training program, a school improvement initiative, and a housing intervention do not fail or succeed in the same way. They operate in different worlds, with different kinds of causation. In some settings, the central issue is whether an intervention works on average. In others, it is why it works here, for these people, under these conditions. In still others, it is whether the chain of assumptions behind the intervention is even plausible.

That is why useful evaluation is not just about finding out whether something happened. It is about building a theory of change that can survive contact with reality.

The real challenge in evaluation is not choosing between numbers and narratives. It is learning how to explain change without pretending the world is simpler than it is.


The temptation to treat programs like machines

Organizations love clean models because clean models reduce anxiety. If a program has a clear input, process, output, and outcome, then accountability seems manageable. You can count participants, track completion, and compare results. The problem is that many social programs are not machines. They are more like ecosystems. People interpret them, resist them, adapt them, and sometimes transform them in ways no design document anticipated.

A job placement program, for example, might look straightforward: train people, connect them to employers, improve employment outcomes. But the actual pathway may depend on trust, local labor markets, transportation, child care, employer bias, and whether participants believe the program is worth staying in. The same intervention can produce very different results in two neighborhoods because the mechanism is not simply “training causes jobs.” It is something closer to, “training may increase confidence and employability when a supportive network, accessible employers, and realistic job pathways exist.”

This is where many evaluations go wrong. They ask only whether an intervention worked, not what kind of work it was doing in a specific setting. That narrow question can be useful for some decisions, but it is often too blunt for learning, adaptation, and scale. A program that appears unsuccessful on average may still work well for a subpopulation. A program that appears successful may only succeed because of unusually favorable conditions that disappear when scaled.

The hidden risk is mistaking a statistical result for an explanation. A result tells you that change occurred. An explanation tells you why it happened, for whom, and under what conditions.


The missing link: explanation as the real currency of evaluation

If you zoom out, the common thread across theory based and realist approaches to evaluation is a deceptively simple idea: evaluation should not merely rank interventions, it should explain them.

That sounds obvious until you try to do it. Explanation is hard because change is not produced by a single cause. It comes from an interaction among three things: the intervention itself, the mechanism it activates, and the context in which that mechanism operates. A workshop does not create behavior change by magic. It may create insight, motivation, social pressure, legitimacy, or a sense of possibility. Those mechanisms only matter if the surrounding context allows them to fire.

This gives us a more useful mental model than the usual input to outcome pipeline. Think of a program as a key, a lock, and a room.

  • The key is the intervention design.
  • The lock is the mechanism it is trying to trigger in the participant or system.
  • The room is the context, which determines whether the door opens into a usable space or into a wall.

A key can be beautifully made and still fail if it does not fit the lock. A lock can be well designed and still not open anything if the room has changed. That is why evaluation must ask not only whether the door opened, but whether the key, lock, and room were aligned.

This is the practical power of a theory based approach. It forces organizations to articulate their assumptions before they collect results. It turns vague optimism into testable claims. Instead of saying, “This mentoring program will help youth succeed,” the theory becomes, “If youth build a trusted relationship with a mentor, then they are more likely to sustain effort, make better decisions, and access opportunities, especially when schools and employers provide real pathways.” Now evaluation can investigate each link in the chain.

Good evaluation is not a verdict. It is a disciplined explanation of how change is supposed to happen, and where that explanation breaks.


Why realist thinking changes the question from “Did it work?” to “What worked, for whom, in what circumstances?”

Realist evaluation sharpens theory based evaluation by insisting that outcomes come from mechanisms operating in contexts. That framing matters because it prevents a common error: treating a successful program as universally portable, or a failed program as universally flawed.

Imagine two cities adopt the same after school tutoring initiative. In City A, students can stay after school because transportation is reliable, parents trust the program, and tutors coordinate with teachers. In City B, students have to leave early for family responsibilities, tutors have little connection to school curriculum, and attendance is erratic. The identical intervention produces different outcomes because the context changes what the program means in practice.

This perspective is especially valuable because many interventions are not deterministic. They do not guarantee a result. They create a possibility structure. They make certain outcomes more likely by changing information, incentives, relationships, or confidence. The evaluator’s job is to identify which conditions are necessary for those possibilities to become real.

A useful way to think about this is through the sentence: “In context C, mechanism M produces outcome O.”

That structure seems simple, but it is revolutionary in practice. It shifts evaluation from universal generalization to conditional learning. Not every program needs the same question answered. Sometimes the key question is:

  • In what contexts is this intervention likely to work?
  • Which mechanisms are actually being triggered?
  • What unspoken assumptions are making the design fragile?
  • What adaptations preserve the core mechanism while fitting local reality?

This is why highly useful evaluation often looks less like pass or fail and more like map making. It draws the terrain. It shows where the road is paved, where it turns to mud, and where a bridge would make all the difference.


The most valuable evaluations do not just judge programs, they improve them

There is a subtle but important shift here. If evaluation is only about accountability, then the end goal is a score. But if evaluation is also about learning, then the end goal is decision quality.

That distinction matters because many organizations already know how to count. What they struggle with is adaptation. They can report enrollment, completion, and short term outcomes, but they cannot tell why one site succeeds and another stalls. They cannot say which component matters most. They cannot distinguish a broken theory from a weak implementation.

This is where theory becomes operational rather than academic. A well built program theory acts like a diagnostic tool. It helps answer four questions that every serious evaluator should ask:

  1. What is the intended mechanism of change?
  2. What contextual conditions must be present for the mechanism to work?
  3. What evidence would show that the mechanism actually fired?
  4. What would count as a sign that the theory is wrong, incomplete, or needs revision?

Notice what this does. It creates a bridge between design, implementation, and interpretation. If participants do not improve, the issue might be poor delivery, wrong assumptions, unsuitable context, or an intervention that never had a credible causal pathway in the first place. Without theory, all failures look alike. With theory, they become distinguishable.

Consider a literacy program in a low resource school. A shallow evaluation might ask whether reading scores increased. A theory rich evaluation might ask whether students received more feedback, whether teachers changed instructional routines, whether attendance improved, whether the classroom climate supported practice, and whether the program altered student confidence. The result is not just a number. It is a diagnosis.

That diagnosis is what allows improvement. It tells you whether to redesign the intervention, strengthen delivery, change the setting, or abandon a false assumption. In that sense, evaluation is not the final step after action. It is part of the action itself.


A practical framework: four questions that make evaluation more useful

The best way to avoid sterile evaluation is to replace vague questions with a sharper sequence. Here is a simple framework that can travel across sectors.

1. What change is the program trying to create?

Start by naming the outcome in plain language. Not “capacity building” or “service enhancement,” but something observable: higher attendance, faster placement, safer housing, better adherence, stronger civic participation.

2. Through what mechanism should that change happen?

Mechanisms are not activities. They are the forces activated by activities. Training is an activity. Confidence, skill, trust, motivation, and reduced friction are mechanisms. If you cannot name the mechanism, you are probably describing logistics, not theory.

3. What context must be true for the mechanism to operate?

This is where many plans become honest or collapse. Ask what has to be in place: staff buy in, transport, institutional support, legal authority, cultural fit, timing, or peer effects. A program is rarely wrong in the abstract. It is often wrong in a specific setting.

4. What would persuade us to revise the theory?

A good theory is not one that can explain everything. It is one that can be tested. If the intended mechanism is not visible, if the contextual conditions are absent, or if outcomes improve for reasons unrelated to the design, the evaluation should not be a rescue operation. It should be a learning device.

This framework does something important psychologically as well. It protects organizations from two dangerous habits: overclaiming success and overexplaining failure. It forces humility without surrendering ambition.


Key Takeaways

  • Stop asking only whether a program worked. Ask what kind of change it was designed to trigger, and under what conditions that change should be expected.
  • Distinguish activities from mechanisms. Training, outreach, and incentives are not the causal engine. Trust, motivation, capability, and reduced friction often are.
  • Treat context as a causal ingredient, not background noise. The same intervention can fail or succeed depending on local constraints, relationships, and institutions.
  • Use evaluation as a diagnostic tool. When results disappoint, theory helps identify whether the problem is the intervention, the implementation, or the setting.
  • Build falsifiable program theories. A useful theory does not explain everything after the fact. It makes clear predictions that can be tested and revised.

The deeper lesson: evaluation is a craft of disciplined curiosity

There is a moral dimension to all of this. When evaluation ignores context and mechanism, it can become a blunt instrument that rewards surface performance and punishes complex reality. It can flatten human systems into metrics and create the illusion that improvement is only a matter of better compliance with a preset model.

But when evaluation is theory driven and realist in spirit, it becomes something more intelligent and more respectful. It acknowledges that people are not passive recipients of interventions. They interpret them. They respond to them. They make them work or not work. That means good evaluation must be curious enough to ask not just what happened, but what the intervention meant inside the real world where it landed.

The most powerful shift is this: success is not the opposite of failure, learning is the opposite of blindness. A program can fail and still teach you something essential. It can succeed and still conceal a fragile theory. The evaluator’s role is not to hand down a final judgment, but to reveal the causal story with enough clarity that the next decision is smarter than the last.

So the next time someone asks whether a program worked, resist the easy answer. Ask instead: worked through what, for whom, and in what world? That question is harder, but it is the one that turns evaluation from a report card into a source of real knowledge.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣