Why Good Theories of Change Fail Unless They Are Treated Like Living Hypotheses

Anemarie Gasser

Hatched by Anemarie Gasser

Jul 31, 2026

10 min read

88%

0

The dangerous comfort of a neat story

Most plans fail for a reason that looks like success: they make the world seem simpler than it is. A theory of change can be beautifully logical, with clean arrows from inputs to activities to outputs to outcomes, and still be wrong in the one way that matters. It can describe a pathway so tidy that everyone involved stops asking the question that should keep the whole thing alive: What would have to be true for this to actually work?

That question is more unsettling than it first appears. It forces us to admit that a theory of change is not a map of reality, but a bet about reality. And a theory based evaluation is not just a technique for measuring progress, but a discipline for interrogating that bet. When these two ideas are combined, they reveal something important: many organizations do not fail because they lack plans. They fail because they confuse a plausible story with a tested one.

The deepest tension here is between coherence and truth. Coherence feels good because it gives us order. Truth is harder because it demands contact with the messy, conditional, often inconvenient world. The real value of theory based thinking is not that it produces prettier diagrams. It is that it turns planning into a form of disciplined doubt.

A good theory of change is not a declaration that the future will behave. It is a hypothesis about the conditions under which the future might cooperate.


The hidden problem with plans that explain everything

Organizations love explanations that are complete enough to be persuasive. A program starts, activities are delivered, people participate, outputs appear, outcomes follow. The story feels satisfying because it gives every stage a place. But completeness can be a trap. The more a theory of change tries to account for everything, the more it can hide the most important variables: context, assumptions, and alternative pathways.

Think of a bridge design. A beautiful blueprint is not the bridge itself. It must still withstand weather, load, soil conditions, and imperfect construction. A theory of change works the same way. It is not enough that the logic looks sound on paper. You need to know whether the assumptions hold under pressure. Are people actually able to engage? Do incentives align? Does the institution have the capacity to deliver? Is the surrounding environment stable enough for the expected causal chain to unfold?

This is where theory based evaluation becomes more than a compliance exercise. It asks evaluative questions that ordinary outcome tracking often misses. Not just: Did results occur? But also: Which links in the causal chain were strongest, which were weakest, and under what conditions did they hold? That shift matters because the most interesting failures are rarely total failures. More often, some parts of a program work while others break, and those fractures are where learning begins.

A campaign to improve public health, for example, may successfully increase awareness while failing to change behavior. A training program may boost confidence but not job placement. A school reform may improve classroom practices but leave achievement unchanged because attendance, home support, or assessment design is the true bottleneck. A simple success or failure label misses this richness. A theory based approach reveals the anatomy of performance.

The point is not to make evaluation more academic. The point is to make it more useful. If you do not know which assumption failed, you will fix the wrong thing.


The real unit of learning is the assumption

The best way to think about a theory of change is not as a sequence of activities, but as a chain of assumptions. Every link says, implicitly, “if this happens, then that becomes likely.” The failure of many strategies is not mysterious once you see it this way. The chain was never broken at the end, it was broken at the assumptions in the middle.

This suggests a different mental model: programs are not machines, they are conditional experiments. A machine produces predictable output if you feed it the right input. A conditional experiment only works when certain social, institutional, and behavioral conditions are present. Education, health, governance, climate action, and community development all operate this way. They are not vending machines. They are ecosystems.

Once you accept that, evaluation changes character. Instead of asking whether the program worked, you ask:

  1. Which assumptions were explicit, and which were only implied?
  2. Which assumptions were validated by evidence?
  3. Which assumptions proved fragile in specific contexts?
  4. Which mechanisms were activated, and which were never triggered?
  5. What did the results mean in light of the theory, not just in light of the targets?

This is a more intelligent form of accountability because it distinguishes between a weak design and a difficult environment. Those are not the same thing. A theory based approach can tell you whether the logic failed, the implementation failed, or the world resisted the logic in ways the design did not anticipate.

Consider a job training initiative aimed at young adults. The theory might say that if participants receive training, then they will gain skills, then employers will hire them, then incomes will rise. But the real bottleneck may be transportation, childcare, discrimination, or lack of employer trust. The theory of change is useful precisely when it is incomplete enough to be tested against reality. It creates a structured place for surprise.

The most valuable part of a theory of change is not the pathway. It is the list of things that must not go wrong for the pathway to matter.


From linear thinking to causal humility

There is a subtle but profound shift buried inside theory based evaluation: from linear certainty to causal humility. Linear thinking assumes that if we do the right things, outcomes follow in a mostly direct way. Causal humility assumes that outcomes emerge through interactions, delays, feedback loops, and context. It does not reject causality. It deepens it.

This matters because institutions often prefer causes that are easy to manage. They like levers. They like inputs they can count and outputs they can report. But many of the most important changes are not lever shaped. Trust, legitimacy, motivation, coordination, and norms are notoriously hard to measure and even harder to control. Yet they often determine whether a program succeeds more than the visible activities do.

Imagine trying to improve a neighborhood not by adding one service, but by changing the social fabric. You might launch youth programs, improve lighting, support local businesses, and coordinate policing. The theory of change might look straightforward. But whether the neighborhood improves depends on how these interventions interact. Youth programs may work only if parents trust them. Lighting may matter only if public spaces are already being used. Business support may matter only if foot traffic increases. The causal path is not one arrow, but a web.

This is why theory based approaches are so powerful when they are used well: they resist the illusion that measurement alone can substitute for understanding. A dashboard may tell you that attendance is up. It cannot tell you whether attendance is up because the intervention worked, because the weather changed, because competing options disappeared, or because participants were selectively retained. The theory of change provides the interpretive frame that turns data into meaning.

Causal humility is not weakness. It is the recognition that causation in social systems is rarely singular. The more complex the problem, the more important it becomes to track not just results, but pathways, mechanisms, and context.


A better framework: treat strategy like a scientific instrument

The most useful synthesis of these ideas is to stop treating a theory of change as a static plan and start treating it as a scientific instrument for decision making. A good instrument does three things. It detects signals. It filters noise. It helps you decide what to do next.

That gives us a practical framework for any serious initiative:

1. Specify the mechanism

Do not begin with activities. Begin with the mechanism of change. What is the supposed causal process? For instance, does the intervention work by increasing knowledge, changing incentives, building trust, reducing friction, or altering social norms? If you cannot name the mechanism, your theory is probably too vague to test.

2. List the critical assumptions

Write down the conditions that must be true for the mechanism to operate. These should include behavioral assumptions, institutional assumptions, and environmental assumptions. If a program depends on digital access, say so. If it depends on participant motivation, say so. If it depends on implementation fidelity, say so.

3. Define the failure modes

Ask how the theory could fail at each stage. This is where the evaluation becomes most valuable. Failure modes are not merely risks to be managed. They are diagnostic clues. They tell you where to look first when results disappoint.

4. Build feedback into the theory

The theory of change should not end at outcomes. It should include learning loops. What will be monitored early? What signals would suggest adjustment? What evidence would cause you to revise the theory itself?

5. Distinguish between adaptation and abandonment

Some theories fail because the idea is wrong. Others fail because the implementation context changed. A mature evaluation practice helps you tell the difference. That distinction prevents two common errors: stubborn persistence in a dead idea, and premature abandonment of a promising one.

This framework turns evaluation from a backward looking ritual into a forward looking intelligence system. It does not eliminate uncertainty. It makes uncertainty workable.


What this changes in practice

The practical payoff of this way of thinking is bigger than many organizations expect. It changes what leaders ask, what evaluators report, and what teams learn.

Instead of asking, “Did we hit the target?” a leader begins asking, “Which assumptions are we most unsure about?” Instead of reporting only outcomes, evaluators report whether the causal chain strengthened, weakened, or shifted. Instead of treating unexpected results as nuisances, teams treat them as clues about the underlying system.

This has a profound cultural effect. It makes honest learning safer. If the language of evaluation is only success versus failure, people hide ambiguity. If the language is mechanism, assumptions, and context, people can say, “The outcome did not appear, but we learned why.” That is not soft accountability. It is stronger accountability, because it improves the odds that future action will be better calibrated.

Here is a simple example. Suppose a nonprofit runs a mentorship program for first generation college students. A superficial evaluation might note that graduation rates did not rise enough. A theory based evaluation would ask whether the intervention actually built belonging, reduced administrative confusion, improved academic confidence, or expanded networks. It might discover that mentoring improved persistence for some students but not others, especially where financial stress remained untreated. The right response is not simply “scale” or “stop.” It is to revise the theory and identify the missing causal piece.

This is the difference between managing performance and improving understanding. Both matter, but only one compounds over time.

Good evaluation does not just judge whether a program worked. It tells you what kind of world the program assumes.


Key Takeaways

  • Treat every theory of change as a hypothesis, not a promise. The goal is not to sound confident. The goal is to make your assumptions testable.
  • Start with mechanisms, not activities. Ask how change is supposed to happen before asking what you will do.
  • Name the fragile links. Identify the assumptions most likely to fail, because those are the places where evaluation is most informative.
  • Use evaluation to explain variation, not just outcomes. A good theory based evaluation helps you understand why something worked in one place and not another.
  • Build revision into the process. The most mature theories are designed to change as evidence accumulates.

The real lesson: plans are not meant to be obeyed, they are meant to be improved

The deepest insight that emerges from combining theory of change with theory based evaluation is simple but easily forgotten: the purpose of a plan is not to predict the future perfectly. It is to make learning possible before reality delivers its verdict in full.

That reframes planning from a performance of certainty into an exercise in disciplined realism. A good theory of change does not hide complexity, it organizes it. A good evaluation does not merely confirm success, it interrogates the causal story behind it. Together, they create a more honest relationship with change itself.

In the end, the question is not whether your theory sounds coherent. The question is whether it can survive contact with the world and improve because of it. That is the difference between a strategy that looks intelligent and one that actually becomes intelligent over time.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣