When a Theory of Change Becomes a Theory of Comparison

Anemarie Gasser

Hatched by Anemarie Gasser

Jul 19, 2026

10 min read

92%

0

The hidden problem is not knowing what works. It is knowing what works compared with what.

Most people assume the hardest part of change is implementation. In practice, the harder problem is much older and more slippery: how do you know whether a change actually changed anything? A program can look successful in isolation and still fail in the world it was meant to improve. A policy can produce uplifting stories and still leave the underlying system untouched. A model can be elegant on paper and useless in practice if it never meets a plausible alternative.

That is where the real tension begins. A theory of change promises direction, a chain from action to outcome. Comparative research promises judgment, a way to tell whether one pathway is better than another. Each by itself is incomplete. Together, they expose a deeper truth: impact is never just about having a plan, it is about having a credible contrast.

Change cannot be understood in the abstract. It only becomes legible when placed beside a counterfactual, a baseline, or another route the world could have taken.

This is why so many well intentioned efforts are misread. We celebrate activity because it is visible, but visibility is not validity. The deeper question is not, “Did we do something good?” It is, “Did this sequence of choices produce a difference that matters, and why this sequence rather than another?”


The illusion of linear progress

A theory of change often feels satisfying because it arranges chaos into a story. Inputs lead to activities, activities to outputs, outputs to outcomes, outcomes to impact. The logic is clean, almost architectural. But social reality is rarely so obedient. People do not move through change like water through pipes. They respond, resist, reinterpret, delay, adapt, and sometimes transform the intervention itself.

This is where comparison becomes essential. Comparison is not a bureaucratic add on. It is the mechanism that prevents a theory of change from becoming a story we tell ourselves after the fact. Without comparison, every outcome can be narrated as expected, even when the path was accidental. With comparison, the map gets tested against terrain.

Think of two city neighborhoods receiving similar youth employment programs. In one, youth employment rises by 12 percent. In the other, it rises by 4 percent. The first reaction is to applaud the stronger program. But comparison asks harder questions: Were the neighborhoods comparable? Did one have better transit access, safer streets, or stronger employer networks? Did the program work because of the intervention itself, or because it landed in a place already primed for success?

That is the deeper function of comparative thinking. It resists the temptation to confuse sequence with causality. A theory of change says, “Here is how we expect the world to move.” Comparative research says, “Here is how to tell whether the movement was truly caused by this pathway.”

The two together create discipline. One gives ambition, the other gives evidence.


Comparison is not a method. It is a moral discipline.

We often talk about comparison as if it were purely technical, a matter of matching cases or checking variables. But comparison has an ethical dimension. It forces humility. It reminds us that every intervention operates among alternatives, and every claim of success implicitly excludes some other explanation.

This matters because change work is full of overclaiming. A community health initiative may reduce hospital visits, but maybe the real driver was a local hiring boom. A girls’ education program may improve attendance, but perhaps the strongest effect came from a broader shift in household expectations. Without comparison, we mistakenly credit what was visible and ignore what was structural.

Comparison also protects against the opposite mistake, underclaiming. Some interventions do work, but their effects are diluted, delayed, or masked by surrounding conditions. A program may not transform the whole system, yet it may still create the first viable foothold for later change. Comparative thinking helps distinguish no effect, weak effect, context dependent effect, and foundational effect.

A useful way to see this is through three questions:

  1. Compared with what? Every claim needs a reference point. Is the baseline doing nothing, doing something else, or doing the same thing differently?

  2. Compared where? Context is not background noise. It is often the main causal ingredient. An intervention can succeed in one place because local conditions supply the missing supports.

  3. Compared over what horizon? Some changes are immediate, others only become visible after institutions, norms, or incentives have time to settle.

This turns comparison into more than a research tool. It becomes a safeguard against narrative drift, the tendency to let a plausible story substitute for a tested one.

The best theories of change do not merely describe movement. They specify the conditions under which movement would count as evidence.


A better framework: the ladder, the lens, and the baseline

The deepest synthesis between planning and comparison is this: a strong change strategy needs three things at once, not one.

1. The ladder: a pathway of plausible steps

This is the theory of change proper. It identifies the sequence from action to result. But the ladder should not be too tidy. Each rung must be a testable hypothesis, not a comforting assumption. For example, “train teachers” is not a complete theory. “Train teachers, which changes classroom practice, which increases student engagement, which improves attendance, which then improves learning” is a hypothesis chain.

2. The lens: a way to see context

Comparison supplies the lens. It asks which features of the environment amplify, suppress, or distort the pathway. A literacy intervention may depend on home language, class size, parental time, digital access, or trust in schools. The lens reveals that an intervention is not a free standing object. It is an object embedded in a field of conditions.

3. The baseline: a credible alternative world

No claim of change is meaningful without a baseline. The baseline may be a control group, a prior period, a neighboring district, a similar community, or a policy that would have happened anyway. The precise form matters less than the principle: impact is difference making, not mere occurrence.

Taken together, these three elements produce a more mature logic of action. The ladder says what should happen. The lens says where it is likely to happen. The baseline says whether it happened because of you.

This framework also explains why many change efforts disappoint. They have ladders without lenses, which means they ignore context. Or they have lenses without ladders, which means they notice complexity but cannot act. Or they have baselines without theories, which means they measure difference without understanding the mechanism. Real rigor requires all three.

Consider an anti poverty cash transfer program. The ladder might suggest that cash increases household stability, which improves school attendance, which eventually raises educational attainment. The lens might show that the effect is stronger where schools are accessible and weaker where fees, transport, or safety remain binding constraints. The baseline might reveal that attendance would have improved slightly anyway because of seasonal work cycles. Only by combining the three can leaders distinguish genuine leverage from wishful thinking.


Why systems change often fails at the comparison problem

Systems change language is seductive because it promises transformation beyond isolated projects. Yet systems are exactly where comparison becomes most difficult. When many causes interact, attribution gets messy. That messiness leads two bad habits.

The first is overgeneralization. A successful pilot is treated as proof of universal validity. The second is fatalism. Because causality is complex, people conclude that nothing can be known with confidence. Both habits are avoidable.

The answer is not to demand perfect certainty. It is to use structured comparison. Structured comparison does not pretend the world is a laboratory. It simply asks that we compare like with like as much as possible, and when we cannot, that we explicitly state the limits of inference.

This matters especially in public policy, where interventions are judged not only by intent but by distribution. Who benefited? Who was left out? What changed in one district but not the other? Comparative analysis reveals that an intervention is rarely neutral. It has winners, losers, and boundary conditions.

Imagine two education reforms. Both increase average test scores. But one narrows inequality between high income and low income students, while the other widens it. A purely aggregate theory of change would call both successful. Comparative thinking would ask the better question: successful for whom, under what conditions, and at what cost?

This is one of the most important intellectual moves in applied work: from average effects to comparative effects. Once you start comparing across groups, regions, and time horizons, you stop treating change as a single number and begin seeing it as a pattern of distributional tradeoffs.

That shift is not academic nitpicking. It is how you avoid building interventions that optimize the visible middle while abandoning the edges.


The practical art of designing for comparison

If comparison is essential, how should it shape action from the start? The answer is to design interventions as if they will need to defend themselves against alternative explanations, because they will.

Here are four design principles that make a theory of change more testable and more useful.

1. Make assumptions explicit

Every pathway rests on assumptions. Parents must trust schools. Teachers must have time. Citizens must believe institutions will respond. State them plainly. Hidden assumptions are the first place theories of change break.

2. Build in variation

If everything happens everywhere at once, nothing can be compared. Staged rollout, pilot zones, phased implementation, and differentiated intensities are not delays. They are opportunities to learn.

3. Define failure and success in more than one way

A program can fail to hit a target and still illuminate a mechanism. It can hit a target but fail to scale. Define outcomes, but also define mechanism markers and equity markers.

4. Treat context as data

Do not record only outputs. Record the enabling and constraining conditions around them. The difference between a usable theory and a decorative one is often the quality of contextual documentation.

A useful analogy is medicine. A doctor does not simply prescribe a treatment and observe whether the patient improves. The doctor compares symptoms, history, test results, and alternative diagnoses. The treatment only makes sense against a model of what would likely happen without it. Social change deserves the same diagnostic seriousness.

Good change work does not ask, “Did we act?” It asks, “What would be different if this action were absent, delayed, or replaced?”

That question may sound technical, but it is actually liberating. It shifts the burden from storytelling to learning. It allows organizations to improve rather than merely justify themselves.


Key Takeaways

  • Always define your counterfactual. Before claiming success, specify what the alternative would have been. No baseline means no credible impact claim.
  • Separate pathway from proof. A theory of change is a hypothesis about how change should happen, not evidence that it did.
  • Compare contexts, not just outcomes. The same intervention can work differently depending on institutions, incentives, and local conditions.
  • Look for distributional effects. Average gains can hide unequal benefits. Ask who improved, who did not, and why.
  • Design interventions to be testable. Stagger rollout, record assumptions, and track context so learning is built in from the start.

The real synthesis: change is a comparison between futures

At the deepest level, theories of change and comparative research are both about possible worlds. One imagines the future you want. The other asks what the future would have looked like otherwise. That is why they belong together. Without a theory of change, comparison is directionless. Without comparison, theory of change is untested aspiration.

The most powerful organizations, researchers, and policymakers do not merely ask how to do good. They ask how to know whether their good is real, relative, and replicable. That question sounds modest, but it is revolutionary. It replaces certainty with discipline, and discipline with learning.

In the end, the purpose of comparison is not to make change smaller. It is to make change more honest. And honest change is the only kind that can scale without self deception.

When you next design a program, evaluate a policy, or tell a story of progress, ask a harder question than “What happened?” Ask, “What did we make more likely than what would have happened otherwise?” That is where explanation begins, and where meaningful change finally becomes visible.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣