When Theories Meet Surprises: A Practical Framework for Reverse Engineering Change
Hatched by Anemarie Gasser
Apr 16, 2026
8 min read
3 views
87%
What if evaluation began with the surprise
Why do so many evaluations explain only what was intended, while the real story of change slips through the cracks? Most evaluation practice oscillates between two instincts: build a theory, then prove it; or collect outcomes, then try to explain them. Each instinct on its own misses half the picture. The richer move is to hold both at once: use theories as provisional maps and surprises as the raw material for refining those maps.
This article offers a clear way to do that. It presents a practical, repeatable cycle for moving from forward looking hypotheses to backward looking evidence. It also gives simple mental tools for turning messy evidence into credible claims about contribution and causation. If you are responsible for designing, evaluating, or learning from programs in complex contexts, read on. This approach will help you make better sense of how change actually happens, and it will make your evaluations more useful for decision makers.
The two instincts and their blind spots
On one side is the strong tradition of theory based evaluation. This begins by articulating a theory of how intervention inputs will produce outputs and then outcomes, usually laid out in a pathway or a theory of change. That theory makes assumptions explicit, creates measurable indicators, and supports planning and accountability.
Strengths: clarity, accountability, and the ability to test specific hypotheses. Weaknesses: in complex settings the pathway is fragile. Assumptions break, unanticipated actors intervene, and emergent outcomes appear that the plan did not foresee.
On the other side is outcome harvesting. Instead of starting with a theory, this practice starts with outcomes. Evaluators collect evidence of changes in behavior, relationships, policies, or practice and then work backwards to assemble plausible contribution claims. This is powerful when results are emergent, messy, or produced by collective action.
Strengths: it captures surprises, unintended consequences, and diverse causal contributions. Weaknesses: contribution claims can be underdeveloped without a theoretical lens; it can be hard to generalize lessons or to predict future performance.
The core tension is simple: theory seeks to predict and explain; harvesting seeks to discover and infer. Both are necessary. Without theory, harvested outcomes float as anecdotes. Without harvesting, theories can become tautologies that only confirm what was written down earlier.
Reverse engineering change: a simple thesis
The central claim here is that evaluation should be organized as an iterative process of forward theorizing and backward harvesting. Treat theories of change as testable hypotheses, and treat harvested outcomes as diagnostic data that revise those hypotheses. Practically, this means moving beyond binary choices of attribution versus contribution and toward a graded practice of causal learning.
Here are the core moves that enact this thesis:
-
Treat every program theory as provisional. Write it so it can be falsified. Include clear assumptions and rival hypotheses. Use the theory to generate evidence needs, not as a final verdict.
-
Harvest outcomes widely. Document not only intended outcomes but also unexpected changes in behavior, relationships, institutions, and narratives. Capture dates, actors, and verifiable evidence for each outcome.
-
Build contribution narratives. For each harvested outcome, reconstruct the causal chain. Identify plausible mechanisms, relevant contexts, and alternative explanations.
-
Score contribution confidence. Use a simple graded scale to express how strongly the evidence supports the contribution claim. Update the theory by keeping what is corroborated and revising what is not.
-
Loop. Use revised theory to design new interventions and to guide focused monitoring. Repeat the harvest at regular intervals to surface new outcomes and test the updated theory.
This is not a single technique but a disciplined practice of alternating inference and hypothesis testing. It turns evaluation into a learning engine rather than a post facto audit.
Practical frameworks that make the idea usable
Below are three practical mental tools that make the cycle operational in real programs.
- The Predictive Responsive Cycle, abbreviated PRC
- Predict: Articulate the theory of change as a compact hypothesis. State the expected outcomes, the key mechanisms, and the critical assumptions.
- Watch: Monitor indicators but also scan for unplanned outcomes and signals of change using qualitative probes and stakeholder interviews.
- Harvest: Collect and document outcome statements with evidence. Include date, actor, observable change, and corroborating data.
- Diagnose: For each outcome, construct a contribution narrative that engages the original theory and alternative explanations.
- Revise: Update the theory. Decide what to keep, what to drop, and what to test next.
Use PRC on a quarterly cadence for programs in fast changing contexts or on an annual cadence for slower efforts.
- The Confidence Ladder for contribution claims
When you explain how a program contributed to an outcome, express confidence using a clear ladder:
- Possible: There is a plausible link, but evidence is thin or alternative explanations are strong.
- Probable: Multiple pieces of evidence point to the program contributing, but gaps remain in the causal chain.
- Strong: Evidence triangulates across sources, process tracing supports a mechanism, and rival explanations are unlikely.
- Confirmed: The causal contribution is well supported and replicated across cases.
This ladder helps communication. Funders and managers can see that an outcome is real even if contribution is only possible. It allows programs to prioritize strengthening evidence for the most important but least proven claims.
- The Causal Mosaic approach
Think of outcomes as tiles in a mosaic. Each tile is an observable change. The evaluator's job is to arrange the tiles to reveal a picture of causation. Some tiles will cluster around a clear mechanism. Others will sit on the margins and point to contextual enablers or inhibitors.
To build the mosaic, assemble diverse evidence types: interviews, documents, timelines, media logs, administrative data. Use pattern matching to connect tiles to mechanisms in the theory. Where tiles do not fit, construct rival hypotheses and test those through focused data collection.
Concrete examples and how to do it step by step
Example 1: A youth employment program
A program offers vocational training and soft skills coaching. The original theory predicts that trained participants will find formal sector jobs. After a year, outcome harvesting reveals that a large share of participants found informal work through family networks, and some started micro businesses that used training in unexpected ways.
How to proceed:
- Harvest: Record outcomes with dates, employers, earnings, and participant statements.
- Diagnose: Ask whether training content mattered, or whether the program mainly served as a signal to networks. Collect interviews with employers and family members.
- Score: If evidence shows employers valued the credential, mark the contribution as probable or strong. If networks were decisive, mark contribution to job creation as possible but contribution to network activation as strong.
- Revise: Modify theory to include signaling and network activation as mechanisms. Adjust monitoring to capture network contacts, not only formal employment.
Example 2: Advocacy campaign that changed a policy unexpectedly
A coalition lobbies for a regulatory change. The theory posits that direct advocacy to policymakers will sway votes. Harvesting finds an unexpected outcome: a private company released data that shifted public opinion, which created pressure on legislators.
How to proceed:
- Harvest: Document the release, media coverage, and legislative timeline.
- Diagnose: Map the relative influence of the coalition versus the data release using timelines and key informant interviews.
- Score: If evidence shows the data release was the tipping point, attribute the policy change to the interaction of advocacy and public data. Mark direct advocacy contribution as possible, and the combined effect as probable.
- Revise: Update the theory to include third party triggers and public opinion as pathways to policy change.
Step by step template you can use immediately
- Write a compact theory hypothesis in one page. Include expected outcomes, three key mechanisms, and three critical assumptions.
- During implementation, maintain an outcomes log. Each entry should answer: who changed what behavior, when, and how do we know? Attach evidence.
- At review points, pick the top 5 outcomes and build contribution narratives. Use timelines and at least two corroborating sources per claim.
- Score each claim on the Confidence Ladder. Note what evidence would move the claim one rung up.
- Update the theory and monitor specifically for the evidence you need.
Evaluation is not about proving a prewritten story. It is about iteratively building a credible account of how change happens in a messy world.
Key takeaways
- Document both planned and unplanned outcomes: keep an outcomes log that records the actor, change, date, and evidence.
- Treat theories of change as testable hypotheses: include rival hypotheses and assumptions you intend to test.
- Use a graded confidence scale for contribution claims: Possible, Probable, Strong, Confirmed.
- Alternate forward planning with backward harvesting on a regular cadence to refine program design and learning.
- Prioritize collecting the evidence that can move the most important claims up the Confidence Ladder.
A final reframing
The dominant metaphors in evaluation are too often either a map or a mirror. A map promises a route to outcomes. A mirror reflects what already happened. Better practice treats evaluation as both cartography and forensic science. You draft a map to navigate the terrain, and you perform a forensic reconstruction when you find footprints that do not match the map.
This reframing has practical implications. It makes room for surprise rather than treating it as a failure. It elevates the role of triangulation, timelines, and alternative explanations. Most important, it shifts the purpose of evaluation from judgment alone to ongoing learning and adaptation.
If you adopt the cycle outlined here, you will stop asking only whether a program worked. You will begin to ask how it worked, for whom, and under what conditions. That question matters more for building durable change in complex systems.
Change is seldom linear. To make sense of it, we must learn to predict and to be corrected. The most useful evaluations are those that welcome contradiction as a source of insight, and that use surprises to sharpen, not to discard, our best ideas.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣