What If the Real Goal Is Not Proving Impact, but Learning What Actually Happened?
Hatched by Anemarie Gasser
Apr 25, 2026
9 min read
1 views
74%
The measurement trap we rarely admit
What if the problem with many evaluation systems is not that they are too weak, but that they are asking the wrong question?
For years, organizations have chased a comforting fantasy: if we can just define the right outputs, we will finally know whether a program worked. Count the trainings, the participants, the sessions, the deliverables, the outputs. Then convert those numbers into a story of success. The trouble is that the world does not cooperate with this fantasy. Human change is messy, indirect, and full of surprises. The thing we most want to measure is often the thing least likely to show up neatly in a spreadsheet.
That is where a different tension emerges. On one side is the urge for proof: clean causal chains, visible indicators, and defensible reporting. On the other side is the need for learning: a disciplined way to understand what changed, for whom, and why. The most interesting insight is that these are not identical tasks. In fact, confusing them may be one of the biggest reasons organizations produce reports that look precise while teaching almost nothing.
The real challenge is not to make change look measurable. It is to make learning rigorous enough to survive the messiness of change.
Why outputs feel safe, but often tell the smallest truth
Outputs are attractive because they are easy to count and hard to argue with. If a program ran 12 workshops, served 400 people, and produced 8 toolkits, there is a reassuring solidity to that ledger. But outputs are only evidence that activity happened. They are not evidence that anything meaningful shifted in the world.
This is the first conceptual trap: activity is not impact. A bridge can have thousands of cars cross it, yet if it bends under load, the real question is not how many cars used it but whether it carried them safely. Likewise, a development program, community initiative, or social intervention can be busy without being consequential. The output report may satisfy an administrative need, but it can also flatten the very complexity that makes change worth studying.
That is why many people feel a quiet frustration with traditional reporting. It often asks, in effect, “What did you produce?” when the deeper question is, “What did that production set in motion?” The second question is harder because it forces us into causality, context, and interpretation. It also admits uncertainty, which is precisely what many institutions try to avoid.
Yet uncertainty is not a weakness in evaluation. It is the terrain.
Causality is not a straight line, it is a pathway through a crowd
If outputs are the ledger, causality is the map. But the map is not a single road from A to B. It is a network of pathways, detours, side effects, and interacting forces. Programs do not act on blank slates. They enter systems already shaped by incentives, history, relationships, and competing pressures.
This is why simple attribution often disappoints. The question “Did this program cause the outcome?” sounds clear, but in real life it is usually too blunt. A better question is, “What role did this intervention play in producing the outcome?” That shift from attribution to contribution matters enormously. It acknowledges that change is usually collaborative, distributed, and contingent.
Imagine a neighborhood where youth violence drops after a series of interventions: mentoring, school reforms, better lighting, community organizing, and police policy changes. No single actor can honestly claim sole credit. But neither should the answer collapse into “everything mattered equally” or “nothing can be known.” Instead, the serious task is to trace the causal pathway. Which changes appeared first? Which actors influenced others? Which conditions made progress possible? Which links in the chain were weak, strong, or missing?
This is where the deepest intellectual shift happens. The point is not to prove that one initiative owns the outcome. The point is to understand how change assembled itself.
The missing middle: from counting results to tracing mechanisms
Most evaluation debates get stuck between two unsatisfying poles. One pole says, “If you cannot measure it precisely, you cannot know it.” The other says, “If change is complex, any attempt to measure it is reductionist.” Both miss the more interesting middle ground: mechanism.
A mechanism is not just a result, and it is not just an activity. It is the process by which something turns into something else. For example, a mentorship program does not succeed merely because it exists. It succeeds if it creates trust, shifts expectations, expands access, or changes decision making. Those are mechanisms. Without them, the intervention is just motion.
This perspective unlocks a more useful evaluation question: not “How many outputs did we produce?” but “What causal ingredients were activated, and how did they combine?” In other words, the value of a program lies not only in what it delivered, but in what it made possible.
A useful analogy is cooking. If you only count ingredients, you still do not know whether the meal was any good. If you only describe the final taste, you do not know what went wrong or what to repeat. Real learning requires tracing the recipe, heat, timing, and interactions. Evaluation works the same way. Contribution analysis asks for the recipe of change, while approaches that capture significance often ask whether the meal mattered to the people who ate it. Together, they move us from crude scorekeeping toward intelligent judgment.
The goal is not to replace numbers with stories, but to build a better theory of how stories become evidence.
Why significance matters more than output volume
A neglected truth in many systems is that not all results are equal. One small change can matter more than a hundred routine outputs. A single policy shift may unlock years of downstream improvement. A quiet change in trust between partners may alter an entire institution’s ability to act. Some outcomes are numerically small but structurally transformative.
This is where the language of significance becomes powerful. Significance asks: what changed that was genuinely consequential? What mattered to the people involved? What shifted in behavior, relationships, confidence, power, or possibility? In that sense, significance is not just a softer metric. It is a more intelligent one.
Consider two community programs. Program A hosts 50 events and reaches 2,000 attendees, but participants leave largely unchanged. Program B holds 6 gatherings, yet those meetings reshape local leadership, create new alliances, and trigger policy reform. Traditional output logic favors Program A. A contribution-focused lens, however, would likely find Program B far more valuable. The difference is not sentimental. It is causal.
This reframes a common organizational mistake: confusing scale of activity with scale of change. The first tells you how much was done. The second tells you what was transformed. They are not interchangeable.
A hybrid model: output, pathway, significance
The best way forward is not to abandon structure, but to broaden it. Think of evaluation as having three layers:
- Output layer: What was delivered?
- Pathway layer: How did change likely happen?
- Significance layer: What difference did that change make?
Each layer answers a different question, and each is incomplete on its own. Output without pathway becomes bookkeeping. Pathway without significance becomes an elegant theory with no human stakes. Significance without pathway can become inspiring but untestable.
This three layer model is useful because it forces discipline without pretending that causality is simple. It says: yes, measure the activity. Yes, trace the mechanisms. Yes, ask whether anyone’s life, system, or decision actually changed in a meaningful way. The strength of the model lies in the fact that each layer checks the others.
For example, a workforce training initiative might report 300 graduates. That is the output layer. But the pathway layer asks whether participants gained confidence, access to networks, or actual hiring opportunities. The significance layer asks whether their earnings rose, whether employers changed their practices, or whether a barrier in the labor market was reduced. Only when all three layers are visible do you have something close to a serious account of change.
The deeper organizational shift: from proving to improving
The most important implication of this synthesis is strategic, not technical. Organizations often treat evaluation as a courtroom. The job is to prove that the work succeeded. But if the real purpose of evaluation is learning, then the better metaphor is the laboratory. The question is not “Can we defend ourselves?” but “Can we understand what happened well enough to do better next time?”
This shift changes behavior. When teams know they will be judged only on outputs, they optimize for visible activity. When they know they will be asked about contribution and significance, they begin designing interventions with stronger theories of change, clearer assumptions, and more attention to feedback loops. They stop asking only how much they did and start asking why certain actions mattered.
That is a healthier culture. It rewards intellectual honesty. It makes space for complexity. It also reduces the pressure to oversell certainty. In complex settings, the most credible claim is rarely “we caused this outcome alone.” It is more often “here is how we contributed, here is what else mattered, and here is why we think this pathway is plausible.” That kind of claim is less theatrical, but far more useful.
In complex systems, humility is not the opposite of rigor. It is part of rigor.
Key Takeaways
- Stop treating outputs as outcomes. Counting activity can be useful, but it cannot tell you whether meaningful change happened.
- Shift from attribution to contribution. In complex systems, ask what role your work played in a larger causal pathway rather than claiming sole credit.
- Look for mechanisms, not just results. Identify the specific processes, such as trust, access, incentives, or coordination, that turn activity into change.
- Measure significance, not just scale. Small changes can be more transformative than large volumes of output.
- Use a three layer evaluation lens. Track outputs, trace pathways, and assess significance together for a more complete picture.
A better question to end with
The biggest mistake in evaluation is not choosing the wrong metric. It is believing that the metric is the reality.
Once you see that, the entire frame changes. The point is no longer to assemble a perfect stack of numbers that proves success. The point is to create a credible account of how change unfolded, what helped it along, and why it mattered. That is a more demanding standard, but also a more human one.
In the end, the deepest value of evaluation may not be control. It may be orientation. It tells us where change is actually happening, which levers are real, and what kind of impact deserves our attention. And perhaps that is the most useful question of all: not how much did we produce, but what became possible because we were there?
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣