When the Metric Is the Message: Why the Best Evaluations Start with Meaning, Not Output
Hatched by Anemarie Gasser
Aug 05, 2026
10 min read
0 views
85%
The problem is not that organizations measure too much. It is that they measure the wrong kind of truth.
What if the most important thing happening in your work cannot be captured by the dashboard you are most proud of? That question sits at the center of a quiet revolution in evaluation. For years, many systems have treated success as something that can be counted, aggregated, and reported upward. Outputs, indicators, targets, percentages. But some of the most meaningful change does not arrive as a neat line item. It shows up as a shift in confidence, a change in relationships, a new way of thinking, or a decision someone makes months later because of an experience that seemed small at the time.
This is where a deeper tension emerges. On one side is causal logic, the desire to explain how an intervention leads to an outcome through a credible chain of influence. On the other is lived significance, the recognition that what matters most to people is often not what fits into a conventional reporting template. Put differently, one approach asks, “What happened because of this?” while the other asks, “What changed in a way that mattered enough to be remembered?” The most interesting work is not choosing between them. It is learning how to hold both without flattening either.
The real question is not whether change can be measured. It is whether measurement can remain faithful to meaning.
Why output reporting often misses the point
Traditional output reporting is comforting because it looks objective. We can count trainings delivered, funds disbursed, households reached, or trees planted. These figures are useful, but they are only the surface of a much deeper story. They tell us what was done, not what was transformed. A school can run ten literacy sessions and still fail to inspire children to read. A health program can reach thousands and still leave communities distrustful. A climate project can plant saplings and still neglect whether people will protect them five years later.
This is not a minor technical limitation. It is a category error. Outputs are not outcomes, and outcomes are not the same as significance. A program may achieve its intended activities while missing the actual change it was designed to create. Conversely, an initiative may appear small on paper yet produce a profound shift in how people relate to one another, advocate for themselves, or navigate future decisions. Conventional reporting often rewards visible activity over invisible transformation.
The consequence is predictable. Teams begin optimizing for what is countable rather than what is consequential. Staff learn to speak the language of forms, not the language of change. In time, the reporting system becomes a performance of certainty, even when the real world is messy, nonlinear, and full of surprise.
A child who gains the courage to speak in class may not alter a dashboard. A caregiver who feels less isolated may not become an immediately reportable statistic. Yet these changes can be the very mechanisms through which long term impact unfolds. If evaluation cannot see them, it is not simply incomplete. It is blind to the roots of change.
The hidden bridge between causal pathways and most significant change
At first glance, a causal pathway approach and a significance based storytelling approach may seem to live in different worlds. One is associated with explanation, plausibility, and contribution. The other is associated with narrative, diversity, and the judgment of what people themselves experience as important. But they are not opposites. They are complementary answers to a harder question: How do we know that change happened, and how do we know what kind of change mattered?
A causal pathway approach helps us avoid magical thinking. It asks us to map the sequence of events, assumptions, and contextual conditions that make change plausible. It insists that interventions are not isolated acts but parts of a chain. That discipline matters because good intentions are not evidence. If we want to claim contribution, we need to understand how outputs may connect to intermediate changes and eventually to broader outcomes.
The most significant change approach, by contrast, helps us avoid bureaucratic emptiness. It recognizes that the most important changes are not always the most easily preselected. Instead of only asking people to confirm predefined indicators, it invites them to describe what changed most for them and why. This opens a space for emergent value, meaning the kind of change that only becomes visible when people are allowed to define significance in their own terms.
The bridge between these two approaches is not methodological. It is philosophical. Both reject the fantasy that evaluation can be reduced to a single number. Both imply that impact is relational, context dependent, and partly interpreted. And both challenge the assumption that the only legitimate evidence is the evidence that looks tidy.
The profound insight is this: causal logic without lived significance becomes sterile, while lived significance without causal logic can become ungrounded. One explains too much and feels the human texture disappear. The other captures the human texture but can struggle to show how change travels through a system. The richest evaluation practice uses each to discipline the other.
A better model: from counting outputs to tracing significance
Imagine you are looking at a river delta from above. Output reporting counts the number of channels, pumps, and gates. Causal pathway analysis maps how water flows from one channel to the next. But the most significant change approach asks a different question: which fields actually received water, which communities felt relief, and which unexpected areas became fertile because of a shift in flow? Together, these views produce a truer picture than any one can alone.
This suggests a practical framework: trace significance through pathways. Instead of asking first, “What can we measure?” ask, “What kind of change would matter enough that people would notice it in their lives?” Then work backward to identify the causal pathway that could plausibly lead there. Finally, look forward again through stories and lived experience to see whether that change actually occurred, whether it was intended, and whether it generated ripple effects you did not anticipate.
This framework does three things at once:
- It protects rigor by requiring a plausible chain of influence.
- It protects relevance by allowing participants to define what counts as meaningful.
- It protects learning by revealing surprises that indicator dashboards routinely miss.
Consider a workforce training program. Output reporting may celebrate the number of participants enrolled and certificates issued. A causal pathway lens asks whether skills improved, whether employers trusted the credential, and whether participants had access to real opportunities. A significance lens asks what changed most in the lives of participants. For one person, it may be income. For another, confidence to negotiate. For a third, the ability to imagine a future beyond survival. Those are not interchangeable details. They are clues to what the intervention actually touched.
This is where the evaluation conversation becomes richer. Instead of treating participants as data points, the process treats them as interpreters of change. Instead of assuming the designed metric is the best one, it tests that assumption against reality. And instead of reducing complexity, it learns from it.
The more important the change, the less likely it is to fit neatly inside the first metric we invented.
What organizations gain when they stop worshipping outputs
Organizations often fear that moving away from output centered reporting means losing control. In practice, the opposite can happen. When teams insist on predefined outputs alone, they create an illusion of control that hides uncertainty. When they make room for significance and causal pathways together, they gain a more realistic form of control, one rooted in learning rather than illusion.
The first gain is better decision making. If only outputs are visible, leaders may continue funding activities that are busy but ineffective. If stories of significance are systematically collected, leaders can spot which changes are durable, which are superficial, and which are entirely different from what was expected.
The second gain is humility. A measurement culture that only validates preselected indicators tends to reward overconfidence. By contrast, a system that asks people what changed most creates intellectual friction. It reminds teams that real life often resists the plan. That friction is not a flaw. It is evidence that the system is learning.
The third gain is strategic clarity. A program can drown in metrics and still not know what it is for. When significant change is named by participants and linked back to plausible pathways, the organization begins to see its true theory of change, not the one written in a proposal, but the one enacted in practice.
The fourth gain is ethical seriousness. Reporting systems can become extractive when they ask people to supply data for someone else’s success story. Inviting people to define what mattered most shifts the center of gravity. It says that the point of evaluation is not merely accountability upward. It is accountability to reality and to the people living inside it.
The deepest shift: evaluation as sensemaking, not surveillance
The most important change in evaluation may be conceptual rather than technical. We need to stop treating evaluation as a surveillance apparatus designed to verify compliance. Instead, it should be understood as a sensemaking practice, a disciplined way of learning what changed, why it changed, and why it mattered.
That shift changes the role of metrics. Metrics are no longer the final word. They become one kind of evidence among several. Stories are no longer decorative anecdotes. They become structured observations about significance. Causal pathways are no longer abstract diagrams for reports. They become hypotheses about how transformation happens in context.
This perspective also changes the role of disagreement. If one group says a program succeeded because it reached many people, while another says it mattered because it transformed a few lives deeply, the disagreement is not a problem to suppress. It is a clue. It may reveal that the initiative is producing broad but shallow effects in one area and narrow but profound effects in another. A mature evaluation culture can hold that complexity without forcing premature simplification.
A useful mental model is to think of three layers of evidence:
- Activity evidence: What was done?
- Pathway evidence: How did those activities plausibly contribute to change?
- Significance evidence: What changed in ways people experienced as important?
When these layers align, confidence increases. When they diverge, learning begins. Perhaps a program is active but not transformative. Perhaps transformation is happening through an unintended pathway. Perhaps the most important result is invisible to the current metrics. In each case, the goal is not to defend the dashboard. The goal is to improve the picture.
Key Takeaways
- Do not confuse outputs with impact. Activities tell you what was delivered, not what changed.
- Ask participants what changed most, not just what you planned to change. This reveals significance that predefined indicators may miss.
- Use causal pathways to test plausibility. A meaningful story of change should still make sense in terms of sequence, context, and contribution.
- Treat stories and metrics as complements, not competitors. Stories reveal value, metrics reveal scale, and pathways connect the two.
- Redesign evaluation as learning. The goal is not to produce a cleaner report, but a truer understanding of how change happens.
Conclusion: the best evidence is the kind that changes what you notice
The real breakthrough is not in inventing a perfect new metric. It is in recognizing that the world does not change in the same units that institutions use to report it. People change through trust, momentum, confidence, dignity, and relationships. Systems change through feedback loops, altered incentives, and newly visible possibilities. These are not decorations around the real work. They are the real work.
Once you see this, reporting looks different. The question is no longer how to force the complexity of change into an output table. The question becomes how to design an evaluation practice that is rigorous enough to trace contribution and human enough to recognize significance. That is a harder standard, but a far better one.
In the end, the most valuable measurement system is not the one that produces the neatest numbers. It is the one that helps you see reality more clearly, act more wisely, and respect the forms of change that matter most but are easiest to miss.
And that changes everything.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣