What If the Most Honest Results Cannot Be Fully Reproduced?
Hatched by Anemarie Gasser
May 27, 2026
9 min read
1 views
71%
The hidden conflict in modern evaluation
What do we really want from a good measurement system: truth, or trust?
That question sits beneath a surprising tension in how organizations learn. On one side is the push for reproducibility, transparency, and cleaner evidence. On the other side is a competing impulse, especially in complex social programs, to capture what changed in ways that cannot be reduced to a spreadsheet, a baseline, or a tidy output metric. One side asks, “Can we verify it?” The other asks, “Did we miss the thing that mattered most?”
This is not just a technical disagreement. It is a philosophical one. It asks whether impact is best understood as a measurable object that can be replicated on demand, or as a living relationship between context, people, and change. The deeper problem is that the most legible outcomes are often not the most meaningful ones, while the most meaningful outcomes are often the hardest to standardize.
The real tension is not between rigor and flexibility. It is between what can be counted reliably and what can be understood responsibly.
That distinction matters because organizations often pretend those two goals are the same. They are not.
Why standard metrics often miss the point
Traditional reporting systems reward what can be extracted from reality quickly. They favor outputs that are easy to name, tally, and compare. How many workshops were held? How many people attended? How many reports were produced? Those numbers are useful, but they are also deceptively comforting. They can create the illusion of understanding while leaving the deepest question untouched: what actually changed in people’s lives, decisions, or behavior?
This is where many evaluation systems break down. They assume that the most significant result will look obvious in advance, and that it will be visible through a conventional indicator. But in social programs, education, community development, and policy work, the most important change is often indirect, delayed, or unexpected. A small shift in confidence. A new relationship across a divide. A change in how a team listens to local voices. These are not “outputs” in the usual sense, yet they may be the very mechanism by which lasting change occurs.
Consider a community health initiative. A standard dashboard may show the number of visits, trainings, or pamphlets distributed. Useful, yes. But it might completely miss the moment when a skeptical local leader becomes a trusted champion, or when residents begin calling the clinic earlier because they now believe it is for people like them. Those are not minor details. They are the difference between a program that looks busy and a program that actually works.
The problem is not measurement itself. The problem is measurement that mistakes administrative clarity for substantive insight.
Reproducibility is necessary, but not sufficient
The call for transparency and reproducibility emerged for a good reason. Too many claims have been made on the basis of hidden methods, selective reporting, or results that cannot be checked. If an evaluation cannot explain how it arrived at its conclusions, those conclusions should be treated with caution. Transparency protects against cherry-picking, wishful thinking, and the quiet laundering of weak evidence into policy confidence.
But reproducibility has a limit. It is powerful when the object of study is stable enough to be revisited under similar conditions. It becomes less straightforward when the thing being evaluated is adaptive, relational, or shaped by local context. In those settings, the question is not simply whether someone else can reproduce the same result. It is whether the evaluation can be audited, understood, and trusted without pretending that context is irrelevant.
Think of the difference between testing a machine and understanding a garden. A machine can be run again under controlled conditions. A garden cannot. You can document the soil, the weather, the seeds, and the watering schedule, but the garden’s life still depends on interactions that cannot be fully standardized. Social change is often closer to gardening than engineering.
This does not mean abandoning rigor. It means broadening what rigor is for. Rigor is not only about producing repeatable numbers. It is also about making the reasoning visible, the assumptions explicit, and the uncertainties honest. A transparent evaluation does not merely say, “Here is the result.” It says, “Here is how we interpreted what happened, what we may have missed, and why this account should or should not be believed.”
The overlooked question: can significance be standardized?
Here is the deeper puzzle connecting these two approaches: can an organization make room for significance without forcing it into a fixed output measure?
This question matters because significance is not always obvious at the start of a project. In many cases, the most important change only becomes visible after people tell stories about what shifted, why it mattered, and how it altered future behavior. A meeting that seems ordinary can, in retrospect, turn out to be the moment a partnership was repaired. A training session that produced no immediate test score increase may still have changed how a local team collaborates and solves problems.
This is where narrative methods become valuable. They do something metrics often cannot: they surface unexpected evidence of change. Instead of asking participants to fit experience into predefined boxes, they ask what happened that felt most important and why. That creates a different kind of evidence, one grounded in judgment, comparison, and collective sensemaking rather than only in numerical aggregation.
Yet narrative methods also face a legitimate challenge: without disciplined process, they can become impressionistic or biased. One person’s powerful story may not represent broader patterns. A vivid anecdote can seduce decision-makers into overclaiming. So the answer is not to replace numbers with stories, but to build a system in which stories and numbers correct each other.
The best evaluation systems do not ask whether stories are more true than numbers. They ask what kind of truth each one is capable of carrying.
That is the key synthesis. Numbers are excellent at scale and comparability. Stories are excellent at significance and context. Each becomes dangerous when it tries to do the other’s job.
A better model: evidence as a two-step conversation
A useful way to think about this tension is to treat evaluation as a two-step conversation.
Step 1: What changed that mattered?
This is the narrative step. It identifies significance before precision. It invites people closest to the work to describe not just activity, but transformation. What surprised them? What felt different? What was the most important change, and why?
This step is especially important when work is exploratory, adaptive, or community-based. It helps surface outcomes that were not anticipated in the original plan. It also protects against the tyranny of pre-defined indicators, which can quietly narrow the meaning of success to whatever was easiest to measure at the outset.
Step 2: What can we verify, compare, and inspect?
This is the transparency step. It asks how the story was collected, who participated, what alternatives were considered, and what evidence supports the interpretation. It does not demand that every insight become a randomized estimate. It demands that the path from observation to conclusion be visible enough for others to evaluate.
Together, these two steps create a healthier standard. They say that a claim about change should be both meaningful and inspectable. If it is meaningful but opaque, it invites skepticism. If it is inspectable but empty, it invites irrelevance.
You can think of this as moving from measurement as scoreboard to measurement as conversation with receipts. The conversation reveals what mattered. The receipts make the conversation credible.
What this looks like in practice
Imagine two organizations evaluating a youth mentorship program.
The first organization tracks attendance, completion rates, and survey scores. These data are easy to reproduce and compare. But when the numbers look flat, the program appears mediocre, even if mentors are helping participants stay in school, rebuild family trust, or imagine a future they had previously excluded.
The second organization begins by asking mentors, students, and parents to identify the most significant change they noticed during the year. One student describes learning how to speak in a meeting without shutting down. A parent describes getting a call from school that felt collaborative instead of punitive. A mentor describes a shift from advice-giving to listening. None of these changes fit neatly into a single output metric, but together they reveal a deeper pattern: the program is changing how people relate, and those relationships are the mechanism of impact.
Now add transparency. The organization documents how stories were gathered, how selection worked, who reviewed the narratives, and what evidence corroborated them. It notes that some participants saw little change, and that not every story was positive. It asks which themes recur across different accounts and which are unique to particular contexts. Now the report is no longer just compelling. It is auditable.
That combination is powerful because it prevents two common failures at once. It avoids the sterility of purely numerical reporting and the looseness of unexamined storytelling.
Key Takeaways
- Do not confuse outputs with outcomes. A high number of activities can coexist with weak or invisible change.
- Treat significance as a first-class object. Ask what changed that mattered before deciding how to measure it.
- Use stories to discover, not to decorate. Narrative should reveal unexpected outcomes, not merely humanize a prewritten conclusion.
- Use transparency to discipline interpretation. Make methods, selection rules, assumptions, and limitations visible.
- Build mixed evidence loops. Let qualitative insight identify what matters, then use quantitative or reproducible methods to test how broadly and reliably it appears.
The deeper lesson: evidence should earn trust, not just precision
The strongest systems of evaluation are not those that can reproduce the same answer forever. They are those that can explain why an answer deserves to be believed, what it cannot claim, and where its limits begin.
That is a more demanding standard than conventional reporting. It asks organizations to stop chasing the comfort of simple metrics and instead build a culture of disciplined interpretation. It also asks them to stop treating stories as soft evidence and start treating them as a source of discovery that must still be handled with care.
In that sense, the most important question is not whether a result is measurable or reproducible. It is whether the system can capture change without flattening it. If an evaluation can only recognize what was already easy to count, it will miss the most consequential shifts. If it can only tell compelling stories without exposing its logic, it will lose credibility.
The future of evaluation belongs to approaches that can hold both truths at once: significance is often subjective before it becomes legible, and rigor is most valuable when it makes that subjectivity inspectable rather than invisible.
So perhaps the real goal is not to prove that every important change can be reproduced in the same form. Perhaps the goal is to build methods honest enough to say: this mattered, here is why we think so, here is how we know, and here is what still remains uncertain.
That is not a compromise. It is a higher form of truth.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣