When the Evidence Is Messy, the Method Must Be Humble
Hatched by Anemarie Gasser
May 13, 2026
9 min read
1 views
76%
The real problem is not ignorance, it is false certainty
What if the biggest threat to good decision making is not that we lack evidence, but that we trust evidence too quickly?
That sounds counterintuitive, because modern evaluation culture often treats certainty as the prize. A program works, an intervention fails, a policy scales, a metric moves, and the story seems complete. But in practice, the most important questions rarely arrive in neat form. Results are partial, contexts differ, outcomes appear late, and the act of measuring itself can change what is being measured. In that world, the pursuit of one clean answer can become a liability.
This is where a deeper tension emerges: we need methods that can discover outcomes we did not anticipate, while also methods that can test whether those outcomes survive scrutiny. One approach is generative, almost investigative. The other is skeptical, almost forensic. The first asks, “What happened that matters?” The second asks, “Would this still hold if we looked harder?”
Together, they point to a more mature idea of evidence: not as a finished verdict, but as a disciplined conversation between discovery and doubt.
Evidence has two jobs, and they are not the same
Most organizations behave as if evaluation has a single job: prove whether an intervention worked. Yet in practice, evaluation has at least two distinct jobs.
The first job is finding meaning. Real-world interventions often create effects nobody predicted. A girls’ education program may improve not only attendance, but also household bargaining power, local norms around marriage, or the confidence of younger siblings. A climate adaptation effort may produce not just fewer losses, but new forms of coordination between farmers and local officials. If you only ask about preselected indicators, you may miss the most important changes entirely.
The second job is testing robustness. Once a result appears, the crucial question is not just whether it is statistically significant, but how fragile it is. Does the conclusion depend on a particular model choice? A specific subgroup? One unusual village? A narrow measurement window? Without sensitivity checks and replications, an impressive finding can be little more than an elegant accident.
The mistake is to treat these jobs as interchangeable. They are not. Discovery without verification produces stories that can seduce. Verification without discovery produces rigorously defended blindness.
Good evidence is not simply evidence with low uncertainty. It is evidence that can survive both surprise and suspicion.
That is the deeper standard these two intellectual traditions, when placed in conversation, reveal.
The hidden symmetry between field discovery and statistical doubt
At first glance, a process that surfaces unexpected outcomes and a checklist that stress tests findings might seem like opposites. One is open ended, the other structured. One looks outward to the field, the other inward to the model. But beneath that surface contrast lies a shared ambition: to make evaluation less arrogant.
Arrogant evaluation assumes the world will fit the questions we asked in advance. Humble evaluation assumes the world may be more interesting than our survey, and more slippery than our regression.
That humility has two dimensions.
1. Humility about what counts as an outcome
Traditional evaluation often begins with a fixed theory of change and a fixed list of indicators. That is useful, but it is also a kind of tunnel vision. Many of the effects that matter most are not initially visible because they are indirect, delayed, or socially embedded. People adapt. Institutions react. Benefits spill over. Harms migrate.
A community irrigation project, for example, may be designed to improve yields. But once water becomes more reliable, women may spend less time walking to distant sources, local disputes may decrease, and labor patterns may shift. If those changes are not actively sought, they disappear from the record. The intervention then looks simpler than it really is.
This is why outcome discovery matters. It allows evaluation to ask not only whether the intended result occurred, but what else changed in ways that are socially relevant.
2. Humility about how certain we should be
Even when the right outcome is identified, a second form of humility is necessary. Findings can be real and still be unstable. A treatment effect may survive one specification and vanish in another. A subgroup effect may look compelling until replication reveals that it was driven by a handful of observations. A policy may work in one setting and fail in another because the mechanism was context dependent all along.
Sensitivity analysis is the discipline of asking, How much would I need to change my assumptions before the conclusion changes? Replication asks an even more brutal question: Does the result persist when the test is repeated? Together, they keep us from mistaking convenience for truth.
The synergy between the two is powerful: discovery expands the map, and robustness checks prevent the map from becoming mythology.
A better mental model: the evaluation funnel
A useful way to connect these ideas is to think of evaluation as a funnel with two gates.
The first gate is breadth. Here the task is to cast a wide net and notice outcomes that matter, including those nobody predicted. This gate rewards attentiveness, qualitative immersion, and openness to surprise. It is the stage where unexpected patterns can emerge from interviews, field observations, administrative traces, or stakeholder narratives.
The second gate is discipline. Here the task is to narrow from plausible signals to credible claims. This gate rewards transparent assumptions, sensitivity analysis, and replication. It is the stage where we ask which findings are stable, which are fragile, and which are artifacts of our own methods.
Most institutions overweight one gate and ignore the other. Some are excellent at generating interesting stories but weak at checking whether those stories hold. Others are excellent at testing predefined indicators but incapable of seeing what falls outside the frame.
The funnel model suggests a different workflow:
- Open the field: ask what changed, not only what you expected to change.
- Cluster the signals: identify recurring outcomes, patterns, and plausible mechanisms.
- Stress test the claim: vary assumptions, sample definitions, model specifications, and time windows.
- Replicate where possible: see whether the effect survives a fresh look.
- Promote only the durable findings: keep the surprising, but only after the suspicious has been addressed.
This is not just a technical sequence. It is a philosophy of inquiry. It says that useful knowledge is not born fully formed. It is first noticed, then interrogated, then earned.
The best evaluations do not merely measure impact. They convert surprise into evidence and evidence into confidence, one filter at a time.
Why this matters in the real world
The appeal of this combined approach becomes obvious when you look at how decisions are actually made.
Imagine a nonprofit launches a youth employment program. The original theory is simple: training plus placements should increase income. After several months, the expected effect on earnings is modest. A shallow evaluation would stop there. But a broader inquiry reveals something more interesting: participants are also building networks that help them find better jobs later, some are delaying migration, and local employers are changing hiring practices after seeing the talent pool.
That is only the first half of the story. Now comes the hard part. Are those additional effects consistent across cohorts? Do they persist after adjusting for attrition? Are they driven by a few unusually active participants? Would a different threshold for “employment” alter the conclusion? A sensitivity analysis might show that the network effect is robust, while the migration effect is fragile and depends on one seasonal survey. Replication might confirm the former and weaken the latter.
The result is not a simplistic yes or no. It is a layered understanding: the program may not have moved earnings immediately, but it did alter social capital in a durable way. That is a more actionable result than a single headline metric, because it tells policymakers what mechanism deserves investment.
The same logic applies in public health, agriculture, education, or governance. In each case, policy is often judged by one visible target, while the real change unfolds through a web of second-order effects. Methods that combine outcome discovery with robustness testing are better suited to that reality because they can distinguish between:
- the intended effect,
- the emergent effect, and
- the fragile effect that only looks real.
Those are not the same thing, and confusing them can be costly.
The deepest lesson: evaluation is an ethics of attention
There is also a moral dimension here that is easy to miss.
When evaluators look only for predefined indicators, they implicitly say that only certain lives and outcomes count. That can erase what communities themselves experience as important. A program may reduce a measurable symptom while intensifying an unmeasured burden. It may hit the target and miss the lived reality.
At the same time, when evaluators announce findings without testing their fragility, they impose certainty on others. Decision makers then allocate resources, scale programs, or close initiatives based on conclusions that may not deserve that authority. In this sense, weak robustness is not just a technical flaw. It is a form of overclaiming.
The combination of outcome discovery and sensitivity analysis offers a healthier ethic. It says: listen broadly, then claim carefully. Notice what matters to people, then verify what you think you heard. Resist the temptation to compress complexity into a single number, but also resist the temptation to call every pattern a finding.
That ethic is especially important in settings where power is unequal. If donors, governments, or researchers define the outcome space too narrowly, marginalized effects are invisible. If they then overstate the certainty of their preferred result, those affected have little room to contest the narrative. Humble evaluation is therefore not only smarter, it is fairer.
Key Takeaways
-
Separate discovery from verification. First ask what changed that matters. Then ask whether the conclusion is stable under different assumptions.
-
Treat unexpected outcomes as data, not noise. The most important effects are often indirect, delayed, or outside the original theory of change.
-
Use sensitivity analysis as a credibility test. If a result disappears when a reasonable assumption changes, it should be reported as fragile, not definitive.
-
Replicate before you scale. A result that matters in one context is not automatically a general rule. Fresh tests reveal whether the mechanism is portable.
-
Adopt the two gate mindset. Broadening the outcome space and narrowing toward robust conclusions are both necessary stages of serious evaluation.
Conclusion: the point is not to be right, but to be trustworthy
The deepest connection between seeing outcomes that were never planned for and stress testing the findings that seem most convincing is this: both reject the fantasy of perfect knowledge.
Real evaluation does not begin with certainty and refine it. It begins with partial vision, then tries to widen what it can see, and finally tests what it thinks it knows until the remaining claims are worth keeping. That may sound slower than the usual hunt for quick answers. It is. But it is also how knowledge becomes durable.
So the question is not whether we should prefer discovery or rigor. The better question is: how do we design evaluation systems that are curious enough to find the unexpected, and skeptical enough to deserve belief?
If we get that balance right, evaluation stops being a ritual of proof and becomes something more valuable: a method for learning what is true in a world that refuses to be simple.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣