Why Good Evaluation Starts by Hunting for Change, Not Measuring Activity

Anemarie Gasser

Hatched by Anemarie Gasser

Jul 11, 2026

10 min read

68%

0

The uncomfortable truth about impact

Most people think evaluation begins with a neat question: Did the program work? That sounds sensible, even scientific. But it is also the wrong starting point in many real-world settings, because it assumes change is obedient, visible, and easy to trace. In complex social systems, outcomes rarely arrive on command. They emerge unevenly, through indirect pathways, with multiple actors shaping what counts as success.

That is why the most useful evaluations do not start by counting activities or even by fixing on preselected indicators. They start with a more difficult question: What changed, for whom, and how do we know our work mattered? The answer often cannot be found in a single survey line or a tidy dashboard. It has to be assembled from evidence, judgment, context, and interpretation.

This is the deeper tension at the heart of outcome evaluation and outcome harvesting. One wants structure, comparability, and rigor. The other insists that some of the most important effects are not known in advance, and may not look like outcomes at all until we learn how to recognize them. The real challenge is not choosing between rigor and openness. It is building a method that can hold both.

The deepest evaluation mistake is to measure what is easiest to count instead of what is actually changing.


Why outputs seduce us and outcomes unsettle us

Outputs are comforting. They are visible, immediate, and easy to assign. Trainings held, reports published, trees planted, loans disbursed, workshops delivered. These are the kinds of things organizations can produce on schedule, with pride and certainty. But outputs are not transformation. They are only evidence that something was done.

Outcomes are harder. They describe changes in behavior, relationships, practices, decisions, and institutional norms. A policymaker revises a rule because a campaign shifted public pressure. Farmers alter planting methods after a demonstration plot changes local expectations. A school administrator adopts a new attendance system because a pilot exposed a persistent bottleneck. These changes matter more than the activity that preceded them, yet they are harder to predict and harder to prove.

This is where conventional evaluation often breaks down. If you begin with fixed indicators too early, you can end up rewarding compliance over curiosity. Organizations optimize for what they already know how to count, not what they are actually trying to change. The result is a familiar but dangerous illusion: a project can look successful on paper while failing to alter the world in any meaningful way.

Outcome harvesting responds to that problem by flipping the sequence. Instead of first asking whether a prewritten plan was implemented, it asks whether any significant changes occurred and then works backward to understand the contribution. That shift may sound subtle, but it is profound. It acknowledges that in complex environments, change is often discovered before it is explained.

Consider a public health initiative in a district with multiple overlapping actors. A rigid evaluation might focus on whether 10 clinics were trained, 5 manuals distributed, and 3 workshops delivered. But the more interesting question is whether health workers began triaging patients differently, whether local leaders started referring women earlier in pregnancy, or whether a neglected referral pathway suddenly became functional. Those are the changes that matter. They are also the changes most likely to be missed if the evaluation only looks for the visible footprint of planned activities.


A better model: evaluation as detective work

The most helpful way to understand outcome harvesting is not as a checklist, but as a kind of forensic inquiry. A detective does not begin with a theory and then ignore the evidence. The detective begins with traces: testimony, anomalies, timelines, documents, and contradictions. The job is to reconstruct a plausible account of what happened and who contributed.

Evaluation in complex systems works the same way. Outcomes leave traces in language, behavior, policy, budgeting, routines, and relationships. They appear in meeting minutes, in changes to procedure, in unexpected alliances, in people saying, “We never used to do it this way.” The evaluator’s job is to identify these traces, verify them, and piece together whether the intervention plausibly contributed.

This framing is powerful because it resolves a false choice. Traditional evaluation often treats rigor as control: define the variables, lock the design, and minimize ambiguity. Harvesting treats rigor as disciplined interpretation: collect diverse evidence, test claims against reality, and be explicit about what is known, inferred, and uncertain. Both care about validity. They just protect it differently.

That difference matters because social change is rarely linear. A mentoring program may not produce immediate test score gains, yet it may shift how teachers collaborate. A civic initiative may not pass a new law, yet it may normalize a vocabulary that later becomes politically powerful. A climate project may not reduce emissions instantly, yet it may create the institutional habit of measuring them seriously for the first time. In each case, the outcome is not merely the end result. It is a shift in the system’s capacity to change.

This suggests a broader insight: outcomes come in layers.

  1. Immediate behavioral outcomes: people do something differently.
  2. Relational outcomes: trust, coordination, and legitimacy shift.
  3. Institutional outcomes: rules, routines, or incentives change.
  4. Cultural outcomes: what seems normal, possible, or acceptable begins to move.

A good evaluation should not collapse all four layers into one metric. It should ask which layer is changing, how, and why that matters.


Contribution is not attribution, and that is the point

One of the most persistent anxieties in evaluation is attribution. Did our program cause the change? Can we prove it? This question is understandable, but in many contexts it is too blunt to be useful. Real-world change is usually produced by many forces at once: policy shifts, economic shocks, local leadership, media attention, timing, and plain chance.

Outcome harvesting offers a different standard: contribution. The question is not whether you were the sole cause, but whether there is credible evidence that your work helped bring about the observed change. That may sound less satisfying to people seeking clean causality, but it is far more honest in complex environments.

Imagine a coalition working to improve girls’ access to secondary education. A year later, several things have changed: a district official introduces a new transport subsidy, school attendance rises, and parents report less skepticism about adolescent girls traveling alone. Did the coalition do this alone? Probably not. But if its advocacy helped create the pressure, language, and alliances that made the subsidy politically possible, then its contribution is real and should be captured.

This is where evaluation becomes more than reporting. It becomes a theory of change made testable through evidence. Not a static theory frozen at the proposal stage, but a living one that can be revised as the world responds. The most useful evaluations do not merely tell us whether a program succeeded. They help us understand what kinds of interventions are capable of producing change in environments like this one.

That is a higher-order question. It turns evaluation into learning, and learning into design intelligence.

Contribution asks a more adult question than attribution: not “Who gets all the credit?” but “What actually helped change happen?”


The hidden discipline inside open-ended methods

A common misconception is that approaches like outcome harvesting are loose because they are open-ended. In reality, they require serious discipline. Openness without discipline becomes storytelling. Discipline without openness becomes bureaucracy. The power lies in combining them.

A strong harvesting process usually demands several kinds of rigor:

  • Outcome specification: the change must be described concretely, not vaguely. “Improved empowerment” is too fuzzy. “Local women began speaking in council meetings and were invited onto a budget committee” is testable.
  • Evidence triangulation: one person’s claim is not enough. Look for documents, interviews, observations, and corroborating testimony.
  • Contribution analysis: trace the sequence of events and ask whether the intervention plausibly influenced the change.
  • Sensemaking: compare multiple outcomes to detect patterns, surprises, and missing links.

This is not just a better methodology. It is a better attitude toward reality. It refuses the temptation to make the world flatter than it is.

There is also an ethical dimension here. When evaluators force complex change into narrow indicators, they can erase the most meaningful effects of a program. A youth initiative may be dismissed because it did not generate quick employment numbers, even if it reduced violence, improved confidence, and built local leadership. A conservation effort may be judged a failure because it did not immediately restore biodiversity, even if it altered land stewardship practices and enabled future recovery. If we only reward short-term, easily measurable outputs, we systematically undervalue the kinds of change that are hardest to generate and most important to sustain.

A more humane evaluation practice would ask: What kind of change is this intervention realistically capable of influencing, and what evidence would reveal that change before it matures into its final form? That question shifts evaluation from policing to learning.


The real prize: seeing change before it becomes obvious

The greatest advantage of outcome-focused evaluation is not accountability, though accountability matters. It is early perception. Organizations that can detect weak signals of change before those signals become obvious gain a strategic advantage.

Think of a gardener who notices tiny green shoots before the whole bed is visible. A novice sees only bare soil and concludes nothing is happening. An experienced gardener sees the first signs of life and knows where to keep watering. Evaluation should work the same way. Its job is not merely to certify success after the fact. It is to help practitioners recognize promising shifts while there is still time to reinforce them.

This is especially important in systems where feedback loops are slow. Educational reform, governance reform, social norm change, and ecological regeneration all move on timelines that are painfully out of sync with funding cycles. If evaluation waits for final results, it often arrives too late to improve practice. But if it can detect emerging outcomes, it becomes a steering mechanism rather than a retrospective audit.

That requires a different mindset from everyone involved.

For funders, it means asking not just for end results, but for credible traces of change and a disciplined account of contribution. For practitioners, it means documenting shifts as they happen, even when they are messy or incomplete. For evaluators, it means listening for surprises, not only for confirmation. And for organizations, it means treating unexpected outcomes as data, not noise.

This may be the most important reframe of all: what does not fit the original plan may be the most important evidence of all. If the intervention produced a change nobody predicted, that does not make the evaluation less valuable. It may make it more valuable, because the world is revealing how it really works.


Key Takeaways

  1. Start with change, not activity. Ask what has actually shifted in behavior, relationships, institutions, or norms before counting how many things were done.
  2. Treat evaluation like disciplined detective work. Look for traces of change across multiple sources, then reconstruct a plausible story of contribution.
  3. Use contribution when attribution is too narrow. In complex systems, the useful question is whether your work helped produce the change, not whether it caused it alone.
  4. Name the layer of change. Distinguish between behavioral, relational, institutional, and cultural outcomes so you do not confuse early signals with final impact.
  5. Document surprises. Unexpected effects are not evaluation errors, they are often the first visible sign that a system is moving.

The final reframing

The most important thing evaluation can tell us is not whether a project followed its plan. It is whether the world is different in ways that matter, and whether our work was part of that difference. That is a harder question, but also a more honest one.

When we stop worshipping outputs and start tracing outcomes, we begin to see that impact is not something an organization manufactures on demand. It is something that emerges through relationships, timing, and contribution. The evaluator’s task is not to simplify that reality until it looks neat. It is to learn how to read it well.

And once you see evaluation that way, the purpose of the exercise changes. It is no longer about proving that you were busy. It is about discovering where change is taking root, so you can nurture it before it disappears into the background. In that sense, the best evaluation is not a verdict. It is an act of attention.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣