The Hidden Discipline Behind Good Questions: How Evaluation Becomes a Theory of Change
Hatched by Anemarie Gasser
Jul 21, 2026
9 min read
1 views
91%
The Question Before the Question
What if the difference between a weak plan and a transformative one is not the size of the goal, but the quality of the questions asked at the start?
Most people treat questions as tools for getting information. But in serious decision making, especially in evaluation and strategy, questions do something more powerful: they define what counts as success, what evidence matters, and which assumptions are allowed to survive. A poorly formed question can make a project look busy while drifting nowhere. A well formed question can force clarity before a single step is taken.
That is why asking meaningful evaluation questions is not a technical afterthought. It is the first act of design. And when that discipline is connected to a theory of change, something important happens. The question stops being a one time instrument for judging outcomes and becomes a way of mapping causality, revealing where a system can actually be influenced.
In other words, evaluation is not just about measurement. At its best, it is a method for thinking.
Why Bad Questions Produce Expensive Confusion
A weak question sounds harmless because it is broad, polite, and easy to agree with. But broadness often hides ambiguity. If you ask, “Did the program work?”, you may get a confident answer that is almost useless. Worked for whom? Compared to what? Over what time horizon? Under what conditions? At what cost?
This is the central danger: bad questions collapse complexity too early. They make it seem as if reality can be reduced to one number, one headline, or one yes or no. Yet most interventions unfold inside messy systems where causes interact, feedback loops appear, and unintended consequences matter as much as intended ones.
Consider a literacy program for children. A weak question might ask whether reading scores improved. A stronger question asks which children improved, by how much, in which classrooms, after which teaching practices, and whether gains persisted six months later. A still stronger question asks what mechanism likely produced the improvement: teacher training, better materials, parental engagement, or increased attendance.
That final shift matters because it changes the value of the evaluation. Now it is not merely judging performance. It is learning how change happens.
The best evaluation questions do not just test an intervention. They expose the causal logic that makes the intervention plausible.
This is where many organizations get trapped. They ask for proof when they really need insight. Proof tells you whether something happened. Insight tells you why it happened, when it happened, and whether it is likely to happen again.
A Theory of Change Is Really a Theory of Questions
A theory of change is often presented as a diagram: inputs, activities, outputs, outcomes, impact. That diagram is useful, but incomplete. Its real purpose is not to decorate a report. It is to surface the assumptions that connect action to result.
A program without a theory of change is like a bridge without engineering calculations. It may look convincing from a distance, but it has not yet answered the critical question: how will weight travel from one side to the other without collapse?
A strong theory of change does not merely say what you will do. It says what must be true for your strategy to work. That means it turns hidden assumptions into testable questions. For example:
- If we train teachers, will they actually change classroom behavior?
- If teachers change behavior, will students respond differently?
- If students respond differently, will family or institutional conditions allow the gains to persist?
- If the gains persist, is the change large enough to matter at the system level?
Each question is a hinge. If the hinge fails, the door does not open. That is what makes the theory of change such a powerful companion to evaluation. It helps distinguish between the parts of a strategy that are merely desirable and the parts that are structurally necessary.
This gives us a deeper framework: evaluation questions should be derived from causal uncertainty. The point is not to ask every possible question. The point is to ask the questions that sit where the chain of change is weakest, most uncertain, or most contested.
That is a radically different mindset from checklist evaluation. Checklist evaluation asks, “Did we complete the steps?” Causal evaluation asks, “Did the steps actually matter?”
The Three Levels of a Meaningful Question
To see the difference, it helps to think of evaluation questions in three levels.
1. Descriptive questions
These ask what happened. They are necessary, but they are only the beginning.
Examples: How many people attended? How many reached the target threshold? What services were delivered?
2. Explanatory questions
These ask why it happened. They are the bridge from reporting to learning.
Examples: Why did one site outperform another? Which conditions made participation higher? What changed in behavior, not just in attendance?
3. Strategic questions
These ask what should happen next.
Examples: Which part of the model deserves scaling? Which assumption is the riskiest? Where would a small change create the largest downstream effect?
Most organizations spend too much time at level one and too little at level three. They collect descriptive data because it is easy to count. But the real value of evaluation appears when evidence helps decide where to invest, what to stop, and what to redesign.
This matters because systems do not reward effort evenly. Some inputs are noise. Some are leverage. Meaningful questions help you tell the difference.
A useful way to think about this is to ask: What would I need to know in order to change my mind or change my strategy? If a question cannot inform a decision, it is probably decoration rather than inquiry.
The Hidden Power of Assumptions
The hardest part of any theory of change is not drawing the arrows. It is examining the assumptions inside the arrows.
Assumptions are the places where confidence often masquerades as logic. They sound like common sense until they are tested. We assume people will use a service once it exists. We assume incentives will change behavior. We assume people understand, trust, and can access what is being offered. In practice, any one of these may fail.
That is why the most important evaluation questions often begin with phrases like:
- What conditions must hold for this outcome to occur?
- Which assumption is most likely to fail first?
- What evidence would show that the causal pathway is breaking down?
- Who benefits, who does not, and why?
These are not merely analytic questions. They are humility questions. They force a team to stop pretending that intention is the same thing as effect.
Take a job training program. The obvious assumption is that training leads to employment. But perhaps the real bottleneck is transportation, childcare, or employer discrimination. If so, the intervention may be perfectly designed for the wrong problem. A meaningful question reveals this by asking not only whether the program worked, but where the pathway constricted.
Strategy improves when assumptions become visible enough to be challenged.
This is one reason evaluation can be uncomfortable. It threatens narratives that rely on linearity and control. Yet that discomfort is a feature, not a bug. A good question should slightly unsettle you, because it is entering the territory where certainty is weakest and learning is most valuable.
From Measurement to Learning: Evaluation as Navigation
The easiest mistake is to think that evaluation happens after action. In reality, evaluation is part of navigation. A sailor does not check the map only after the ship arrives. They continuously compare position, course, wind, and destination.
That analogy matters because it changes how we think about evidence. Evidence is not a verdict handed down at the end. It is a feedback signal used during motion. And the best feedback systems are built around questions that are specific enough to guide action but flexible enough to learn from surprise.
For example, instead of asking, “Did the initiative reduce poverty?”, a more navigable question might ask:
- Which subgroup experienced the largest improvement?
- What changed first: income, stability, or resilience?
- Did local conditions amplify or weaken the effect?
- What part of the model appears most transferable?
Now the evaluation becomes operational. It informs where to steer next.
This is especially important in complex systems, where the same intervention can produce different results in different contexts. In one neighborhood, a community health initiative may succeed because local leadership is strong. In another, it may stall because trust is low. A meaningful question does not pretend these differences are inconvenient. It treats them as the point.
The deeper lesson is that context is not a complication added to the model. Context is part of the model.
The Most Useful Questions Are Designed Like Experiments
One way to improve evaluation questions is to think like an experimentalist, even when running a non experimental program. Not because everything must become a randomized trial, but because experiments force clarity.
An experiment asks: what specific change am I trying to detect, and what evidence would make that change credible?
That discipline can transform vague questions into sharp ones.
Instead of: Did the mentorship initiative help?
Ask: Did participants show greater persistence, better problem solving, or stronger belonging than similar participants who did not receive mentorship?
Instead of: Was the policy effective?
Ask: Which behavior changed, in which group, under which implementation conditions, and at what cost relative to alternative approaches?
This kind of questioning has a hidden benefit: it reveals when the intervention itself is not the central variable. Sometimes the core issue is not the program design but the fidelity of implementation. Sometimes it is not fidelity but uptake. Sometimes it is not uptake but the surrounding system.
That means good evaluation questions do double duty. They assess outcomes, and they diagnose where the causal chain is fragile.
A practical mental model is to ask four questions in sequence:
- What changed?
- For whom did it change?
- Why did it change?
- What does that imply for future action?
This sequence moves from observation to interpretation to decision. It is simple enough to use and deep enough to prevent superficial reporting.
Key Takeaways
- Start with causal uncertainty, not curiosity alone. Ask where your biggest assumptions might fail.
- Use three layers of questions: descriptive, explanatory, and strategic. Do not stop at the first layer.
- Make context visible. If an intervention works in one setting and not another, context is part of the explanation.
- Treat evaluation as navigation. Evidence should help you steer, not just judge.
- Design questions so they can change decisions. If a question would not alter what you do next, refine it.
The Real Test of a Good Question
A meaningful evaluation question is not one that makes you sound rigorous. It is one that makes your work more truthful.
Truthful work is harder to manage because it refuses the comfort of simple narratives. It may tell you that your favorite activity is not moving the needle. It may show that a small, boring process change matters more than a high profile initiative. It may reveal that the intended impact is real, but only for a subset of people, under specific conditions.
That is not failure. That is information.
And once you see evaluation this way, the theory of change stops being a static diagram and becomes a living discipline of inquiry. You stop asking, “Did we do what we planned?” and start asking, “What must be true for this to work, where is that truth weakest, and what would we need to learn next?”
That shift is bigger than better measurement. It is the difference between performing accountability and practicing intelligence.
The most powerful systems are not the ones with the loudest answers. They are the ones that know how to ask the next necessary question.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣