The Questions You Ask Decide What the Data Can Ever Mean
Hatched by Anemarie Gasser
May 29, 2026
10 min read
5 views
58%
The hidden problem is not data scarcity, it is question poverty
Most debates about evidence start in the wrong place. We argue about whether we have enough data, whether the data are reliable, or whether the numbers are biased. Yet the deeper failure often comes earlier: we ask thin questions, then expect rich answers.
That is the real tension at the heart of evaluation and policymaking. Data does not merely answer questions. It is shaped, narrowed, and sometimes distorted by the question that called it into being. A budget spreadsheet, a household survey, a clinic report, a community interview, and an observation of daily practice each reveal different realities, but only if the guiding question is precise enough to let each source speak.
This is why so many evaluations produce technically correct but practically useless conclusions. They measure what is easy to count, not what must be understood. They confuse the availability of indicators with the significance of insight. The result is an evidence culture that can become fluent in metrics while remaining strangely mute about meaning.
The quality of evidence is not determined first by how much you collect, but by how well you frame what matters.
The deepest challenge, then, is not choosing between qualitative and quantitative data, or between official records and lived experience. It is learning how to ask questions that are simultaneously answerable, consequential, and honest about uncertainty.
Why the wrong question can make good data useless
Consider a simple policy question: Did a job training program work? It sounds reasonable, even responsible. But it may already be too blunt to be useful. Work for whom, in what time frame, compared with what alternative, and by what definition of success? A participant may not get a job immediately but may gain credentials, networks, confidence, or a higher wage six months later. Another may get a job quickly but remain trapped in unstable work. If the question is too coarse, the answer will be misleading no matter how sophisticated the analysis.
This is why meaningful evaluation begins with question design, not data collection. A strong question is not merely a curiosity. It is a decision tool. It clarifies what tradeoffs matter, which outcomes count, whose perspective is being centered, and what kinds of evidence are needed to reduce ignorance rather than simply decorate a report.
A useful way to think about this is to separate three layers:
- Descriptive questions: What is happening?
- Causal questions: What difference did this intervention make?
- Normative questions: What should we value, protect, or improve?
Many evaluation failures happen because these layers are collapsed into one another. A program may be described as effective because one metric rose, but that says little about whether the change was caused by the program, whether it mattered in people’s lives, or whether it produced unacceptable side effects. Likewise, a community may report strong satisfaction, but that does not automatically tell us whether outcomes improved or whether the benefit was equitably distributed.
When questions are poorly formed, data often becomes a kind of alibi. It gives the appearance of rigor while avoiding the harder task of judgment.
Triangulation is not just about agreement, it is about perspective
The common story about triangulation is that multiple data sources increase confidence when they point in the same direction. That is true, but incomplete. The deeper value of triangulation is not simply convergence. It is perspective correction.
Imagine trying to understand a city by looking only at satellite imagery. You would know road layouts, land use, and perhaps traffic patterns, but not how safe the streets feel at night, how accessible the bus stops are, or which neighborhoods are socially connected despite being physically close. Now imagine adding commuter surveys, mobile phone traces, school attendance patterns, clinic data, and neighborhood interviews. These do not merely confirm one another. They reveal different dimensions of the same urban life.
This is the mistake many institutions make: they treat different kinds of evidence as if they were interchangeable when they are actually complementary. Administrative records excel at scale and consistency, but they often miss informal activity and excluded populations. Surveys can capture broad patterns, but they depend on question wording, recall, and response categories. Interviews can uncover meaning, but they are not designed to estimate prevalence. Observations can reveal practice, but they are limited in scope. Each source is a lens, not a mirror.
A mature evidence practice asks a better question than “What data do we have?” It asks, “What aspect of reality is each source structurally capable of seeing, and what does it systematically miss?” That question changes everything. It moves evaluation away from a simplistic hunt for one definitive dataset and toward a more disciplined ecology of evidence.
Good triangulation does not ask different sources to say the same thing. It asks them to reveal what the others cannot.
This is also why disagreement across sources should not be treated as failure. If survey data says service coverage is high, but interviews reveal persistent barriers for marginalized groups, that is not a contradiction to be papered over. It is a clue. The divergence points to a gap between formal access and lived access, between policy on paper and policy in practice.
The real unit of analysis is often the question, not the metric
Metrics are seductive because they feel concrete. Numbers compress complexity into something manageable. Yet any metric is only as meaningful as the question that defined it. If we ask about attendance, we may measure presence. If we ask about learning, we may need retention, application, confidence, and long term outcomes. If we ask about health care quality, we may need not just throughput but trust, continuity, and safety.
This suggests a powerful shift in mindset: a metric is not an answer, it is a proxy for a question. That distinction matters because proxies are always partial. They substitute for something more complex than they can fully capture. The mistake is to treat proxy success as conceptual success.
Think of a school district celebrating higher graduation rates. That could mean students are learning more, or it could mean standards changed, grade inflation rose, or credit recovery improved completion without deep mastery. The same number can sit atop very different realities. The number itself is not lying, but it is speaking a reduced language.
The task of evaluation is therefore not to maximize the number of indicators. It is to improve the mapping between the question and the evidence. That mapping becomes stronger when three conditions are met:
- The question is specific enough to guide measurement.
- The evidence source is appropriate to the type of claim being made.
- The interpretation makes explicit what the data cannot show.
This is a more honest form of rigor. It accepts that not all important things are equally measurable, and not all measurable things are equally important.
A policymaker who understands this can ask better questions such as:
- Which groups are benefiting, and which are still excluded?
- What changed because of the intervention, and what would likely have changed anyway?
- Which outcomes matter immediately, and which require time to emerge?
- What unintended effects are likely to be invisible in routine monitoring?
These questions do not eliminate ambiguity. They organize it. They create a structure in which multiple forms of evidence can be read intelligently rather than defensively.
A practical framework: ask in layers, then evidence in layers
The most useful evaluation questions are rarely single questions. They are layered questions. A layered question begins with the outcome we care about, then moves backward to causality, mechanism, equity, and context. This prevents the common error of asking one broad question and then expecting one dataset to answer it completely.
Here is a simple framework that can be applied to programs, policies, and organizational decisions:
1. What changed?
This is the descriptive layer. It asks for patterns, trends, distribution, and variation. Administrative data and routine monitoring are often strongest here.
2. For whom did it change?
This is the equity layer. It asks whether the effect is concentrated among already advantaged groups or whether benefits reached those most in need. Disaggregated data, targeted surveys, and qualitative accounts are especially useful here.
3. Why did it change?
This is the mechanism layer. It asks what process, behavior, or institutional feature produced the observed shift. Interviews, process tracing, and observation often reveal what dashboards conceal.
4. What else changed alongside it?
This is the systems layer. It asks about spillovers, unintended consequences, and tradeoffs. A program can improve one outcome while worsening another. Evidence must be wide enough to detect that.
5. What should count as success?
This is the normative layer. It asks whose values are being used to define success, and whether the definition itself needs to be debated. This layer is often omitted, but it is the most politically important.
This layered structure changes evaluation from a retrospective audit into an inquiry. It acknowledges that evidence is not just about proving something worked. It is about making wiser decisions under conditions of imperfect information.
A health intervention, for example, may show reduced hospital readmissions. That is valuable, but incomplete. Did it reduce readmissions by improving care coordination, or by making access harder? Did it help the sickest patients or mainly those already able to navigate the system? Did it reduce costs while increasing stress on caregivers? Each layer requires a different kind of evidence.
The same is true in education, social services, climate policy, and organizational change. The better the question, the more the data can do. The worse the question, the more even excellent data will disappoint.
The most important skill is not data literacy, it is evidence judgment
Data literacy teaches people how to read charts, interpret regression outputs, or understand margins of error. Necessary as that is, it is not enough. The deeper competence is evidence judgment, the ability to decide what kind of evidence is appropriate for what kind of claim.
This is a distinctly human skill because it involves values, context, and epistemic humility. A randomized trial can estimate average treatment effects, but it cannot tell you whether a result matters enough to justify adoption in a particular setting. A focus group can surface hidden concerns, but it cannot establish how common those concerns are. A dashboard can show performance trends, but it cannot explain why staff are gaming a metric or why users have stopped showing up.
Evidence judgment means learning to ask:
- What claim am I actually making?
- What would count as strong evidence for that claim?
- What alternative explanations remain plausible?
- What voices are missing from the dataset?
- What would I still want to know before making a consequential decision?
This kind of thinking resists the false comfort of certainty. It treats evidence as a conversation among methods, not a contest in which one method defeats the others. It also pushes institutions to become more transparent about uncertainty. That is not weakness. It is maturity.
A hospital deciding whether to adopt a new discharge protocol does not need a single number that settles everything. It needs a coherent evidentiary picture: readmission data, patient interviews, staff workflow observations, subgroup analyses, and cost implications. The point is not accumulation for its own sake. The point is fit.
When evidence fits the question, it becomes actionable. When it does not, it becomes theater.
Key Takeaways
- Start with the question, not the dataset. If the question is vague, the evidence will be vague too.
- Treat every metric as a proxy. Ask what it captures, what it misses, and what it may be distorting.
- Use triangulation for perspective, not just confirmation. Different sources should reveal different parts of the same reality.
- Ask layered questions. Separate what changed, for whom, why, what else changed, and what should count as success.
- Practice evidence judgment. Match the type of claim to the type of evidence, and be explicit about uncertainty.
Conclusion: the question is the intervention
We often think of evidence as something that arrives after action, a tool used to evaluate whether the action worked. But in practice, the question asked at the start is already an intervention. It shapes what gets measured, what gets noticed, whose experiences count, and which truths remain invisible.
That is why the most important decision in evaluation is not choosing a dashboard, a survey, or an interview protocol. It is deciding what kind of world you are trying to see. Once that is clear, data becomes more than a pile of indicators. It becomes a disciplined way of learning from reality.
The next time a policy report promises answers, pause before the charts. Ask whether the real issue is not the lack of data, but the lack of a question worthy of the evidence. Because in the end, good decisions are not made by collecting more information alone. They are made by asking better questions, then letting the right kinds of data argue back.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣