Before AI Finds the Pattern, Decide What Counts as Evidence

Ilaria Vergine

Hatched by Ilaria Vergine

Aug 08, 2026

10 min read

91%

0

What if the most dangerous mistake in qualitative research is not misunderstanding the data, but asking artificial intelligence to find patterns before deciding what counts as relevant?

A modern research system can identify which codes occur together, surface recurring themes, compare large bodies of text, and even help draft an interpretation. That speed is genuinely useful. Yet speed creates a subtle epistemic danger: the faster we discover relationships, the easier it becomes to mistake visible relationships for meaningful ones.

The problem is not that AI sees patterns. The problem is that patterns acquire meaning only inside a defined field of attention. Before asking what appears together, a researcher must decide whose experiences belong in the inquiry, which phenomenon is being examined, what context matters, and what evidence is allowed to speak.

This leads to a broader thesis: the quality of AI assisted qualitative research depends less on the sophistication of its pattern detection than on the discipline of its inclusion criteria. AI can accelerate interpretation, but it cannot substitute for the prior act of deciding what the interpretation is about.

The Pattern Is Not the Phenomenon

Imagine feeding an AI system thousands of interview transcripts about access to healthcare. It reports that the codes “transportation,” “missed appointments,” and “anxiety” frequently occur together. This sounds like a useful finding. Perhaps transportation barriers produce anxiety, which then contributes to missed appointments.

But the same cooccurrence could mean several different things. Patients may be anxious because they fear being judged for arriving late. They may miss appointments because public transportation is unreliable. They may mention transportation and anxiety together because interviewers routinely asked about both. Or the pattern may be concentrated among rural participants and disappear entirely in urban settings.

The machine has detected a relationship in language. It has not yet established the phenomenon that relationship represents.

This distinction is easy to overlook because cooccurrence feels objective. A code appears beside another code, perhaps hundreds of times. A visual network forms. A cluster grows darker and more central. The output looks less interpretive than a human conclusion, but the apparent neutrality is misleading. Every pattern depends on earlier choices about which documents were collected, which participants were included, how concepts were coded, and what counts as a meaningful connection.

AI can tell us what is adjacent in the record. Research must determine what is significant in the world.

This is why inclusion criteria are not administrative details. They are the architecture of the question. They determine the boundaries within which a pattern can be interpreted responsibly.

Inclusion Criteria Are a Theory of Relevance

In qualitative research, inclusion criteria are often treated as practical filters. Researchers specify eligible participants, settings, concepts, outcomes, and sources of evidence. The language can sound procedural, but the underlying act is philosophical: we are declaring what belongs to the problem and what does not.

Suppose the research question concerns how first generation university students experience academic belonging. If the study includes all students, regardless of educational background, the resulting themes may be rich but conceptually blurred. If it includes only students who have completed at least one year, it may miss the shock of entering university. If it excludes part time students, working students, or students studying remotely, it may confuse one institutional experience with the experience of belonging itself.

Each boundary changes the possible findings. It does not merely reduce the sample. It changes the shape of reality that becomes visible.

The same principle applies to the concept under examination. A study about “trust” must clarify whether it means trust in clinicians, institutions, scientific evidence, or personal networks. Without that clarification, an AI system may combine statements that use similar language but refer to different phenomena. It might discover that “trust” cooccurs with “family,” but the relationship could represent reliance, skepticism toward institutions, cultural obligation, or emotional support.

A broad concept can produce a large number of associations while producing little understanding. A precise concept may generate fewer patterns, but those patterns are easier to interpret and test.

This suggests a useful mental model: the research question acts like a lens, while inclusion criteria determine the lens’s aperture. A wide aperture lets in more experiences, contexts, and forms of evidence. That can reveal unexpected variation, but it can also reduce contrast. A narrow aperture allows detailed focus, but risks excluding the very differences that matter.

Good research does not seek the widest or narrowest aperture. It seeks a deliberately chosen one, with the reasons made visible.

The Context Problem: Same Words, Different Worlds

AI is particularly good at finding linguistic regularities across large collections of text. But qualitative meaning is often contextual, and context can be hidden inside what appears to be a common code.

Consider the word “safety.” In a workplace study, it may refer to physical protection from injury. In a study of domestic violence, it may refer to secrecy, housing, financial independence, or the ability to communicate without surveillance. In a study of online communities, it may mean protection from harassment or the freedom to disclose an identity. The same code can point to different realities.

Geographic location, culture, race, gender, age, institutional setting, and historical moment can all alter the meaning of a theme. If these dimensions are not defined before analysis, an algorithm may flatten them into a single category. It can treat repeated language as shared experience when the repetition actually conceals a crucial difference.

For example, suppose a qualitative analysis of maternal healthcare identifies frequent cooccurrence among “respect,” “communication,” and “choice.” In one group, “choice” may mean being offered treatment options. In another, it may mean being listened to when refusing an intervention. In a third, the absence of choice may be tied to language barriers or racialized assumptions about pain. The codes overlap, but the mechanisms do not.

Context therefore should not be treated as metadata added after the analysis. It is part of the meaning of the data. A research system that records the statement but loses the setting may preserve words while discarding evidence.

A practical consequence follows: whenever an AI system identifies a prominent cooccurrence, the researcher should ask three questions.

  • For whom does this pattern appear?
  • Under what conditions does it appear?
  • What alternative meaning could the same words have in another context?

These questions convert a pattern from a conclusion into an invitation to investigate.

AI as a Cartographer, Not a Judge

The most productive role for AI in qualitative inquiry is not to act as an oracle, and not even primarily as a coauthor. It is to function as a cartographer of the evidence.

A cartographer can map roads, elevations, boundaries, and population density. A map can reveal an unexpected route or show that two places are closer than assumed. But a map cannot decide whether a road is safe, whether a boundary is just, or why residents avoid a particular neighborhood. Those judgments require history, situated knowledge, and human accountability.

Similarly, AI can help researchers identify code cooccurrences, unusual cases, dense clusters, changes across time, and passages that deserve closer reading. It can reduce the burden of searching and make large qualitative collections more navigable. It can also expose the researcher’s assumptions by finding evidence that does not fit the emerging interpretation.

But AI should not be allowed to silently determine the scope of inquiry. If it proposes categories, those categories need to be examined against the research question. If it synthesizes themes, the researcher should be able to trace each theme back to participants, contexts, and contradictory cases. If it drafts prose, the prose should remain subordinate to the evidence rather than becoming a polished substitute for it.

A useful division of labor looks like this:

  • The researcher defines the phenomenon. What exactly is being studied, and why?
  • The research design defines the field of relevance. Which participants, contexts, and forms of evidence belong?
  • AI searches and compares within that field. What patterns, absences, and cooccurrences merit attention?
  • The researcher interprets the pattern. What could explain it, and what evidence would challenge that explanation?
  • Participants and affected communities constrain the conclusion. Does the interpretation make sense in the lives it claims to describe?

This is not a rejection of AI. It is a refusal to confuse computational assistance with epistemic authority.

The Hidden Cost of Unexamined Scope

Poorly defined inclusion criteria create two kinds of error. The first is obvious: relevant evidence is excluded. A study of digital exclusion that considers only broadband access may miss device sharing, digital literacy, disability, language, and the bureaucratic design of online services.

The second error is more subtle: irrelevant evidence is included and then allowed to distort the pattern. If a study combines patients, clinicians, administrators, and policymakers under a single code for “barriers,” an AI system may produce a highly coherent theme. Yet coherence can be manufactured by collapsing distinct viewpoints into one analytical category.

This is the false coherence problem. The more varied the evidence, the more tempting it is to search for a unifying theme. But a theme that explains everything may actually explain nothing. It may be a broad label placed over disagreements that should remain visible.

The remedy is not to eliminate variation. It is to model it. Researchers can ask AI to compare code relationships by subgroup, setting, time period, or source type. They can distinguish a pattern that is widespread from one that is concentrated. They can look for negative cases, where the expected relationship is absent, and disconfirming passages, where participants explicitly reject the researcher’s interpretation.

This produces a more rigorous workflow: define, map, stratify, challenge, revise.

First, define the concepts and boundaries. Next, map the patterns inside those boundaries. Then stratify the findings by the contexts that might alter meaning. Challenge the most attractive interpretation with contrary evidence. Finally, revise the scope or the explanation when the evidence demands it.

The key is that AI enters most powerfully in the mapping and challenging stages, not as a replacement for defining or judging.

A Practical Protocol for Responsible Pattern Discovery

Researchers and research teams can apply this approach immediately by treating every AI generated theme as a claim with conditions attached.

Before analysis, write a scope statement that answers five questions:

  1. Who is included, and why? State the relevant participant characteristics and explain exclusions.
  2. What concept is being examined? Define its boundaries and distinguish it from neighboring concepts.
  3. What phenomenon is of interest? Describe the experience, process, or interaction the inquiry aims to understand.
  4. Which contexts matter? Specify geography, culture, institutions, identities, and time periods that may shape meaning.
  5. What evidence can speak? Decide whether the corpus includes interviews, observations, letters, guidelines, prior reviews, or other forms of evidence.

During analysis, ask the AI system not only for frequent cooccurrences but also for:

  • cooccurrences by participant group and context
  • codes that should appear together but do not
  • passages that contradict the dominant theme
  • examples where the same code has different meanings
  • sources that are overrepresented or underrepresented
  • cases that fall just outside the inclusion boundary

After analysis, require an evidence trail. Every major theme should show where it came from, which participants expressed it, which contexts shaped it, and what evidence complicates it. A polished synthesis without that trail is rhetorically impressive but methodologically fragile.

The deepest discipline is to preserve the distinction between discovery and justification. AI is excellent at helping researchers discover possible relationships. Justification requires a transparent scope, contextual interpretation, comparison with alternatives, and an honest account of uncertainty.

Key Takeaways

  • Define relevance before searching for patterns. Inclusion criteria determine what an AI system is allowed to treat as evidence.
  • Treat cooccurrence as a clue, not a conclusion. Two codes appearing together may reflect causation, shared context, interview design, or entirely different meanings.
  • Make context analytical. Geographic, cultural, racial, gendered, institutional, and temporal differences can transform the meaning of the same code.
  • Use AI to map and challenge interpretation. Ask it to find contradictions, absences, subgroup differences, and alternative readings, not only dominant themes.
  • Demand traceability. A credible synthesis should connect every important claim to participants, contexts, source types, and disconfirming evidence.

The future of qualitative research will not be decided by whether machines can recognize themes. They already can. The more important question is whether researchers can maintain a disciplined relationship with the boundaries that make themes meaningful.

A pattern is never simply found. It is produced by a meeting between evidence and attention. Inclusion criteria decide what enters that meeting. AI can illuminate the relationships inside the room, but it cannot decide who was left outside, what the room represents, or whether the conversation has been understood.

The most advanced research system is still vulnerable to a primitive mistake: answering a precise question about the wrong world.

The responsible researcher therefore does not ask AI to replace judgment. They ask it to make judgment more visible, more testable, and more difficult to fool.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣