The Health Data Paradox: Better Detection Begins by Admitting What We Cannot See
Hatched by SEAN SYLVIA
Aug 10, 2026
11 min read
1 views
93%
What if the most dangerous mistake in public health is not a wrong diagnosis, but a confident count?
A disease surveillance system may report ten cases when one hundred people are sick. An artificial intelligence system may label thousands of chest X rays with remarkable consistency, yet learn from labels that were never clinically verified. In both situations, the central problem is not simply missing data or imperfect algorithms. It is that health systems turn messy reality into official categories, and those categories acquire authority as they move through institutions.
This creates a deeper question: How can we know whether a health signal reflects what is happening in the world, or merely what our systems are capable of noticing?
The answer requires a shift in perspective. Surveillance and medical AI are often treated as separate domains: one belongs to epidemiology and government, the other to computation and clinical practice. But both are versions of the same enterprise. They are machines for converting uncertain observations into decisions. Their reliability depends less on the volume of information they collect than on the quality of the path connecting reality to action.
A health signal is never just found. It is produced by a chain of observation, classification, validation, and response.
Once that chain becomes visible, several familiar assumptions begin to look dangerous.
The invisible filter between illness and evidence
Imagine a person with tuberculosis who lives far from a clinic. The person develops symptoms, delays seeking care, reaches a facility with limited diagnostic capacity, receives an inconclusive test, and never appears in a national reporting system. From the perspective of official surveillance, the case may not exist.
This is not simply underreporting in the narrow sense. It is a measurement problem shaped by access, technology, incentives, and definitions. A surveillance system does not observe disease directly. It observes people who enter particular institutions, are examined with particular tools, satisfy particular case definitions, and are subsequently reported by people operating under resource constraints.
The difference matters because counts are often interpreted as if they were transparent windows onto reality. Yet a rise in reported pneumonia may mean that transmission has increased. It may also mean that more people visited clinics, laboratories received new equipment, clinicians changed their coding practices, or reporting became more complete. A fall may indicate improvement, or a breakdown in services.
The reported number is therefore better understood as the output of several interacting factors:
Observed cases = underlying illness multiplied by access, detection, classification, and reporting.
This is not a literal equation that can solve every epidemiological problem. It is a mental model that prevents a common error: treating observed cases as equivalent to actual cases. If diagnostic ability changes, the meaning of the count changes with it.
The same logic applies inside a medical image archive. A chest X ray does not arrive with its diagnosis visibly printed on the pixels. Its label may be derived from a radiologist's report, a billing code, a later clinical note, or an automated extraction from text. Each step introduces interpretation. Some public image collections contain labels that are inaccurate, incomplete, inconsistent, or never compared with a rigorous reference standard.
An AI model trained on such labels can become highly proficient at reproducing the archive's habits without becoming proficient at recognizing disease. It may learn that certain institutions use a particular phrase, that certain patient groups are more likely to receive a particular workup, or that an image from a particular ward carries a statistical clue. The model can be accurate relative to the labels while being wrong about the patient.
This is the medical equivalent of mistaking the number of reported cases for the number of infections. In both cases, the system is judged against a representation that may already contain the bias and uncertainty we hoped the system would overcome.
Two kinds of ignorance, one shared structure
Public health surveillance has developed an important distinction between indicator based data and event based data.
Indicator based systems collect standardized reports: diagnosed cases, laboratory confirmations, syndromes, and other defined measures. They are structured, comparable, and suitable for tracking trends. Their strength is consistency. Their weakness is that they can only see what enters the reporting pipeline.
Event based systems look for unusual signals outside formal reporting channels. They monitor media, online discussions, community reports, animal deaths, and unexpected clusters. A volunteer who notices poultry dying across several villages may identify a threat before laboratories confirm anything. A moderated online report or a local rumor may be clinically ambiguous, but it can still be valuable as an early warning.
These approaches are not competitors. They answer different questions. Indicator based surveillance asks, “How many confirmed or classified events meet our definition?” Event based surveillance asks, “What unusual pattern might deserve investigation?” One provides measurement; the other provides sensitivity to surprise.
Medical AI faces a parallel choice. A model can be trained to imitate existing labels, or it can be designed to expose uncertainty, compare images with reference patterns, and allow users to set a threshold for similarity or confidence. A model that labels a large archive automatically may increase efficiency, but its value depends on whether it helps distinguish known disease from unfamiliar cases.
This suggests a useful framework: every health information system needs both a census function and a search function.
The census function produces standardized, comparable records. It supports trend analysis, resource allocation, and accountability. The search function looks for anomalies, weak signals, and cases that do not fit the existing categories. It is designed not to describe the known world perfectly, but to notice when the known world may be incomplete.
A system with only a census function becomes blind to what its categories exclude. A system with only a search function becomes overwhelmed by noise. Robust health intelligence requires both, connected by a process of investigation.
Why explainability is more than a feature
Explainability is often presented as a matter of trust. If users can understand why an AI system produced a label, they may be more willing to use it. That is important, but insufficient. Explainability also functions as a quality control mechanism.
Suppose an AI model labels a chest X ray as showing pulmonary edema. If the system can identify comparable reference images and show which features influenced the result, a radiologist can ask whether the reasoning is clinically plausible. Perhaps the model is responding to a genuine pattern of fluid accumulation. Perhaps it is relying on an unrelated artifact, such as a marker placed on the image or a feature associated with a particular hospital.
The explanation does not automatically make the prediction correct. It makes the prediction inspectable.
That distinction has a direct analogue in surveillance. A weekly disease count without information about diagnostic methods, reporting completeness, geographic coverage, and recent system changes is difficult to interpret. A sudden increase is meaningful only when analysts can inspect the pathway that produced it.
A well designed surveillance report should therefore answer more than “how many?” It should make visible:
- Which populations were covered?
- What case definition was used?
- How many reports came from laboratory confirmation, clinical judgment, or syndromic classification?
- Did testing capacity or access change?
- Were duplicate reports removed consistently?
- Which areas or groups are likely to be missing?
- What unusual signals have not yet been confirmed?
These questions resemble the questions asked of an explainable model because they serve the same purpose. They reveal the difference between a conclusion and the evidence chain behind it.
Explainability is not merely a way to make a system persuasive. It is a way to make its failure visible before failure becomes policy.
The most useful systems do not hide uncertainty behind a single score. They organize uncertainty so that a human can decide what to investigate next.
The danger of clean data
There is a seductive idea that better health intelligence comes from cleaning the data until it is standardized. Cleaning is essential. Duplicate records, coding errors, and inconsistent classifications can distort any analysis. But cleaning can also erase information that looks inconvenient.
A rumor is not the same as a confirmed case. An animal death is not proof of a human outbreak. An ambiguous radiology label is not equivalent to a specialist verified diagnosis. Yet uncertainty is not the same as uselessness. A weak signal may be the earliest indication of a serious event.
The real goal is not to remove ambiguity. It is to label ambiguity accurately and route it appropriately.
Consider three layers of health information:
- Observation: someone reports unusual symptoms, an animal die off, or an unexpected image pattern.
- Interpretation: analysts or models classify the observation using available evidence.
- Confirmation: experts, laboratories, or additional data establish whether the interpretation holds.
Many systems collapse these layers into one field called “case” or “diagnosis.” Once that happens, users lose the ability to distinguish suspicion from confirmation. A more honest system preserves the layers and records how an observation changed status.
This principle can improve both surveillance and AI development. In a disease reporting network, an unverified community report should not be counted as a confirmed outbreak, but it should remain visible as a lead. In an imaging archive, an automatically inferred label should not be treated as a platinum standard, but it can support preliminary organization if its provenance and confidence are retained.
The difference is architectural. Instead of asking whether data are clean or dirty, ask: clean for what purpose? Data suitable for estimating prevalence may be unsuitable for detecting anomalies. Data suitable for training a retrieval system may be unsuitable for validating a diagnostic device. Fitness for purpose must be explicit.
From data pipelines to learning loops
The most reliable health systems are not static pipelines that collect information and produce reports. They are learning loops.
A pipeline moves in one direction: observation becomes record, record becomes statistic, statistic becomes decision. A learning loop adds feedback. Decisions alter testing and reporting. Investigations confirm or reject early signals. Confirmed findings improve case definitions, training data, and model evaluation. Failures are not merely errors to suppress; they are evidence about where the system cannot see.
A practical learning loop has five stages:
Sense: gather both formal measurements and informal signals.
Separate: preserve distinctions between observation, interpretation, and confirmation.
Triangulate: compare independent sources, such as laboratory results, clinical reports, community observations, and imaging evidence.
Act: direct scarce attention toward the signals with the greatest potential consequences, not merely the highest confidence scores.
Update: revise definitions, training data, workflows, and resource allocation based on what investigation reveals.
Triangulation is especially important because correlated sources can create an illusion of confirmation. Ten reports copied from the same news story are not ten independent observations. Likewise, an AI model validated against labels generated from the same flawed reporting process has not truly been tested against reality.
Independence is therefore a hidden asset in health information. A community report, a laboratory result, and a radiologist's review carry more evidentiary value when they arise through different pathways. Agreement across independent pathways is stronger than agreement within one administrative system.
This also changes how performance should be evaluated. Sensitivity, specificity, accuracy, and completeness remain important, but they are not enough. Systems should also be assessed for:
- Coverage: who and what can the system observe?
- Latency: how long between event and signal?
- Provenance: where did the information originate?
- Calibration: does confidence match reality?
- Resilience: what happens when a laboratory, network, or reporting channel fails?
- Discoverability: can the system notice events outside its established categories?
A model that is slightly less accurate but well calibrated and transparent may be more useful than a higher scoring model that fails silently on unfamiliar cases. A surveillance network that reports fewer confirmed cases but clearly marks gaps may support better decisions than one that publishes precise numbers without revealing its blind spots.
Key Takeaways
- Treat every health number as a measurement, not a fact. Ask what access, diagnostic capacity, definitions, and reporting practices shaped it.
- Build both census and search functions. Standardized indicators track known patterns, while event based systems detect anomalies that formal categories may miss.
- Preserve uncertainty instead of laundering it into certainty. Distinguish observations, interpretations, and confirmations in both databases and AI outputs.
- Make provenance and reasoning visible. Explanations, reference comparisons, and reporting metadata allow users to detect plausible looking errors.
- Evaluate the whole information chain. Measure coverage, latency, calibration, resilience, and ability to discover the unfamiliar, not only predictive accuracy or data volume.
The deepest lesson is that better health intelligence will not come from choosing between human judgment and automation, formal statistics and rumors, or clean data and messy data. It will come from designing systems that know what each kind of evidence can and cannot establish.
A community volunteer may notice the first sign of an outbreak. A laboratory may confirm it. A statistical dashboard may reveal its geographic spread. An imaging model may identify patients who need urgent review. None of these signals is sufficient alone. Their power comes from being connected without being confused with one another.
The future of public health and clinical AI therefore depends on a form of disciplined humility. We need systems that can count accurately, search widely, explain their conclusions, and expose the conditions under which those conclusions become unreliable.
The question is not whether our machines can see more. It is whether we can build institutions capable of recognizing what their machines, definitions, and databases are still unable to see.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣