Why Better Diagnosis Requires Less Trust in the Diagnosis
Hatched by SEAN SYLVIA
May 29, 2026
10 min read
3 views
88%
The Hidden Problem Is Not Just Bias, It Is Thresholds
What if the biggest source of error in medicine is not that clinicians disagree, but that they disagree for reasons they cannot see? A cardiologist, a radiologist, and now a model can all look at the same case and produce different recommendations. The instinct is to blame prejudice, inconsistency, or poor judgment. But there is a deeper possibility: some of that variation is not moral failure at all. It is diagnostic thresholding, the unseen rule each decision maker uses to decide when evidence is “enough.”
That matters because thresholds are shaped by more than knowledge. They are shaped by fear of missing disease, fear of overtesting, training quality, and sometimes bias. A clinician who is too eager to rule out disease may order more tests than necessary. A clinician who is too cautious may miss a dangerous condition. A model can inherit the same problem, but with a new twist: it may appear objective while quietly reproducing the biases and threshold habits embedded in the data and prompts it learned from.
The real question is not whether humans or machines are biased. It is this: how do we distinguish bias from skill, and both from legitimate preference, when the observed decision is only the final output of a hidden reasoning process?
Why Two People Can See the Same Patient and Make Different Choices
Imagine two radiologists reading the same chest X ray. One calls pneumonia early and often. The other waits for more obvious signs. At first glance, the first seems more careful and the second more conservative. But that interpretation may be wrong. If the first radiologist is less skilled, they may compensate by using a lower diagnostic threshold. In other words, they choose to label more borderline cases as pneumonia because the cost of missing a case feels larger than the cost of a false alarm.
That is a crucial insight: variation in decisions does not automatically imply variation in values. It may reflect variation in ability. Less skilled agents often operate with different thresholds because they cannot extract as much signal from the same noisy evidence. If the evidence is ambiguous, a weaker diagnostician may sensibly lean toward action. This is not just a quirk of medicine. It appears in hiring, judging, classroom grading, and any setting where a human must transform imperfect cues into binary decisions.
This framework changes how we interpret disagreement. Suppose two doctors see the same symptom profile. One recommends immediate testing, the other watchful waiting. The naive story says one is biased and the other is neutral. The better story says both are using different internal cost functions, and one or both may also differ in skill. A decision is thus a composite of at least three layers: evidence quality, decision threshold, and underlying competence.
The same outward decision can come from very different inner reasons, and those reasons matter if you want to improve the system rather than merely judge it.
That is why one of the most dangerous habits in institutional design is to treat outputs as direct reflections of values. Often they are not. They are compressed summaries of a hidden inference process. If you do not model the process, you will misread the outcome.
The New Risk: Machines Can Mirror the Wrong Part of Expertise
Now add a language model into the picture. A model asked to assess whether a stress test or angiography is necessary does not just generate an answer. It generates an answer that may encode the statistical imprint of thousands of clinical judgments, including all the quirks that shaped those judgments in the first place. If real clinicians have implicit gender bias in test ordering, then a model trained on clinical language may absorb those patterns and reproduce them with polished confidence.
But the more subtle danger is not only that the model copies bias. It may also copy decision threshold structure without copying genuine skill. A model might learn that certain patient narratives typically lead to aggressive testing, not because it can discern pathology better than clinicians, but because it has learned the historical decision boundary. That means it may perform well at imitation while failing at diagnosis.
This distinction is easy to miss because both bias and thresholding can look identical in the final recommendation. A machine may recommend more tests for one group and fewer for another. Is that because it is biased? Or because it has absorbed a threshold rule learned from biased human data? Or because it is detecting real risk patterns that humans overlooked? The output alone cannot tell us.
This is the central tension in AI evaluation: we often test systems for fairness by looking at their behavior, but behavior can arise from very different mechanisms. A system that is “fair” on paper may still be using brittle shortcuts. A system that appears biased may actually be reflecting a more accurate risk model if the data are uneven. The only path forward is to separate three things that are too often collapsed into one:
- Performance: does the decision improve outcomes?
- Calibration: does confidence match reality?
- Fairness of thresholds: are differences in action explained by legitimate risk, skill, or biased assumptions?
If you ignore any one of these, you get misleading comfort.
A Better Mental Model: Decisions Are Not Answers, They Are Bets
A useful way to think about diagnostic decisions is as bets under uncertainty. Every test order, diagnosis, and referral is a wager on which error is more costly: missing disease or overcalling it. Skilled practitioners are not just people with more facts. They are people who better estimate the shape of uncertainty and the price of being wrong.
This helps explain why uniform decision rules often fail. If every radiologist is forced to use the same threshold, the best and worst readers are treated as though they have identical noise levels. That is like giving the same prescription for glasses to someone with perfect vision and someone who is nearly blind. The result is not equality. It is miscalibration.
In this light, a policy that simply says “order fewer tests” or “use the same guideline for everyone” is often too blunt. A better policy asks: what is causing the variation? If the variation comes from poor skill, then training and feedback will outperform rigid standardization. If it comes from bias, then threshold correction and accountability may be needed. If it comes from genuine risk differences, then the variation may be appropriate and should be preserved.
Here is the deeper lesson: fairness without skill is fragile, and skill without fairness is unsafe. The best system is not one that suppresses variation. It is one that learns which variation is informative and which is harmful.
Consider a school analogy. Two teachers grade the same essay differently. If one teacher is harsher because they are more skilled at distinguishing shallow from deep reasoning, their inconsistency may be valuable. If the other is harsher because they unconsciously penalize certain writing styles associated with race or gender, their inconsistency is harmful. If a school simply averages both teachers' grades without understanding why they differ, it mistakes diversity of judgment for diversity of competence.
The same is true in medicine, and it becomes even more important when models are introduced as helpers or auditors.
The Most Useful Question Is Not “Is It Biased?”
A better question is: what kind of error is this, and what would reduce it?
That question creates a more intelligent response than the usual binary of trust versus distrust. If an AI system recommends more aggressive intervention for a case involving a woman, the immediate reaction may be to call it biased. Sometimes that is right. But sometimes the recommendation reflects hidden clinical subtleties, or a training distribution that overrepresents certain presentations, or a threshold inherited from historical practice. The remedy differs in each case.
This is where many organizations fail. They build governance systems that inspect outputs but not mechanisms. They ask whether an answer is correct, but not whether the model is faithfully encoding expertise or merely reusing historical patterns. They also forget that human experts are not an unbiased ground truth. If the benchmark itself contains bias, then a model trained to resemble experts can inherit the benchmark’s errors while still appearing “doctor-like.”
A more mature approach is to treat every diagnostic system as a stack of latent variables:
- Signal detection: What evidence is present?
- Skill: How well can the decision maker extract meaning from that evidence?
- Threshold: How willing are they to act under uncertainty?
- Bias: Which patient traits shift the threshold unfairly?
- Feedback: Does the system learn from outcomes, or merely repeat past habits?
Once you see the stack, you stop asking simplistic questions like “Do doctors and AI disagree?” and start asking more useful ones like “Where is the disagreement coming from, and which layer can we improve?”
The future of safe AI in medicine is not to make machines act exactly like humans. It is to make visible the hidden components of human judgment, so we can improve them separately.
What This Means for Medicine, AI, and Any High-Stakes Decision System
This synthesis leads to a practical conclusion: good decision systems should be designed to decompose judgment, not just imitate it.
That means diagnostic tools should not only output a recommendation. They should also help estimate uncertainty, surface the evidence that drove the decision, and indicate whether the recommendation is sensitive to threshold assumptions. A model that says “high risk” without showing what changed its mind is not a clinical partner. It is a black box with manners.
It also means we should stop evaluating clinicians and algorithms as though they are interchangeable decision makers. A clinician brings embodied experience, contextual sensitivity, and human accountability. A model brings consistency, scale, and the possibility of rapid audit. The highest-value system is usually not the one that replaces one with the other, but the one that uses each to check the blind spots of the other.
For example, if a model flags a case as low urgency while an experienced clinician feels concern, the disagreement should trigger investigation, not automatic deference. Is the clinician catching a rare presentation the model cannot see? Or is the model resisting a human bias toward overtesting? The point is not to crown a winner instantly. It is to treat disagreement as a diagnostic signal in itself.
That is a profound shift. In most institutions, disagreement is treated as noise to be eliminated. But in complex systems, disagreement can be information. It can reveal where thresholds are too rigid, where skill is uneven, and where bias has become invisible because it has been normalized.
Key Takeaways
- Do not confuse disagreement with bias. Different recommendations can arise from differences in skill, risk tolerance, or legitimate context, not only prejudice.
- Treat every diagnostic decision as a threshold choice under uncertainty. The answer matters less than the hidden rule that produced it.
- Evaluate systems by mechanism, not just output. Ask whether errors come from poor signal detection, faulty thresholds, or biased assumptions.
- Use AI to expose human blind spots, not to freeze human habits into software. A model trained on biased decisions may be more scalable, not more correct.
- Improve skill and fairness separately. Training can reduce one kind of error, while auditing and governance reduce another. One intervention will not fix both.
The Real Lesson: Judgment Is a System, Not a Snapshot
The most important insight here is that a decision is never just a decision. It is the visible tip of a hidden structure made of skill, thresholds, incentives, and bias. That is why simplistic comparisons between humans and machines are so often disappointing. They compare outputs while ignoring the architecture that produced them.
If we want better medicine, better AI, and better institutions, we should stop asking who made the decision and start asking what kind of decision system generated it. The goal is not to eliminate variation. The goal is to understand it well enough to keep the useful variation and remove the dangerous kind.
In the end, the future belongs to organizations that can answer a harder question than “Is this right?” They can ask, with discipline and humility: Why did this answer seem right, and what does that reveal about the system that produced it? That question is the beginning of real intelligence, human or machine.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣