When AI Passes the Exam, Who Holds the Bag?
Hatched by SEAN SYLVIA
Apr 18, 2026
11 min read
7 views
88%
The strange new problem with competent machines
What happens when an AI can pass the medical licensing exam, but the clinician beside it may still be the one who gets blamed when something goes wrong?
That question sounds like a legal technicality, but it is really a preview of a much bigger shift. We are moving from a world where machines struggled to meet professional standards to a world where they can sometimes clear the test, yet still lack the kind of understanding, context, and accountability that make judgment meaningful. The result is a dangerous mismatch: capability is becoming machine-like, while responsibility remains human-like.
That mismatch is not just unfair. It may be structurally unstable.
A system in which AI generates recommendations and humans are expected to rubber-stamp, override, or merely “monitor” them creates a peculiar trap. The machine becomes smarter in the narrow sense of producing answers, while the clinician becomes weaker in the practical sense of exercising their best abilities. The person is left holding legal and moral liability for decisions they did not fully shape, cannot fully inspect, and may not even have the time or training to challenge. In other words, AI can turn the clinician into a liability sink.
The deeper issue is not whether AI can score well on an exam. It is whether we are designing systems that confuse passing tests with bearing responsibility.
Passing the test is not the same as earning trust
The fact that a chatbot can pass all three parts of a medical licensing exam is genuinely impressive. It signals that these systems can absorb patterns, synthesize medical knowledge, and produce responses that look plausibly expert. For many people, that is enough to trigger a leap from amazement to alarm: if a machine can pass the test, why not let it decide?
But exams measure only a slice of competence. A medical licensing exam can test recall, reasoning, and pattern recognition. It cannot fully test whether a system understands the full messiness of a patient’s life: family pressures, language barriers, nonadherence rooted in poverty, fear of side effects, religious concerns, or the subtle clues that emerge only in conversation. It also cannot test what a system does when its training data is weak for a particular group, or when it confidently generalizes from patterns that do not apply.
That is why the exam result is both meaningful and misleading. It tells us that AI is becoming good at medical language and probabilistic reasoning. It does not tell us that the system deserves to be treated like a physician, a teammate, or even a reliable subordinate. The distinction matters because medicine is not just an information retrieval problem. It is an obligation to act wisely under uncertainty, in the presence of values, context, and consequences.
A machine can be right often enough to impress us, yet still be the wrong kind of system to trust with a human life.
There is a temptation to believe that competence automatically generates accountability. In human institutions, that is partly true. A doctor who gains expertise also gains authority, and authority comes with responsibility. But AI breaks this moral bargain. The machine can influence the decision deeply while remaining legally and ethically opaque. If something fails, the human at the bedside may inherit the consequences without inheriting real control.
That is the heart of the problem: not intelligence, but asymmetric responsibility.
The liability sink: when oversight becomes a trap
In many current healthcare workflows, AI is positioned as decision support. In theory, the clinician remains in charge. In practice, the clinician is often asked to do an awkward double job. First, they must evaluate the AI output. Second, they must act as a safeguard against its mistakes. That sounds reasonable until you realize what is being asked of them.
The clinician may not know how the model was trained, what biases live in the data, or which patients are systematically underserved by its predictions. They may not have enough time to audit the recommendation properly. They may even become less confident in their own judgment because the machine’s output carries an aura of objectivity. This is the classic automation problem, but in medicine it has a special sting: the human is still the one who signs the chart.
Here is a simple way to see the asymmetry. Imagine a weather app that predicts rain with high confidence, but when it is wrong, you are the one fined for carrying an umbrella. That would be absurd. Yet healthcare can drift into a version of this logic: the AI supplies the prediction, the clinician carries the blame, and the patient experiences the consequences.
The phrase liability sink captures something important because it echoes another familiar image: a heat sink in electronics absorbs unwanted heat to protect the system. But this is no protection. It is a transfer of burden. The clinician absorbs the legal and emotional heat created by the system, even when the system itself is what generated the risk.
This matters because accountability only works when it is coupled to agency. If you want someone to be responsible, they must have genuine power over the outcome. If you want them to supervise, they must have enough visibility to supervise effectively. If you want them to learn from errors, they must be able to understand what happened. Without those conditions, “human oversight” becomes a slogan that masks a liability transfer.
Responsibility without control is not ethics. It is outsourcing risk to the nearest person.
Why AI errors are harder to learn from than human errors
Human medicine is full of mistakes, but human mistakes have one advantage: they can often be explained. A missed diagnosis, a bad handoff, a dosing error, a failure to listen. These can be painful, but they are legible. That legibility allows institutions and individuals to learn. It also allows clinicians involved in an adverse event to make meaning out of the experience, however imperfectly.
AI complicates that learning loop. If a model recommends insulin when it should not, and the result is harm, the natural question is: why did it happen? With a human error, the answer might be found in a rushed note, a misread lab, or a cognitive bias. With AI, the answer may be buried in a model that no one can fully inspect. Was it the training data? The feature weighting? The way the system handled ethnicity, age, or comorbidity? Was the recommendation wrong in general, or only in this patient’s context?
When the reason for the error is unclear, the post-incident process changes. Instead of reflective learning, the likely response is avoidance. Clinicians may decide the safest course is to trust the AI less, use it less, or treat it as a decorative extra rather than a real aid. That is not because they are technophobic. It is because people do not build trust in systems that make them responsible for black box mistakes.
This creates a cruel irony. The more powerful the AI becomes, the more we may need human judgment. But the more the AI shapes decisions in hidden ways, the harder it becomes for humans to learn from failure and preserve confidence in their own expertise. The clinician can become both overburdened and deskilled at the same time.
Think of an air-traffic controller whose screen offers highly confident guidance but no explanation of its logic. If the controller is blamed for a collision, while the system cannot show its work, the controller is not really supervising. They are absorbing risk. Over time, such a role will either become demoralizing or purely ceremonial.
That is why AI in medicine cannot be judged only by accuracy metrics. It must also be judged by how it redistributes learning, authority, and blame.
The real design question: who should own the uncertainty?
The common framing asks whether AI is accurate enough to be trusted. A better question is: who owns the uncertainty when the system is wrong?
This is the missing unit of analysis. Every high-stakes system has uncertainty, but not every system assigns it fairly. In a well-designed human workflow, the same person or institution that creates risk also has enough control to manage it. In a poorly designed AI workflow, uncertainty is created upstream, hidden in the model, and then dumped downstream onto the clinician.
A more robust mental model is to think in terms of four roles:
- Creator: the people who build and train the model.
- Approver: the people who validate it for use.
- Operator: the clinician who uses it in a live case.
- Owner of consequences: the person or institution that carries the harm when it fails.
In too many current systems, these roles collapse onto the operator. That is convenient for procurement and regulation, because it makes deployment easier. It is also ethically suspect. If the AI is responsible for shaping the recommendation, then the creators and approvers must share in the burden of its failures. Otherwise the system socializes the prestige of innovation and privatizes the liability of harm.
This is not an argument against AI. It is an argument for matching responsibility to influence.
If a model nudges a doctor in one direction, the designer of that nudge should not disappear from the accountability chain. If a model works worse on some ethnic backgrounds, that risk should be visible in deployment, not hidden in a footnote no one reads. If a human clinician must override the model, then the interface should be designed to make meaningful override possible, not psychologically exhausting.
The principle is simple: the more the machine shapes the decision, the more the system must disclose, constrain, and account for that influence.
What a healthier AI clinical system would look like
A better system would not ask whether humans or AI should be in charge. That is too crude. It would ask how to preserve the strengths of both without turning the human into a scapegoat.
In practice, that means moving away from a model where AI emits a recommendation and the clinician merely accepts or rejects it. Instead, the AI should be treated as one instrument in a larger diagnostic and conversational process. It should expose uncertainty, show comparative cases where possible, identify known blind spots, and make its confidence and limitations legible to the clinician and patient alike.
For example, imagine a diabetes recommendation tool that does not simply say “start insulin,” but instead says: “Based on coded records, insulin is suggested, but confidence is lower for patients with incomplete medication histories, and performance has been weaker in specific demographic groups. Consider confirming recent weight changes, home glucose data, and patient preferences before acting.” That is not perfect transparency, but it is far better than an unexplained command.
The larger point is that good AI should expand judgment, not replace it with theater. It should help the clinician see more, not merely decide faster. It should support the conversation between clinician and patient, not flatten it into a one-click approval process.
There is also an institutional requirement. Health systems need policies that recognize AI as part of the causal chain. Incident review should examine model behavior, training data, deployment practices, and update cycles, not just the final human action. If an AI contributes to harm, then the investigation should treat that contribution as real and specific, not as a vague background tool no one is accountable for.
That shift would have a cultural benefit too. Clinicians are more likely to adopt AI when they feel it improves their agency rather than threatens their professionalism. People accept tools that make them better. They resist systems that make them responsible for outcomes they cannot control.
Key Takeaways
- Do not confuse exam performance with real-world trustworthiness. A model that passes a licensing exam may still fail at context, equity, and accountability.
- Accountability must track influence. If AI meaningfully shapes a decision, the people who built and approved it should share responsibility for failures.
- Human oversight is not enough if it is poorly supported. Clinicians need visibility into model limits, biases, and uncertainty to supervise effectively.
- Learning from error must remain possible. If AI failures are opaque, clinicians are more likely to avoid the tool instead of improving through it.
- Ask who owns the uncertainty. If the operator carries all the blame while others hold the design power, the system is morally out of balance.
The real test of AI is not whether it can answer, but whether it can be answered for
The most seductive story about medical AI is that competence will solve the problem. Build a model that performs well enough, and the rest will follow. But healthcare has never been only about performance. It is about judgment, responsibility, and the moral structure of care.
A machine that can pass an exam may deserve respect. It does not automatically deserve authority. And a clinician who uses that machine should not automatically become the person society blames when things go wrong. If we are not careful, we will build systems where the smartest part is the one that cannot be held to account, and the accountable part is the one with the least control.
That is the real warning hidden inside this moment. The central question is no longer whether AI can think like a doctor on paper. It is whether we can design institutions where intelligence does not escape responsibility and responsibility does not become a trap.
Because in medicine, the ultimate test is not only whether a system can produce an answer. It is whether, when the answer harms someone, we can still tell who truly owned the decision.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣