The Black Box Problem: When Prediction Replaces Judgment in Public Life
Hatched by SEAN SYLVIA
Jul 27, 2026
10 min read
4 views
86%
What happens when the system gets better at predicting and worse at caring?
The most dangerous feature of a black box is not that it is opaque. It is that it can be highly accurate at the wrong job.
That is the unsettling common thread between modern causal inference and the growing use of AI in health insurance: both begin with the same promise of technical sophistication, but they diverge sharply in what they are meant to protect. One is a discipline built to ask, as carefully as possible, “What caused this outcome?” The other is a business practice increasingly asking, “What denial can be justified with the least visible friction?”
The difference sounds subtle until you see the consequences. In one world, a model is useful only if it helps us understand reality, identify treatment effects, and avoid being fooled by selection bias, confounding, or false certainty. In the other, a model can be celebrated precisely because it delivers a clean answer at scale, even if that answer quietly narrows care for the sickest people. The irony is brutal: the same language of rigor, optimization, and data driven decision making can be used either to illuminate causality or to industrialize exclusion.
That is the real tension here. The problem is not AI itself. The problem is who gets to define success.
Prediction is not the same thing as accountability
A predictive model is excellent at finding patterns in past behavior. It can say, with impressive confidence, that a patient with a certain diagnosis, length of stay, or treatment trajectory is likely to need more care. But prediction answers a narrow question: what tends to happen next? It does not answer the more important one: what should happen now, given the patient’s actual medical need and the rules that are supposed to govern access to care?
That gap is where harm enters. When insurers use algorithmic tools to determine whether treatment will be paid for, the model is not merely observing the world. It is actively shaping it. The tool may appear neutral because it is statistical, but the decision it produces is value laden. It determines whose pain is treated as legitimate, whose recovery is considered worth the cost, and which forms of suffering can be quietly delayed behind administrative language.
Think of the difference between a weather forecast and a fire department. A forecast can help you prepare for rain. But if the forecast were used to decide which neighborhoods deserve emergency response, then prediction would have become policy. That is the key failure mode in health coverage: a model built to estimate becomes a proxy for permission.
This is why technical sophistication can become a moral shield. Once a denial is wrapped in algorithmic language, it acquires the aura of inevitability. “The system said no” sounds more objective than “we decided not to pay.” But those are not the same statement. One describes a model output. The other describes a human and institutional choice.
The more automated a denial becomes, the more important it is to ask who wrote the rules, what the model was optimizing, and what it was never allowed to see.
The public often imagines that the danger of AI lies in error, but the deeper danger is alignment. A model can be very accurate at serving the wrong incentive structure. If the incentive is to reduce spending, the model will learn where spending can be cut without immediate reputational cost. If the goal is to protect seniors, the model should be evaluated very differently. The unsettling fact is that the same algorithm can be praised as smart in one setting and condemned as abusive in another, depending entirely on the purpose it serves.
The real issue is not opacity. It is asymmetry.
We often talk about black boxes as if the problem were ignorance alone. But opacity is only part of the story. The deeper issue is asymmetry of power.
A senior facing an appeal does not have the resources to interrogate a proprietary algorithm. A physician trying to get a treatment approved is not negotiating with a peer who shares the same incentives and obligations. A regulator may see the broad structure of the system but not the hidden thresholds, internal criteria, or model updates that determine outcomes in practice. The result is a one way system: the insurer sees everything, the patient sees almost nothing.
That asymmetry matters because due process is not just about the right to object. It is about the ability to understand the basis of a decision well enough to contest it meaningfully. If a denial is generated by rules no outside party can inspect, then appeal becomes theatre. The patient is invited to challenge a decision without access to the logic that produced it.
This is where the lesson from causal inference becomes unexpectedly relevant. Good empirical work is built on skepticism toward hidden structure. Researchers ask whether the estimate is really causal or whether some unobserved force explains the result. They worry about selection bias because they know that apparent patterns can be artifacts of the data generating process. That mindset should not stay inside academia. It should shape public oversight.
In health insurance, the crucial question is not simply whether the model predicts denials efficiently. It is whether the denial process has been designed to respect the underlying legal and clinical framework. If traditional Medicare would cover a service, then a private plan should not be permitted to invent a more restrictive private standard and call that innovation. That is not better science. It is regulatory arbitrage with a dashboard.
This is the central asymmetry: insurers can learn from all the data, but patients must live with the consequences of models they cannot inspect. When one side has computational power and the other has only procedural vulnerability, the system is not optimized. It is tilted.
Why statistical sophistication can make bad incentives harder to see
There is a peculiar temptation in data rich environments: to treat measurement as morality. If a denial process can be quantified, it starts to look disciplined. If an algorithm can be trained, tested, and deployed, it starts to look responsible. If the resulting reductions in spending can be measured, it starts to look like efficiency.
But efficiency is not justice. And measurement is not legitimacy.
Here is a useful mental model: the optimization trap. Whenever an institution can measure one part of a problem more easily than another, it will tend to optimize the measurable part and neglect the rest. In insurance, spending is easy to measure. Emotional distress, delayed recovery, family burden, and lost function are much harder to compress into a dashboard. So the system learns to call the easier thing success.
That is why AI can intensify old abuses rather than replace them. It does not invent the incentive to deny care. It scales it. A human claims reviewer might apply pressure inconsistently. A model can make the same pressure appear systematic, rational, and therefore harder to challenge. The old practice was discretionary. The new one is industrial.
This is also why the appeal of algorithms is so strong to institutions under scrutiny. A proprietary model gives decision makers something that resembles objectivity while preserving discretion. It provides a language of technical necessity. The denial is no longer a choice by a person with a balance sheet. It becomes the output of a system. Yet systems are built, selected, tuned, and governed by people.
The danger is not that machines have replaced judgment. It is that machine language is being used to disguise judgment from accountability.
There is a lesson here for anyone who works with data, not just in health care. A model can be analytically elegant and institutionally corrupt. The more complex the model, the easier it is for leadership to point to technical sophistication as a substitute for ethical scrutiny. That is why any high stakes deployment should be evaluated not only for accuracy, but for the incentive structure it normalizes.
If a model improves cash flow by making justified care harder to obtain, then the model has been perfectly tuned to the wrong target.
The missing question: what is the model allowed to optimize?
The most important question in these debates is rarely asked directly: what should the model be allowed to optimize, and what must remain outside its reach?
That question is more practical than philosophical. Every system has boundaries. The key is whether those boundaries are explicit. In health coverage, the lawful and humane boundary should be simple: if a treatment is covered under established Medicare rules, then proprietary criteria should not override that protection. Any algorithm used in adjudication should be constrained by public rules, not by hidden internal standards.
This is where governance must become more than after the fact review. A yearly committee report is not enough if the underlying architecture still rewards denials. Oversight must be built into the model’s purpose. That means three things:
-
The objective function must be public or at least reviewable. If a model is designed to reduce unnecessary care, who defines unnecessary? If it is designed to predict which claims can be denied without immediate pushback, that should be named as such.
-
The decision rule must be bounded by external standards. A model should not be free to invent stricter criteria than the underlying public program allows.
-
The contestability of the decision must be real, not ceremonial. Patients and clinicians need a path to challenge not just the outcome, but the basis for the outcome.
This is not anti innovation. It is pro legitimacy.
Here causal inference offers a deeper metaphor. The best empirical research does not worship the model. It respects identification. It asks whether the causal claim survives scrutiny from alternative explanations. Public institutions should adopt the same discipline. A denial should be treated like a causal claim: if you say this patient does not need care, show the basis in a way that can be examined, not just asserted.
The more consequential the decision, the more the institution should be forced to move from black box inference to explainable justification. This is especially true when the decision affects older adults, serious illness, and access to treatment that can determine whether someone recovers or declines.
Key Takeaways
- Do not confuse prediction with permission. A model that forecasts likely outcomes should not automatically decide access to care.
- Ask what the system is optimizing. If the hidden goal is cost reduction, then “efficiency” may simply be denial at scale.
- Treat opacity as a power problem, not just a technical problem. Black boxes become dangerous when one side can inspect and tune them while the other side cannot challenge them.
- Demand bounded algorithms. In high stakes settings, AI should operate inside public rules, not above them.
- Make contestability real. A patient appeal process that cannot meaningfully inspect the basis of a denial is not due process, it is paperwork.
The deeper lesson: institutions reveal their values through their models
There is a reason these issues feel larger than health insurance. They point to a broader transformation in modern life: we increasingly ask systems to decide, but we rarely ask what those systems are for. We celebrate quantification because it promises consistency, speed, and scale. Yet when quantification is detached from accountability, it does something subtler and more corrosive. It turns values into procedures and then hides the values inside the procedures.
That is why the clash between causal rigor and algorithmic denial matters so much. One tradition tries to learn how the world works without fooling itself. The other too often uses technical form to obscure institutional intent. Both can speak the language of science. Only one is trying to reduce self deception.
The real battle is not between humans and machines. It is between public standards and private optimization. If we allow high stakes models to define their own success, they will always find ways to look efficient while externalizing pain. But if we insist that models remain answerable to law, medicine, and human judgment, then data can serve care rather than consume it.
In the end, the question is not whether algorithms will shape our institutions. They already do. The question is whether we will let them become substitutes for responsibility, or whether we will force them back into their proper role: tools that inform judgment, not machines that evade it.
That is the lesson hidden inside the black box. The most important thing it conceals is not the code. It is the choice.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣