Why Modern Medicine and Modern AI Both Fail by Solving the Wrong Layer First
Hatched by Emil Funk Vangsgaard
May 23, 2026
10 min read
3 views
72%
The common mistake hidden in two very different revolutions
What do a lupus drug trial and a state of the art AI model have in common?
At first glance, almost nothing. One is about a devastating autoimmune disease that attacks the body across multiple organs. The other is about a language model that can build web apps, solve pull requests, and generate animated explanations. Yet both reveal the same uncomfortable truth: progress usually arrives faster than understanding, and the hardest part is not making something work once, but knowing what it should optimize for in the first place.
In lupus, that dilemma appears in the gap between biological insight and clinical reality. B cells are clearly implicated in disease, so anti CD20 drugs make intuitive sense, and yet trials can fail on their primary endpoints while off label use continues because doctors still see benefits in practice. In AI, the same tension shows up when a model suddenly becomes capable of dazzling feats, but those feats are impressive precisely because they expose how many different kinds of intelligence are being bundled together and judged through crude benchmarks.
The deeper question connecting these worlds is this: when a system is too complex to be captured by one clean metric, how do we decide what counts as real improvement?
That question matters far beyond lupus or AI. It sits at the center of medicine, engineering, policy, and even organizational decision making. The easiest thing to measure is often not the thing that matters most. And the thing that matters most is often distributed across many dimensions, making it visible only to practitioners who have learned to distrust simple success signals.
Complex systems punish single metric thinking
Lupus is not one disease in the simple sense that a fever or a fracture is one problem. It is a system wide pattern of misrecognition, where the immune system turns its attention inward and different organs bear the cost differently. Skin, kidneys, joints, blood, fatigue, and inflammation all move on different timelines and respond to different pressures. That is why the disease resists easy treatment logic: there is no single lever that neatly fixes every manifestation.
This is also why clinical trial design becomes almost philosophical. If a drug improves skin lesions dramatically but misses a composite endpoint, is it a failure? If it helps the kidney but not enough patients satisfy a specific response threshold, what exactly has been disproved? Endpoint selection is not a technical footnote. It is the hidden place where a field defines what it values.
AI evaluation has a strikingly similar problem. A model can excel at reasoning, coding, multimodal interaction, and UI generation, yet one benchmark may still fail to capture its real usefulness. A model can look weaker on a narrow test and still be far more capable in practice. In both cases, the score is not the same thing as the system.
This is the first mental model worth keeping: complex systems require portfolio evaluation, not monoculture evaluation. A monoculture asks one question and mistakes the answer for the whole truth. A portfolio asks several questions, then accepts that tradeoffs are real. In lupus, that means distinguishing organ specific benefit from global disease control. In AI, that means distinguishing coding ability, reasoning ability, tool use, and reliability instead of assuming one benchmark can proxy them all.
When a system expresses itself in many dimensions, the right question is not, “Did it win?” It is, “Which kinds of value did it create, for whom, and at what cost?”
That question becomes even more important when safety and cost enter the picture.
Why good mechanisms still fail in the real world
One of the most revealing details in lupus treatment is the persistence of drugs that failed randomized trials. Rituximab, which targets CD20 on B cells, did not meet primary endpoints in key studies. And yet it remains in use because the biological story is compelling, clinicians keep seeing responses, and the alternatives are imperfect. This is not irrationality. It is what happens when evidence, mechanism, and lived patient need do not align neatly.
The market reality intensifies the problem. Biologics are expensive, payers resist broad access, and patients often need to fail standard therapies before receiving newer options. So the treatment landscape becomes a layered negotiation among biology, regulation, reimbursement, and physician judgment. A therapy may be scientifically attractive and clinically useful while still being administratively hard to deploy.
AI has reached a comparable stage of friction. A model may be astonishing in demo form, but the moment it is asked to operate inside workflows, it confronts latency, integration, cost, trust, and domain specific failure modes. A system that can generate a beautiful artifact may still struggle to sustain a robust product. In other words, capability is not deployment.
That distinction matters because both fields are tempted by an illusion: the belief that if we understand the mechanism, outcomes will follow automatically. But real systems are mediated by incentives and constraints. In lupus, B cell targeting makes sense, but the disease remains multi organ and heterogeneous. In AI, scale and reasoning improvements are real, but usefulness depends on whether the model fits human processes.
This is the second mental model: mechanism is necessary, but not sufficient, for adoption. A drug can have a clean biological rationale and still fail because the disease is too broad, the endpoint too blunt, or the cost too high. A model can have dazzling internal capabilities and still fail because the user journey is brittle, the outputs are hard to trust, or the interface is not aligned with a real task.
The most common mistake in both medicine and AI is the same: solving the biological or algorithmic layer first, while leaving the coordination layer untouched.
The real frontier is not better answers, it is better resolution
The most intriguing lupus insight is not just that B cells matter. It is that the field is moving toward more precise ways of reading disease itself. There is interest in biomarkers, urine molecules, and better ways to distinguish lupus nephritis from other manifestations. That desire points toward a deeper shift: from treating a disease as a single label to treating it as a dynamic pattern across tissues, time, and response profiles.
That same shift is happening in AI. The excitement around new models is not only about raw capability, but about the emergence of artifacts, applications, visual explanations, and interactive systems. The model is becoming less like a black box oracle and more like a flexible substrate for many kinds of output. That, too, is a search for resolution: not just “Is it smart?” but “What kind of work can it help us do, and how transparently?”
This suggests a powerful reframing: the frontier is not simply stronger interventions, but finer observation.
In medicine, finer observation means recognizing that patients with mild or moderate disease are not just “less severe” versions of the same case. They may represent a huge, under addressed population whose needs are invisible when trials focus only on the most obvious extremes. In AI, finer observation means recognizing that “reasoning” is not one thing. A model can write code, reason through a diagram, explain a concept, or complete a workflow, and each task stresses a different facet of capability.
Consider an analogy. If you use one blurry camera to study a city, you might see where the skyline is, but you will miss the neighborhoods, traffic patterns, and hidden bottlenecks. You can still produce a map, but it will be a coarse one. Many current systems, in both healthcare and AI, are like that blurry camera. The system works, but the signal is too low resolution to guide the most important decisions.
The practical lesson is that innovation should increasingly focus on instrumentation: better biomarkers, better endpoints, better telemetry, better feedback loops. Not because measurement is glamorous, but because without it, high potential interventions get trapped in ambiguity. The wrong measure can make a useful treatment look marginal. The wrong benchmark can make a transformative model look ordinary.
Progress does not only come from stronger interventions. It comes from sharper ways of seeing the system those interventions act on.
The hidden economics of ambiguity
There is a reason both fields keep running into the same wall. Ambiguity is expensive, and whoever pays for ambiguity gets to shape the future.
In lupus, payers want clear justification for costly biologics, clinicians want flexibility, patients want relief, and regulators want evidence that can generalize. But the disease itself refuses to simplify. This creates a market where access is gated by imperfect proof, and proof is gated by imperfect measurement. The result is not just slow adoption. It is a structural bias toward treatments that are easiest to validate, not necessarily those that best match patient heterogeneity.
AI has its own version of this. Organizations often adopt the tools that are easiest to demo, easiest to integrate, or easiest to budget for, even when those are not the tools with the highest long term value. A model that can do many impressive things may still be hard to justify if its outputs are not measurable inside existing workflows. Ambiguity becomes a tax on usefulness.
This is where the two domains really converge. In both, the winners are not always the most advanced systems. They are often the systems that make uncertainty cheaper. A biologic that can be tracked with a better biomarker becomes easier to prescribe. A model that can show its work, edit artifacts, or operate through structured workflows becomes easier to trust.
That observation leads to a third mental model: the bottleneck is often not competence, but legibility.
A competent system that cannot be read, audited, or placed in context will underperform. In medicine, this means a therapy needs a measurable path from patient profile to expected benefit. In AI, it means the model needs a legible interaction layer so users know when to rely on it and when not to. The future belongs to systems that are not only powerful, but interpretable enough to fit into human institutions.
What to do with this insight
If you zoom out, lupus treatment and frontier AI are both teaching the same discipline: stop asking whether a system is broadly good, and start asking which layer of the problem it truly improves.
Sometimes the answer is biology or model architecture. But often the real leverage is one layer deeper or higher. Better disease stratification. Better endpoint design. Better access pathways. Better evaluation loops. Better interfaces. Better feedback from real use back into development.
This matters because many high stakes fields are still optimized for artifacts of clarity rather than clarity itself. They reward neat trials, elegant demos, and tidy narratives. But the hardest real world problems are messy for structural reasons. They do not become tractable because we wish them to be. They become tractable when we build systems that respect their complexity.
A useful rule of thumb is this: if a problem has many failure modes, the solution should probably have many success signals.
That is the bridge between these two seemingly unrelated stories. In lupus, the ideal next generation therapy is not just “more effective.” It is more selective, more measurable, more tolerable, and more usable within real care pathways. In AI, the ideal next generation model is not just “smarter.” It is more reliable, more legible, more context aware, and more embedded in actual workflows.
The deeper ambition in both fields is the same: move from dramatic promises to structured usefulness.
Key Takeaways
- Do not confuse a strong mechanism with a complete solution. Biological plausibility or model capability is only the first layer of value.
- Use multiple success signals for complex systems. One endpoint or benchmark rarely captures what really matters.
- Measure legibility, not just performance. A therapy or model that can be understood and audited will travel farther than one that only looks good in isolation.
- Focus on resolution. Better biomarkers, better telemetry, and better task decomposition often unlock more progress than another marginal increase in raw power.
- Optimize for deployment, not just demonstration. Real impact happens when a system survives cost, workflow, trust, and human constraints.
Conclusion: the future belongs to systems that can be seen clearly
The most important lesson connecting lupus medicine and modern AI is not about biology or software. It is about epistemology, the art of knowing what is actually happening inside a complex system.
We are entering an era where power is easy to show and hard to evaluate. A drug can change a symptom without solving the disease architecture. A model can amaze without fitting the workflow. In both cases, the temptation is to mistake visible brilliance for durable value.
The better question is more demanding, but more useful: what kind of system becomes easier to understand after it succeeds?
That is the real test. Not whether it dazzles. Not whether it scores well on a narrow metric. But whether, after all the excitement, we can see the problem more clearly than before. Because the deepest innovations do not merely act on the world. They improve our ability to read the world back.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣