Why Better Decisions Need Fewer Experts and More Tests

SEAN SYLVIA

Hatched by SEAN SYLVIA

Jul 22, 2026

10 min read

87%

0

The hidden problem with expertise

What if the real bottleneck in both medicine and artificial intelligence is not intelligence, but confidence in systems we do not sufficiently test?

That question sounds counterintuitive because both domains are surrounded by impressive machinery. Medicine has journals, trials, specialists, guidelines, and a vast share of national income. AI governance has technical labs, policy teams, safety researchers, and increasingly sophisticated alignment methods. Yet in both cases, the decisive question is not whether the system looks expert, but whether it actually improves outcomes under real-world conditions.

That is where the deeper tension emerges. In medicine, we often assume more care should mean more health. In AI, we often assume more technical control should mean better alignment. But the most sobering evidence in both domains points to the same possibility: more intervention can produce more activity without producing more value.

This is not an argument against expertise. It is an argument against confusing expertise with epistemic honesty. The better question is not, “Who is in charge?” It is, “What feedback system can reveal when our confident theories are wrong?”


When more medicine does not mean more health

The medical world offers a brutal lesson in how much our intuitions can fail. A country can spend a fifth of its GDP on medicine, produce mountains of research, and still discover that many of its largest interventions barely move the outcomes people care about. Randomized aggregate studies keep uncovering the same unsettling pattern: more insurance, more visits, more treatment, and then, often, no clear improvement in overall health.

That is not a small technicality. It is a diagnosis of a whole style of reasoning. We look at a sick person getting care, then another, then millions, and we tell ourselves the aggregate must be good because the parts feel obviously beneficial. But an individual encounter is not the same as a population effect. A hospital can be busy and still not be saving lives at the margin. A medicine can be active and still not matter much once you account for selection, confounding, overdiagnosis, placebo effects, defensive medicine, and the fact that many conditions improve on their own.

The deeper lesson is that high-volume intervention can hide low-value impact. When the system is noisy, biased, and full of incentives, the visible outputs become seductive. More scans, more prescriptions, more procedures, more paperwork, more certainty. Yet the real measure is not motion, but delta. Did the patient live longer, suffer less, recover better, or function more fully because of the intervention?

A system can become very good at producing signs of activity while remaining mediocre at producing the thing it claims to optimize.

That is the medical paradox. And it is exactly the kind of paradox that should make us suspicious whenever another domain claims that complexity itself is progress.


AI alignment has the same blind spot

AI governance faces a related temptation. When a technology becomes powerful, the instinct is to believe that only highly technical experts can decide how it should be controlled. But that assumption can be dangerously incomplete, because the core problem is not only technical correctness. It is value legitimacy.

If a system will shape speech, labor, education, persuasion, surveillance, or public knowledge, then the question of alignment is not just, “What can the model do?” It is also, “What should it do, and whose judgments count?” That is where collective intelligence enters. The idea is simple but disruptive: structured public input can be more than a ceremonial checkbox. It can surface information that specialist teams overlook, clarify tradeoffs that engineers flatten, and create accountable decisions that reflect actual human values rather than merely institutional habits.

This matters because a model can be technically aligned to a narrow specification and still be socially misaligned. It can satisfy a policy written by experts while failing ordinary users, minority communities, or future stakeholders who bear the costs. In other words, alignment is not only a calibration problem. It is a governance problem.

The deepest parallel to medicine is that both fields are prone to false confidence from local expertise. Doctors can know a lot about pathology and still miss what improves population health. AI researchers can know a lot about training and inference and still miss what humans collectively want from the system. In both cases, the question is not whether experts matter. They do. The question is whether expertise alone can see the whole picture. Often it cannot.

A model trained to maximize a narrow metric may look brilliant in lab conditions and fail in the wild. A health system optimized for throughput may look efficient and fail in human outcomes. The same structural error appears in both places: we confuse optimization for orientation.


The real unit of intelligence is the feedback loop

The unifying insight is that neither medicine nor AI should be evaluated primarily by its inputs, credentials, or technical sophistication. The true unit of intelligence is the feedback loop.

A feedback loop asks three things:

  1. What did we try?
  2. What changed?
  3. What did we learn that should alter the next decision?

Medicine often fails when it treats treatment as proof rather than hypothesis. AI governance often fails when it treats expert consensus as conclusion rather than experiment. In both cases, the cure is not more faith in authority. It is better measurement, better comparison, and better mechanisms for learning from outcomes.

Think of a thermostat. It is not impressive because it contains a lot of knowledge. It is effective because it continuously checks the room temperature and adjusts itself. Now compare that with a bureaucracy that collects thousands of reports but never updates its assumptions. The bureaucracy may look more sophisticated, but the thermostat is smarter because it is connected to reality.

This is why collective intelligence methods are so compelling. They are not merely democratic rituals. At their best, they function like an upgraded sensor network for institutions. They widen the funnel of evidence, capture lived experience, expose hidden tradeoffs, and make it harder for a small group to mistake its own perspective for the truth.

Likewise, the best medical evidence is not the most dramatic case study or the most elegant theory. It is the evidence that survives contact with randomized trials, long-term follow-up, and aggregate outcomes. That is because medicine, like AI governance, is full of local stories that can mislead. A treatment may help one person, a model may impress one evaluator, a policy may please one committee. The hard part is finding out whether the system works when scaled across real populations.

The most important question in any complex domain is not, “What do smart people believe?” It is, “What does the system learn when reality pushes back?”


From expert control to epistemic humility

This suggests a very different model of progress. Instead of asking institutions to become more certain, we should ask them to become more corrigible.

Corrigibility means being designed to notice error, absorb correction, and change course without collapsing into defensiveness. A corrigible health system is one that can admit a popular intervention does not help much, even if it is profitable, prestigious, or emotionally satisfying. A corrigible AI governance system is one that can admit that a technically elegant alignment scheme is incomplete if it excludes meaningful public input.

This is not anti-expert. It is anti-myth. The myth says that better outcomes come from concentrating authority in the hands of those who know the most. Sometimes that helps. But often the bigger gain comes from building institutions that know when they do not know.

There is a useful distinction here between knowledge work and judgment work. Knowledge work is about facts, methods, and technical constraints. Judgment work is about tradeoffs, legitimacy, values, and priorities. Medicine needs both, because deciding how to allocate finite care is not just a scientific question. AI governance needs both, because deciding what a model may do is not just a coding question. When institutions pretend judgment can be deduced entirely from expertise, they create a vacuum where public trust decays and mistakes compound.

A practical way to see this is to imagine two hospitals. In Hospital A, the staff is brilliant, but outcomes are rarely audited and patients are treated as passive recipients. In Hospital B, the staff is slightly less glamorous, but every major intervention is tracked, patients’ experiences are incorporated, and low-value procedures are aggressively reconsidered. Over time, Hospital B is more likely to improve, not because it knows everything, but because it learns faster.

The same contrast applies to AI. A lab can have world-class researchers and still miss critical societal harms if the only feedback comes from internal testing. A broader decision process, if well designed, can reveal risks, priorities, and failure modes that a closed expert loop cannot see.


What a better institution would look like

If the core problem is weak feedback, the solution is not to worship consensus or to reject expertise. It is to design institutions that combine technical rigor with plural perspective.

That means three things.

First, treat every intervention as a hypothesis. In medicine, this means asking not only whether a treatment is biologically plausible, but whether it improves outcomes at the margin. In AI, it means asking not only whether a safety protocol is elegant, but whether it actually produces better behavior under adversarial and real-world conditions.

Second, separate activity from impact. Busy systems love vanity metrics. Hospital beds filled, model evaluations passed, meetings held, policies drafted. But none of these automatically means lives improved or harms reduced. Better institutions track the thing they claim to care about, even when it is harder to measure and less flattering.

Third, widen the circle of judgment without dissolving standards. Collective intelligence is powerful when it is structured, not sentimental. Public input needs facilitation, informed framing, clear tradeoffs, and mechanisms that convert discussion into decision. Otherwise it becomes a popularity contest. But when done well, it can expose blind spots that technical teams systematically miss.

This is the synthesis hidden in both domains: the best decisions come from systems that are both empirically disciplined and socially accountable. Expertise tells us what is possible. Collective intelligence helps tell us what is worth doing. Randomized evidence tells us what works. Democratic input tells us what counts.

A society that forgets either half will drift into error. A technocracy without feedback becomes arrogant. A populism without evidence becomes reckless. The task is to build institutions that can hold both realities at once.


Key Takeaways

  1. Do not confuse activity with impact. More treatment, more policy effort, or more technical intervention does not necessarily produce better outcomes.

  2. Treat complex systems as learning systems. Ask what feedback loop will reveal whether an intervention actually worked.

  3. Distinguish expertise from legitimacy. Experts know a lot, but they do not automatically know the values, tradeoffs, or social priorities that should govern decisions.

  4. Measure what matters at the aggregate level. Individual success stories can mislead when the real question is population effect or system-wide behavior.

  5. Build institutions that can admit error. The best systems are corrigible, meaning they can update when reality contradicts their assumptions.


Conclusion: the future belongs to institutions that can learn

The deepest connection between medicine and AI is not that both are technical fields. It is that both reveal the limits of expertise when it is cut off from feedback, accountability, and plural judgment. In medicine, we keep discovering that more care does not always mean more health. In AI, we are starting to learn that more sophistication does not always mean more alignment.

That should change how we think about progress. The goal is not to build systems that merely look smart. It is to build systems that stay honest under pressure. The highest form of intelligence is not prediction alone, and not authority alone, but the capacity to revise itself in light of evidence and human values.

The future will not belong to the most confident institutions. It will belong to the ones that can say, with rigor and humility, “We tested it. We listened. We learned. And we changed.”

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣