The AI Progress Trap: Why We Need Less Faith in Magic and More Diversity in Ideas

Thomas Hirschmann

Hatched by Thomas Hirschmann

Aug 15, 2026

11 min read

92%

0

What if the greatest obstacle to better artificial intelligence is not a lack of computing power, data, or public enthusiasm, but our tendency to mistake familiarity for understanding?

The current AI moment is shaped by two seemingly opposite facts. The public is often most receptive to AI when it understands the technology least. At the same time, the industry is investing most heavily in a single family of systems, even as each new increase in scale produces less dramatic improvement than the last.

These facts belong to the same story. In both cases, people are responding to an illusion of inevitability. Users see a system perform a surprising task and imagine an almost magical general intelligence behind it. Researchers see a successful architecture improve with more resources and imagine that continuing along the same path will eventually resolve every limitation.

The deeper problem is not that people believe too much in AI. It is that wonder and convergence are being confused with understanding and progress.

The double illusion: magic for users, inevitability for builders

Consider a person who asks a language model to write a sonnet, explain a legal concept, or generate a plausible business strategy. The system produces an answer in seconds, using abilities that appear to involve reasoning, creativity, or expertise. If the user has little knowledge of how such systems work, the result can feel magical. The less familiar the mechanism, the more the performance seems to reveal an inner intelligence.

That reaction is not irrational. A child does not need to understand a combustion engine to be astonished by an airplane. Human beings routinely judge capabilities by observing outcomes, not by reconstructing the machinery that produced them. When the outcome touches a domain we associate with uniquely human abilities, the astonishment becomes even stronger.

But the same psychological shortcut appears inside the technology industry. A model that becomes more capable after receiving vastly more parameters, data, and computation can create the impression that scale itself is the hidden principle of intelligence. The system improves, so the strategy appears validated. Investment follows improvement, and further investment produces more improvement. This feedback loop can make a particular approach look less like one promising path and more like the path.

Yet a larger engine is not necessarily a new kind of vehicle. If a model grows dramatically while its practical capabilities improve only modestly, the gap between more and better becomes impossible to ignore. Increasing the size of a system may yield useful gains, but it does not guarantee a conceptual breakthrough. A hundred times the machinery does not imply a hundred times the intelligence.

This is where the two illusions meet. Public awe can encourage overconfidence in what current systems are. Industrial success can encourage overconfidence in what current methods will become. In each case, a visible performance is mistaken for evidence of a complete theory.

The public sees a miracle and assumes there is a mind behind it. The industry sees an improvement and assumes there is a destination ahead.

Both assumptions can be wrong.

Why monocultures become dangerous precisely when they succeed

A monoculture is not simply a situation in which one technology is popular. It is a situation in which one technology becomes the default lens for defining the problem itself.

When nearly every major question is phrased as, “How can a language model do this better?” other questions disappear. Could a symbolic system handle the task more reliably? Would a planning architecture help? Is a causal model needed? Could a system combining perception, memory, simulation, and language outperform a larger text predictor? Does the task even require a generative model?

The danger is not merely financial. It is intellectual. A dominant method attracts talent, benchmarks, infrastructure, and prestige. As these resources accumulate, alternatives begin to look less promising, not necessarily because they are weaker, but because they lack the surrounding ecosystem that makes progress visible.

This is how a technological monoculture can conceal its own diminishing returns. The field keeps producing impressive demonstrations, but the demonstrations are increasingly optimized for the strengths of the dominant method. The system becomes excellent at passing tests designed around its existing capabilities, while deeper limitations remain untreated.

Imagine trying to improve transportation by making one kind of horse larger and faster. For a while, the gains would be real. Better breeding, better nutrition, and better equipment would produce impressive results. The success could attract more investment and persuade observers that the future of transportation had been identified. But no amount of optimization would produce the capabilities of an automobile, a railway, or an aircraft.

Scientific progress often depends on such discontinuities. Breakthroughs do not always arrive as smoother versions of the current paradigm. They may emerge from a different representation of the problem, a different architecture, or an entirely different question. The history of computing is full of moments when progress came not from extending the dominant machine indefinitely, but from changing what counted as a machine and what it was for.

This makes diversity more than a portfolio strategy. Diversity is a search strategy for escaping the limits of the current map.

A research ecosystem containing language models, embodied agents, causal systems, symbolic reasoners, evolutionary methods, neuromorphic hardware, and hybrid architectures is not merely hedging against failure. It is increasing the probability that someone will discover a better coordinate system for intelligence.

The important distinction is between incremental improvement and paradigm discovery. Incremental improvement asks how to make the current system more efficient or capable. Paradigm discovery asks whether the current system has defined intelligence too narrowly.

The paradox of literacy: knowledge can reduce excitement while improving judgment

Public understanding introduces a related paradox. People with lower AI literacy tend to be more receptive to AI, partly because they are more likely to perceive its capabilities as magical. Demystification can reduce the emotional force of a demonstration. Once someone learns that a model generates outputs by predicting patterns from vast training data, the result may seem less like thought and more like sophisticated machinery.

That reduction in awe is not necessarily a social problem. It may be a cognitive improvement.

A magician who explains a trick usually makes the trick less mysterious, but the audience may leave with a better appreciation of the actual skill involved. The disappearance of magic can make room for a more precise form of admiration. We can be impressed by an illusion without believing the performer possesses supernatural powers.

The challenge is that institutions often want both responsible understanding and rapid adoption. They may tell people that AI is merely a tool when discussing safety, then present it as almost magical when encouraging use. The result is a confused public posture: skepticism in theory, enchantment in practice.

This confusion matters because people do not evaluate every AI application equally. They are more likely to overtrust a system when it performs a task associated with human identity, such as composing, advising, diagnosing, or reasoning. A fluent answer can feel like evidence of comprehension even when the system lacks stable knowledge, reliable goals, or a grounded model of the world.

The answer is not to eliminate wonder. Wonder is a powerful engine of curiosity, experimentation, and adoption. The answer is to separate wonder as motivation from wonder as evidence.

A person may be amazed that a system can summarize a complicated document. That amazement can motivate exploration. But it should not by itself justify using the system to make a medical decision, evaluate a job candidate, or determine whether a legal argument is sound.

This distinction can be expressed as a simple rule:

Let magic open the door, but never let it certify what lies behind the door.

The same rule applies to research. A spectacular demo should attract attention, not settle the question of whether the underlying method is sufficient.

A better model: cultivate two kinds of diversity

The future of AI requires diversity in two different domains, and confusing them creates new problems.

The first is architectural diversity. This means pursuing multiple technical approaches rather than assuming that one family of models will absorb every important capability. It includes investment in systems that reason through explicit structures, learn from interaction with environments, maintain durable memory, represent causality, plan across long time horizons, and use language as one component rather than the entire substrate of intelligence.

The second is epistemic diversity. This means preserving multiple ways of interpreting what an AI system is doing. Engineers may analyze probabilities and optimization. Users may experience fluency and wonder. Social scientists may study trust and adoption. Domain experts may focus on failure modes that benchmarks miss. Regulators may ask who bears responsibility when the system is wrong.

These perspectives should not be collapsed into one official narrative. Each reveals something the others can miss.

A purely technical view may underestimate how strongly presentation shapes trust. A purely psychological view may treat the system as a social object while ignoring its architecture. A purely commercial view may measure adoption but overlook whether users are becoming dependent on unreliable outputs. A purely skeptical view may dismiss genuine utility because the system does not think in a human way.

The goal is not to choose between enchantment and disenchantment. It is to create a layered understanding in which different descriptions operate at different levels.

For example, a language model can be described as:

  • A statistical system that predicts sequences from learned patterns.
  • A useful interface for drafting, translation, and exploration.
  • A social technology that changes how people distribute authority and expertise.
  • A component that may eventually cooperate with memory, planning, tools, and perception.

None of these descriptions alone is complete. The mistake is treating one description as the entire reality.

This layered model also changes how organizations should evaluate AI. Instead of asking only whether a system performs impressively, they should ask four separate questions:

  1. Capability: What can the system do under favorable conditions?
  2. Reliability: How does it behave when the prompt, data, or environment changes?
  3. Interpretability: Do users understand the basis and limits of its output?
  4. Substitutability: What other technical approaches could perform the same task, perhaps more safely or efficiently?

The fourth question is especially neglected. If an organization assumes that every task must be solved with the dominant architecture, it may never discover a simpler or more dependable alternative.

What to do when the future is not yet decided

For individuals, the practical lesson is to become neither an evangelist nor a reflexive skeptic. Use AI enthusiastically for low consequence experimentation, but increase scrutiny as the cost of error rises. Treat fluency as a user interface feature, not as proof of understanding.

For companies, avoid building an AI strategy around a single vendor, model family, or benchmark. Establish small comparison projects using different methods. Test whether the system improves outcomes rather than merely producing outputs that appear sophisticated. Make room for teams whose job is not to deploy the leading approach, but to challenge its assumptions.

For educators, AI literacy should include both mechanism and psychology. Students need to know how models generate responses, but they also need to understand why fluent performance triggers trust. The critical skill is not memorizing that AI can be wrong. It is learning to recognize the situations in which confidence, speed, and humanlike language are especially misleading.

For researchers and funders, the central question should be whether a project expands the space of possible architectures and representations. A less fashionable approach may be more valuable if it attacks a limitation that scaling cannot solve. The most important result may not be a better score on an existing benchmark, but evidence that the benchmark itself is measuring the wrong thing.

Key Takeaways

  • Separate awe from evidence. A surprising output is a reason to investigate a system, not a reason to trust it.
  • Diversify technical bets. Compare language models with causal, symbolic, embodied, planning, memory based, and hybrid systems before assuming one architecture is universally appropriate.
  • Measure diminishing returns honestly. Track practical capability, reliability, and cost, not just model size or benchmark gains.
  • Teach the psychology of AI trust. Understanding why fluency feels like intelligence is as important as understanding how prediction works.
  • Preserve productive mystery without surrendering judgment. Curiosity accelerates adoption and discovery, but decisions should rest on tested performance and known limitations.

The deepest lesson is about how progress works. Progress does not require the disappearance of uncertainty. In fact, uncertainty can be productive when it keeps multiple possibilities alive. The danger begins when mystery is converted into certainty, whether the certainty belongs to a user who sees a magical mind or to an industry that sees an inevitable scaling curve.

The future of intelligence will not be determined by how convincingly machines imitate the appearance of thought. It will be determined by whether we remain capable of asking what kinds of thinking our current systems cannot represent.

That question demands two forms of humility. Users must remember that an impressive performance may conceal profound limits. Builders must remember that a successful method may still be only one chapter in the history of intelligence.

The most advanced AI ecosystem, then, will not be the one with the single largest model or the most enchanted audience. It will be the one that can sustain wonder while resisting worship, pursue scale while funding alternatives, and recognize that the next breakthrough may come from an idea that today looks less magical because it has not yet been imagined.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣