Why Machines Pass the Test That Matters Least and Fail the Ones That Matter Most

Alessio Frateily

Hatched by Alessio Frateily

May 04, 2026

10 min read

89%

0

The Strange Problem With “Smart” Systems

What if the biggest mistake we make about intelligence is assuming it looks the same everywhere? A system can write a polished essay, ace a hard exam, and hold a conversation that feels uncannily human, yet still stumble on a simple visual logic puzzle. At the same time, a completely different technology, one that seems almost primitive by comparison, can shine light through the forehead and appear to support recovery in disorders as varied as stroke, depression, dementia, and traumatic brain injury.

These two facts point to the same uncomfortable truth: intelligence and healing are both domain specific before they are general. We keep hoping for a single sign that tells us whether a machine is truly smart or a treatment is truly effective. But reality is messier. The tests that are easiest to observe are often the least revealing, while the mechanisms that matter most are hidden in the background, where performance depends on context, embodiment, and the structure of the system itself.

That is the deeper connection between a conversational AI and photobiomodulation. Both force us to stop asking, “Does it work?” in the abstract, and start asking a better question: What kind of work, in what system, under what constraints?

Surface Fluency Is Not the Same as Deep Competence

We are seduced by outputs that look polished. A fluent paragraph feels like understanding. A coherent dialogue feels like reasoning. A confident medical explanation feels like expertise. But fluency can be a costume, and the costume can be remarkably convincing.

The same trap appears in how we evaluate brain interventions. A treatment that sounds biologically plausible and produces encouraging reports can be mistaken for a universal lever, when in fact it may be acting on a narrow set of pathways, cell types, or injury states. Red or near infrared light is not magic. It does not “fix the brain” in a single stroke. Rather, it may alter energy metabolism, reduce inflammation, or support vulnerable tissue in ways that depend on the condition being treated and the timing of the intervention.

This creates a useful mental model: surface competence versus substrate competence.

  • Surface competence is what we can easily observe, such as fluent speech, test scores, or symptom reduction.
  • Substrate competence is what is actually being transformed underneath, such as causal reasoning, neural resilience, or tissue repair.

A chatbot can achieve surface competence in language without necessarily possessing the kind of grounded world model we associate with robust understanding. Likewise, a light based intervention can produce substrate effects without looking dramatic from the outside. The danger in both cases is the same: we overgeneralize from visible success.

The most misleading systems are often the ones that are easiest to admire.

That is why these two topics belong together. They expose a common epistemic error: mistaking performance on a narrow proxy for broad capability or broad cure.

The Real Test Is Not “Can It Perform?” but “What Kind of System Does It Need?”

The old fantasy of intelligence testing assumes there is a clean, portable essence called “smartness.” If you are smart, you should do well on exams, conversation, puzzles, and planning. But real systems are more specific. A human chess master, a bilingual poet, a surgeon, and a philosopher can all be brilliant and still excel in different ways. The same is true of machines: a model can be outstanding at language and poor at visual reasoning because those are different computational terrains.

This is where the visual puzzle example matters. A simple logic grid, one that a child can solve with a pencil, can expose a machine’s weakness even after it has already dazzled us with prose. Why? Because the puzzle demands a tightly constrained relationship between representation, spatial structure, and stepwise consistency. It is not enough to sound right. The system must maintain an internal scaffold that remains stable across transformations.

Photobiomodulation teaches a parallel lesson in biology. The brain is not one homogeneous organ that responds uniformly to any stimulating input. It is a layered, highly conditional system. A weak signal can matter a great deal if it reaches the right place at the right time and alters a fragile process that is already on the edge. Light applied to the head is interesting precisely because it does not fit the simplistic image of a dramatic intervention. It is subtle, targeted, and dependent on context.

This suggests a deeper framework: every powerful intervention is a match between a mechanism and a vulnerability.

In AI, language models exploit the structure of text, syntax, and statistical regularities. They are strong where patterns are rich and repetitive. They are weaker where the task requires grounded spatial reasoning, persistent memory, or physical interaction with the world.

In medicine, red or near infrared light may be helpful where tissue is stressed, energy metabolism is compromised, or inflammation is contributing to decline. It is not an all purpose cure. It is a conditional amplifier.

The lesson is not that these systems are flawed in embarrassing ways. The lesson is that strength is always local before it is global.

The Illusion of a Universal Metric

We love universal metrics because they simplify the world. One score, one test, one verdict. But universal metrics tend to reward whatever is easiest to measure, not whatever is most essential.

For AI, the temptation is to use humanlike conversation as the measure of intelligence. If a system can imitate the style of thinking, we assume it has thinking. Yet the moment we ask it to reason about a diagram, navigate a physical environment, or keep track of multiple interacting constraints, the illusion can break. The conversation was never the whole mind. It was only one interface.

For brain treatment, the temptation is to use symptom change as the entire story. If mood lifts or function improves, we may treat the intervention as broadly restorative. But that can hide a more nuanced reality: an effect may be modest, temporary, highly specific, or only beneficial for certain conditions and protocols. The intervention may be meaningful without being universal.

This is where a more disciplined way of thinking becomes essential. Consider three questions:

  1. What is being measured?
  2. What is being hidden by the measurement?
  3. What kind of system would make this outcome likely?

These questions matter because both intelligence and healing are multi layer phenomena. A fluent answer does not prove deep reasoning. A promising signal does not prove comprehensive recovery. If we insist on universal metrics, we end up rewarding mimicry in one field and overclaim in another.

A better approach is to use diagnostic pluralism. Instead of one master test, use a battery of probes that stress different dimensions of the system. For AI, that means not just language, but reasoning under ambiguity, visual grounding, long horizon planning, and robustness to adversarial shifts. For photobiomodulation, that means not just subjective improvement, but biomarker changes, functional outcomes, timing effects, dose response, and population specific responses.

A system becomes legible only when you test the edges of what it can do.

Light, Language, and the Hidden Architecture of Change

There is a surprising philosophical symmetry here. Language models and light based therapies both challenge our intuition that the most effective forces must be the most obvious ones. We expect intelligence to announce itself through grand reasoning. We expect treatment to work through dramatic intervention. Instead, both domains reveal the power of small signals applied to complex systems.

A large language model does not reason like a human in the way we might intuitively imagine. It predicts, recombines, and generalizes over vast linguistic patterns. Yet out of that machinery emerges surprisingly coherent prose, coding help, and knowledge synthesis. Likewise, light at specific wavelengths does not “heal the brain” in a mystical sense, but it may influence cellular processes that ripple outward into larger patterns of resilience and function.

This is not a coincidence. Complex systems often change not through brute force but through selective perturbation. A small nudge can reorganize a larger structure if the system is already poised near a threshold.

Think of a traffic jam. Adding more cars does nothing. But changing one signal light at the right time can clear a blockage. Or think of a jazz ensemble. A single drummer shifting the timing slightly can transform the feel of the entire piece. In both cases, the leverage comes from understanding the architecture of the system, not from applying more energy.

This is the bridge between AI evaluation and brain intervention. The question is not whether the input is dramatic. The question is whether it is structurally aligned with the process you want to change.

That insight should make us more humble in two directions at once.

First, we should be more humble about machine fluency. Being impressed by text is like being impressed by a mirror that has learned to speak. The reflection can be useful, but it is not the thing itself.

Second, we should be more humble about treatment claims. When a subtle intervention seems to help, that does not mean we have discovered a universal key. It may mean we have found a well placed lever in a very specific machine.

What This Means for How We Build and Judge the Future

If we take this seriously, the practical consequences are large. The next generation of AI evaluation should resemble a clinical workup more than a single exam. A good clinician does not ask one question and declare the patient healthy. They triangulate across symptoms, history, behavior, and targeted tests. Similarly, we should evaluate AI systems across a diverse set of challenges that reveal different failure modes.

The same rigor should apply to interventions like transcranial photobiomodulation. If a method appears promising, we should ask not only whether it works, but for whom, under what conditions, through what mechanism, and with what tradeoffs. That protects us from hype while preserving genuine signal.

Here is the more general principle: do not confuse an interface with an essence.

A chatbot is an interface to language competence, not necessarily to worldly understanding. A light based therapy may be an interface to biological modulation, not a universal cure. In both cases, the interface can be extraordinarily powerful, but only when we respect the limits of what it reveals.

This should change how we build products, design experiments, and interpret success. It should also change how we talk about intelligence and health. The question is never simply whether something is impressive. The question is whether its impressiveness is portable, causal, and robust.

Key Takeaways

  • Stop using one proxy as the whole truth. Fluent conversation, exam performance, or symptom improvement can each hide important weaknesses.
  • Ask what substrate is being changed. Surface outputs matter, but durable progress depends on deeper mechanisms, whether in a model or in the brain.
  • Use multiple probes, not a single score. Robust evaluation should test different dimensions, especially the edges where systems tend to fail.
  • Look for conditional power, not universal claims. The most interesting technologies and therapies often work best in specific contexts rather than everywhere.
  • Be suspicious of impressive interfaces. A beautiful output can reflect genuine competence, but it can also be a convincing mask.

The Deeper Lesson: Intelligence and Healing Are Both About Fit

The most important insight here may be that we have been asking the wrong kind of question about both AI and brain interventions. We ask, “Is it smart?” or “Does it work?” as if there were a single answer that applied across all contexts. But complex systems are not judged that way. They are judged by fit: fit between task and architecture, between perturbation and vulnerability, between signal and state.

That is why a system can pass the Turing test and still fail a visual puzzle. It has achieved fit in one representational domain but not another. That is also why light applied to the head can be scientifically intriguing without being a magic wand. It may fit a specific biological need in a way that is subtle but meaningful.

The world is full of systems that look more general than they are. The real art, whether in technology or medicine, is learning to see the hidden boundaries of their power. Once you do, you stop asking which thing is “really intelligent” or “really therapeutic.” You start asking a better, more adult question: What is this system actually good for, and what does it reveal about the structure of the world?

That shift in question does more than improve evaluation. It changes your sense of what progress means. Progress is not the arrival of one universal test or one universal cure. Progress is learning to recognize the precise place where a small signal can unlock a large change.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣