Why the Hardest Systems Are the Ones We Cannot Fully Explain
Hatched by Thomas Hirschmann
May 25, 2026
10 min read
3 views
84%
The Strange Problem of Improving a Mind You Cannot Read
What do a brain stimulation protocol and an AI spam filter have in common? At first glance, almost nothing. One tries to nudge the prefrontal cortex into regulating emotion more effectively. The other tries to help humans evaluate a machine that sorts email. But both reveal the same unsettling truth: the most important systems are often the hardest to judge from the outside.
That difficulty is not an inconvenience. It is the core problem. In both cases, we are dealing with systems whose behavior is real, measurable, and consequential, yet whose internal logic is only partly visible. You can sometimes see the result. You can sometimes measure improvement. But the mechanism, the thing you would most like to understand, remains slippery.
This creates a deeper question that matters far beyond neuroscience or interface design: How do you improve, evaluate, or trust a system when the most valuable effects are partly hidden?
The answer is not to demand perfect transparency. It is to build a better relationship between visible outcomes, hidden mechanisms, and human judgment. That is where these two domains unexpectedly meet.
The Hidden Layer Is Where the Real Work Happens
Emotion regulation is a useful place to start because it exposes a recurring illusion. People often assume that if a technique changes behavior, then it must also change the underlying emotional process in a simple, direct way. But the evidence points to a more complicated picture: interventions can enhance explicit regulation, the kind of deliberate, reportable effort people can describe, without necessarily transforming implicit regulation, the automatic layer beneath awareness.
That distinction matters because it reveals a general principle: the visible layer of performance is not the same thing as the operating layer of the system.
Imagine an email spam filter. A user sees fewer junk messages and concludes the system has become smarter. But what exactly improved? Did the model learn to detect deceptive phrasing? Did it adapt to a new class of sender behavior? Or did it simply get better at a narrow set of patterns while still being vulnerable elsewhere? The surface outcome, fewer spam emails, tells you something important, but not everything.
Now consider the challenge faced by an evaluator using human AI guidelines. Some heuristics are concrete and readily testable. Others are abstract, phase dependent, or hard to apply in a live setting. This is not a flaw in the heuristics. It reflects the deeper truth that evaluation is often an exercise in inference under partial observability.
That is exactly what makes the analogy with emotion regulation so illuminating. In both domains, there is a temptation to confuse what is easy to observe with what is truly transformed. The brain stimulation case warns us that a system can show meaningful gains in one mode of operation while leaving other modes largely untouched. The AI design case warns us that a system can be difficult to evaluate not because it lacks structure, but because the structure is distributed across phases, contexts, and hidden interactions.
When a system changes at one level, it does not follow that every level has changed in the same way.
This is the first lesson: improvement is often partial, asymmetric, and context bound.
Why Evaluation Breaks at the Edges
The hardest part of evaluating complex systems is not the center, where behavior is obvious. It is the edges, where behavior depends on context, timing, and interpretation.
A spam filter is easy to praise when it catches obvious junk. It is much harder to evaluate when it must distinguish between a legitimate newsletter, a suspiciously formatted invoice, a phishing attempt, and a message that is bad for one user but useful for another. Likewise, emotion regulation is easy to applaud when someone can clearly label a feeling or intentionally reframe a thought. It is much harder to know whether deeper automatic responses have changed, especially when the person cannot directly access them.
This is why heuristic evaluation can feel both powerful and frustrating. Heuristics give you a language for inspection, but not all heuristics are equally measurable. Some guidelines can be applied as if they were a checklist. Others require judgment, contextual knowledge, and a tolerance for ambiguity. That difficulty is not a failure of the method. It is a sign that the system’s intelligence is not located in a single visible point.
A useful mental model here is the difference between a thermostat and a weather system. A thermostat is easy to evaluate because its input, logic, and output are tightly coupled. A weather system is harder because local changes, delayed feedback, and nonlinear interactions make prediction difficult. Many human AI systems, and many mental processes, are weather systems disguised as thermostats.
That is why simple evaluation often misleads us. We ask whether the output looks good, but the real question is whether the system has become more robust, adaptable, and aligned across conditions. This requires looking at failure modes, not just average performance.
For example, a spam filter that performs well on ordinary inbox traffic but fails catastrophically on edge cases is not really “good” in a meaningful sense. Likewise, an intervention that helps a person manage stress in a controlled setting but does not generalize to real life may be useful, but its scope is limited. The central issue is not whether change occurred. It is where the change lives.
The Mistake of Asking for One Level of Truth
One of the most common errors in both technology and mental health is the demand for a single, totalizing measure of success. We want one score that tells us whether the system works. But complex systems rarely cooperate with that wish.
A better approach is to think in terms of layers of truth:
- Behavioral truth: What changed in observable output?
- Procedural truth: How was the change produced?
- Transfer truth: Did the change generalize to new situations?
- Structural truth: Did the underlying capacity actually shift, or did the system merely learn a workaround?
In emotion regulation, someone may become better at naming feelings and using deliberate strategies. That is behavioral and procedural truth. But if stress still hijacks them under pressure, transfer truth is weak. If the deeper automatic response remains intact, structural truth is limited.
In AI evaluation, a spam filter may reduce junk mail, which is strong behavioral truth. But if evaluators cannot explain why certain messages are flagged, procedural truth is weak. If the system fails on new scam patterns, transfer truth is weak. If it relies on brittle shortcuts rather than generalizable understanding, structural truth is weak.
The power of this framework is that it prevents false confidence. It tells us that a system can be locally effective and globally fragile. That may sound pessimistic, but it is actually liberating. Once you accept that a single metric cannot capture the whole system, you can design better tests, ask better questions, and stop overcrediting superficial gains.
This is especially important when the system is meant to support human judgment rather than replace it. In such cases, the goal is not full automation or total control. The goal is reliable collaboration across uncertainty.
Designing for What Can Be Seen and What Cannot
The most important design principle that emerges from this synthesis is simple: good systems make the hidden partially legible without pretending it is fully visible.
That means two things at once. First, systems need mechanisms for exposing enough structure that humans can evaluate them intelligently. Second, evaluators need frameworks that respect what remains inaccessible.
A spam filter, for instance, should not merely say “spam” or “not spam.” It should provide cues: suspicious domain patterns, unusual sender behavior, a confidence score, or examples of similar messages. Those signals do not eliminate uncertainty, but they give the human something to reason with. Similarly, if an intervention is intended to change emotional responses, it is useful to distinguish between the person’s conscious strategy use and the automatic effects that may or may not have shifted.
The practical lesson is that explanation should be enough to guide judgment, not so complete that it becomes a fiction. Too little explanation leaves users blind. Too much explanation invites overtrust, as if the system’s internal reality has been fully captured in a neat narrative.
This is why phase based guidelines matter. Systems behave differently at different moments: before use, during interaction, after failure, during adaptation, and in long term maintenance. A system that is easy to assess at deployment may be hard to assess after users have adapted to it. A system that appears helpful in the short term may produce hidden costs later. Likewise, a technique that improves deliberate self control may not alter the emotional reflexes that shape future behavior.
The deeper design challenge is therefore not to make everything transparent. It is to create strategic visibility, enough to support wise decisions when complete visibility is impossible.
The best evaluation does not eliminate uncertainty. It makes uncertainty navigable.
A Better Way to Think About Improvement
Most people think of improvement as a straight line: measure, intervene, improve, repeat. But the two domains here suggest a different model. Improvement is often layered, selective, and negotiated.
First, layered: different aspects of the system respond at different speeds. Second, selective: some capacities change while others stay stable. Third, negotiated: human judgment remains necessary because no single measurement settles the question.
This reframes what success looks like. Success is not the absence of ambiguity. It is the ability to act intelligently within it.
For researchers, this means designing evaluations that distinguish explicit from implicit effects, surface behavior from deeper adaptation, and short term performance from long term resilience. For builders of AI systems, it means treating heuristics not as a substitute for understanding, but as scaffolding for it. For users and decision makers, it means resisting the urge to trust the outcome just because the outcome looks good.
A useful analogy is medical testing. A single blood pressure reading matters, but it does not tell the whole story. Trends, context, symptoms, and risk factors all matter. Complex human AI systems and complex human minds are similar. They require multiple lenses, not one verdict.
That is the real synthesis here. The brain stimulation story shows that a technique can move one kind of regulation without fully touching another. The AI evaluation story shows that heuristics can guide us, but only imperfectly, because systems unfold across contexts and phases. Together they teach a broader principle: the deepest challenge in complex systems is not making them do something, but knowing what, exactly, has changed.
Key Takeaways
-
Do not confuse visible improvement with complete improvement. A system can look better on the surface while leaving deeper mechanisms unchanged.
-
Evaluate across layers. Ask what changed in behavior, how it changed, whether it generalizes, and whether the underlying structure actually shifted.
-
Treat heuristics as guides, not verdicts. Guidelines help humans inspect complex systems, but they rarely settle the full question on their own.
-
Look for edge cases and failure modes. The real test of a system is often how it behaves when conditions are messy, ambiguous, or unfamiliar.
-
Aim for strategic visibility. Good design reveals enough to support judgment without pretending the hidden layer can be fully captured.
Conclusion: The Real Question Is Not Whether It Works, But Where
We often ask whether a system works. That is the wrong first question. The better question is: Where does it work, under what conditions, and at what level of the system?
That shift changes everything. It turns evaluation from a binary verdict into an inquiry about layers, contexts, and mechanisms. It reminds us that the most consequential systems, whether brains or AI tools, are rarely fully legible from the outside. And it suggests a humbler, more powerful standard for progress: not perfect understanding, but disciplined partial understanding that improves our decisions.
In the end, the deepest connection between these seemingly unrelated domains is this: the hardest systems to improve are the ones that force us to become better judges. The work is not only to change the system. It is to learn how to see what has changed, what has not, and why that difference matters.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣