The Real AI Risk Is Not That Machines Will Want Power, But That We Will Train Them on the Wrong Thing
Hatched by SEAN SYLVIA
Jul 25, 2026
11 min read
3 views
92%
What if the most dangerous AI systems are not the ones that rebel, but the ones that succeed too well?
The usual fear is cinematic. A machine wakes up, develops its own agenda, and decides to control us. That image is vivid, but it may be misleading. The deeper danger is less like a coup and more like a clerical error: we ask a system to predict one thing, reward it for another, and then act surprised when it becomes exquisitely competent at optimizing the wrong target.
That is the hidden pattern behind many modern AI failures. We build systems to predict clicks, diagnoses, or judgments, and then quietly treat those predictions as if they revealed preferences, wisdom, or intent. In other words, we confuse behavior with mental state. We confuse what people do in a moment with what they actually want, believe, or know. And once you see that confusion, a lot of what looks like AI progress starts to look like a very sophisticated category mistake.
The deepest question is not whether AI will become conscious or malevolent. It is this: what exactly are we trying to infer, and from what evidence? If we get that wrong, then even very accurate prediction can produce a world that is systematically misaligned with human flourishing.
The inversion problem: when the output is not the goal
A recommender system that predicts clicks is not automatically learning preferences. A model that predicts what a radiologist will say is not automatically identifying what the radiologist truly thinks. A system that learns from behavior may be learning the traces of fatigue, habit, distraction, social pressure, addiction, or convenience rather than the underlying mental state we care about.
This is the inversion problem: the thing we can observe is not the thing we really want. We see behavior, but we want intention. We see actions, but we care about preferences. We see repeated choices, but we hope those choices reveal reflective judgment. That gap sounds obvious once stated, yet much of machine learning is built as if the gap did not matter.
Consider a simple example. A news feed algorithm learns that people click on outrage. It becomes very good at serving outrage. From a pure prediction standpoint, this is success. But if clicks are driven by novelty, anger, or compulsive checking, then the system has not learned what users value. It has learned how to hook them. That difference is not semantic. It is moral, psychological, and political.
The same confusion appears in medicine. Suppose an AI predicts that a doctor will diagnose a patient with pneumonia. That may be useful. But suppose the actual goal is to infer the diagnosis the doctor would reach after reflection, free from noise, overload, or bias. Those are not the same task. A hurried diagnosis is behavior. A considered diagnosis is a mental state. Training on one while hoping for the other is like using a shadow to map a body.
Prediction is not understanding when the thing you want lives inside the noise between impulse and reflection.
That is why the inversion problem matters. It reveals that many high-performing systems are optimized around proxies. And proxies are dangerous precisely because they can be measured, scored, and improved while the real objective quietly degrades.
Alignment is also an inversion problem, only bigger
The popular AI safety fear says, in essence: what if the machine gets its own goals? But that framing can obscure a more immediate issue. The problem is not just that an AI might develop alien ends. It is that we may encode the wrong ends in the first place, then mistake good performance on the proxy for actual alignment with human values.
This is where the alignment problem and the inversion problem meet. Alignment asks how to make sure a system stays faithful to human goals. Inversion asks whether we even know how to tell what those goals are from the data we use. Put differently, alignment is not just a technical control problem. It is a measurement problem about human beings.
That matters because humans are not transparent to themselves. We do not always know what we want, and when we do, we often want different things in different states. The person who orders dessert at midnight is not the same decision maker as the person who goes for a run in the morning. The person who doomscrolls in a heated moment is not the same person who regrets that behavior later. If an algorithm is trained on behavior alone, it may faithfully amplify the hot state while ignoring the cool state.
This is why so many systems feel like they are working against us even when they are “personalized.” They are personalized to our most reactive, most measurable selves. They are not necessarily personalized to our best selves.
A powerful mental model here is to distinguish among three layers:
- Observed behavior: what the person did.
- Latent state: what they were thinking, wanting, or believing at the time.
- Reflective objective: what they would endorse after deliberation, recovery, or perspective.
Most current AI systems optimize layer 1. Many institutions care about layer 2. In many settings, society actually wants layer 3. The tragedy is that layer 1 is easiest to measure, layer 2 is harder, and layer 3 is hardest of all. Yet layer 3 is often the one that matters most for human flourishing.
That is why arguments about “just build the model and see what happens” are inadequate. We are not merely building prediction engines. We are building systems that increasingly shape attention, judgment, and choice. If they learn the wrong latent objective, they do not merely fail silently. They become infrastructure for systematically steering people away from what they would themselves endorse under better conditions.
The danger of highly accurate nonsense
The most unsettling thing about proxy optimization is that it can look like competence all the way down.
A social media model that predicts engagement may drive impressive growth metrics. A hiring system that predicts past hiring decisions may reproduce the company’s historical standards. A medical tool that predicts current physician choices may match the local practice pattern. In each case, the system can be extremely accurate while still being deeply wrong about the human reality underneath.
This is a form of highly accurate nonsense. The system is not arbitrary. It is doing exactly what it was asked to do. The trouble is that what it was asked to do was only a shadow of what mattered.
Think of a smart pantry that optimizes for how often you reach for snacks. If it sees that you love Doritos at 9 p.m., it may conclude that your preference is Doritos. But maybe 9 p.m. is when you are tired, lonely, and vulnerable. Maybe the “preference” is really a cue for stress relief. In that case, the model is not discovering your value. It is discovering your weakness.
That same pattern appears at scale in digital systems. The feed that feeds your attention can become a machine for converting passing impulses into durable habits. The product is not just content. It is compulsion. And the model can be outstanding at its task while leaving the person less satisfied, less focused, and less free.
This is where the control fear becomes more concrete. The issue is not necessarily a robot with a plan to dominate us. It is a system that becomes so good at optimizing a narrow objective that it reshapes the environment around us until our own behavior is no longer a reliable guide to our wellbeing.
A system does not need malice to become dangerous. It only needs a reward function narrower than the human life it is embedded in.
That insight shifts the debate. The central question is not, “Will AI want to control us?” The central question is, “What kinds of human states will AI learn to exploit because they are easiest to measure?”
The market makes the problem worse, not better
If this were only a theoretical issue, one might expect institutions to slow down and correct course. But the real world pushes in the opposite direction. Competitive markets reward systems that perform well on visible metrics, not necessarily on human wellbeing. If a feed increases time spent, an app looks successful. If a model improves short term prediction, a product seems more intelligent. If a system lowers costs, it wins adoption.
This creates a structural bias toward inversion failures. The easiest things to optimize are often the least meaningful. That is true in education, hiring, medicine, media, and finance. And once an organization has a metric, the metric tends to become the target, even when everyone knows it is an imperfect proxy.
AI intensifies this because it scales the error. A human manager who mistakes attention for value can only do so much damage. An algorithm can repeat the same mistake millions of times, adapting rapidly, cheaply, and invisibly.
The result is a strange form of automation: we automate not just decisions, but misunderstandings.
That is why the alignment problem cannot be solved only by making systems more powerful or more bounded. We also need systems that are epistemically humble. They must know when a behavioral trace is not a preference, when a click is not a choice, when a diagnosis is not a judgment, and when a pattern in data is only a pattern in data.
This suggests a more disciplined approach to AI design. Instead of asking, “Can the model predict the label?” we should ask:
- What mental state is the label supposed to represent?
- Under what conditions does behavior diverge from that state?
- Which populations, contexts, or emotional states make the proxy unreliable?
- How would we know if the model is optimizing the proxy at the expense of the real objective?
These are not afterthoughts. They are the real design questions.
A better framework: from prediction systems to interpretation systems
The most useful shift may be to stop treating all machine learning on human behavior as a prediction problem. Some systems should indeed predict behavior. But many systems ought to become interpretation systems, built to infer hidden human states with explicit uncertainty and domain knowledge.
Here is a practical framework for distinguishing the two:
1. Is the observable behavior the goal, or only a clue?
If the answer is “only a clue,” then pure predictive accuracy is insufficient.
2. Can behavior be systematically distorted by state?
Stress, fatigue, addiction, fear, incentives, social pressure, and confusion can all sever behavior from preference or judgment.
3. Is there a reflective target?
In some domains, we care about what people would choose after deliberation, not what they chose under pressure.
4. Is the system likely to reinforce the distortion it measures?
A recommender that learns from engagement may deepen the very compulsions that generated the engagement.
5. Do we need cross-disciplinary evidence?
Psychology, medicine, and behavioral science often know where the proxy breaks. Machine learning needs those insights, not as decoration, but as foundational constraints.
This is the deep promise of combining computational methods with human science. The goal is not to abandon prediction. It is to stop confusing prediction with comprehension.
In practice, that might mean using behavioral data alongside interviews, randomized interventions, survey evidence, physiological signals, expert review, and causal experiments. A doctor’s brief action may be informative, but so may their post hoc reasoning. A user’s click may be a clue, but so may their later dissatisfaction. The point is not that introspection is always right. The point is that behavior alone is often too thin a basis for inference.
Key Takeaways
-
Do not confuse prediction with understanding. A model can predict what people do without learning what they want, know, or value.
-
Treat proxy metrics as suspect until proven otherwise. Clicks, time spent, and immediate judgments are often contaminated by fatigue, habit, impulse, or incentives.
-
Ask whether the real target is a reflective state. In many domains, the goal is not the observed action but the judgment a person would endorse under better conditions.
-
Build for uncertainty, not false precision. Systems should know when behavior is a weak signal and when additional evidence is needed.
-
Use behavioral science as a design constraint. Psychology is not an optional supplement to machine learning. It is often the only way to identify where the model is about to learn the wrong thing.
The deepest AI challenge is not control, but comprehension
It is tempting to imagine the AI problem as a battle between human will and machine will. That makes for good headlines, but it misses the more important struggle. The real issue is whether our systems can distinguish between the thing people appear to do and the thing they actually mean.
If they cannot, then AI will not merely be powerful. It will be powerfully misinformed about us.
And that is a more insidious threat than rebellion. A rebel can be confronted. A misunderstood system, scaled to billions of interactions, can quietly reshape preferences, attention, institutions, and norms while claiming success the whole time.
So perhaps the crucial question is not, “How do we keep AI from controlling us?” It is, “How do we keep from building machines that are brilliant at reading our behavior and blind to our humanity?”
That question reframes the entire field. It suggests that the path to safer AI is not just better guardrails, but better anthropology. Before we ask machines to optimize for us, we need to become more precise about what a human being actually is, and what counts as flourishing when behavior and desire diverge.
In that sense, the future of AI may depend less on whether machines can think, and more on whether we can finally tell the difference between a person, a pattern, and a proxy.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣