When Interfaces Lie: Why Good AI Design Starts with Words, Not Scores
Hatched by Thomas Hirschmann
Jun 05, 2026
9 min read
4 views
87%
The Hidden Failure in AI Design
What if the biggest risk in building AI systems is not that they are too smart, but that we evaluate them too mechanically? It is tempting to believe that if we can define enough requirements, assign enough points, and compare enough concepts, we will arrive at the best system. But AI is not just a technical artifact that can be ranked like office furniture. It is a language machine, a meaning machine, and a social machine. That means the way it speaks, frames, and responds is not decorative. It is the product.
This creates a dangerous illusion: a polished evaluation score can make a concept look reliable even when its language model is shaping user behavior in ways the scoring rubric never anticipated. A system that seems optimal on paper can still fail in practice if its words mislead, overpromise, flatten nuance, or create false confidence. In AI, the interface is not a wrapper around intelligence. It is where intelligence becomes legible, usable, and sometimes dangerous.
The deeper question is not simply which concept scores highest. It is: What are we actually measuring when we evaluate an AI system, and what are we missing when we reduce language to a checklist?
The Seduction of the Score
Scoring systems are powerful because they promise fairness, comparability, and clarity. If you have a set of requirements, you can assess each concept against the same criteria, assign values, and choose a preferred option with confidence. This works well in domains where the object being evaluated is stable, measurable, and relatively independent of context. A bridge either meets load requirements or it does not. A motor either meets efficiency targets or it does not.
AI systems are different. They do not merely perform tasks. They interpret prompts, generate language, and influence human judgment through phrasing, tone, and framing. A system can satisfy a requirement on a spreadsheet while still creating confusion in the hands of real users. For example, a customer support chatbot may score well for response speed and completeness, yet still frustrate users if it speaks with false certainty, avoids admitting uncertainty, or uses jargon that sounds authoritative but is semantically vague.
The problem is that scores often capture the visible shell of performance rather than the lived experience of interaction. A model can be rated highly for correctness in test cases and still fail when a user asks an ambiguous question, because the system's language makes a guess sound like a fact. In other words, evaluation can reward the appearance of competence instead of the practice of communication.
This is where language enters as a design issue, not a cosmetic one. If words shape how users understand an AI system, then words also shape what the system is allowed to become in practice.
A score can tell you whether a concept satisfies requirements. It cannot tell you whether the system tells the truth in a way people can actually use.
Why Words Are Not Just Output
It is easy to treat language as the final layer of an AI system, the place where the machine finishes its work by putting results into sentences. But language is not the end of the process. It is the mechanism through which the system defines reality for the user. Every choice of wording sets a boundary around what the system claims to know, what it invites the user to do, and what kinds of errors become likely.
Consider two systems that both provide medical triage guidance. One says, “You may want to seek care if your symptoms worsen.” The other says, “You are not in danger.” The first phrase is cautious and descriptive. The second is definitive and potentially irresponsible. Both may be generated from the same underlying confidence estimate, yet the second turns uncertainty into authority. The linguistic difference is not minor. It changes the moral profile of the system.
This is why linguistics matters so much in automated systems. Words carry assumptions, not just information. Terms like “risk,” “confidence,” “recommendation,” and “explanation” are not neutral labels. They encode relationships between system and user. They tell users whether to trust, whether to verify, whether to act, and whether the system is speaking probabilistically or absolutely.
A language-aware design process would therefore ask questions that scoring rubrics often ignore:
- Does the system speak in a way that makes uncertainty visible?
- Does it distinguish between inference and fact?
- Does it use the user's language or force them into the system's vocabulary?
- Does it sound confident because it is correct, or because it is optimized to sound persuasive?
These are not stylistic questions. They are questions about how automation can quietly govern human understanding.
The Real Tension: Measuring Capability Versus Shaping Meaning
The central tension is this: evaluation frameworks want objectivity, but language-based systems live in interpretation. Requirements specifications are essential because they force designers to be explicit about what matters. Yet explicit criteria can also create a false sense of completeness. They can make it seem as though every meaningful property of an AI system can be anticipated in advance and scored in isolation.
That is rarely true.
An AI system does not just fulfill a task. It enters a relationship with a user. In that relationship, meanings shift depending on context, trust, stakes, and prior experience. A simple example is autocomplete. Technically, it can be scored on relevance, latency, and accuracy. But a user experiences autocomplete as an invitation, a suggestion, or sometimes an intrusion. The same output can feel helpful in one context and manipulative in another. A score may identify technical superiority while missing social friction entirely.
This is where a more sophisticated design mindset becomes necessary. Instead of asking only, “Which concept is best according to our criteria?” we should also ask, “Which concept creates the healthiest interpretive environment for users?” That phrase matters because AI systems do not merely solve problems. They create environments in which problems are understood.
Think of a thermostat. If it is poorly designed, it gives inaccurate temperature readings and wastes energy. But an AI writing assistant is more like a co-author with a voice. It can nudge style, censor ambiguity, embellish confidence, or standardize expression. The failure mode is not just inaccuracy. It is the subtle reshaping of how people think, write, and decide.
That is why a purely numerical approach can be incomplete. Not because numbers are bad, but because numbers can only score what the team has already named. Linguistics asks: what unnamed assumptions are hiding inside those names?
A Better Model: Evaluate Systems on Semantics, Not Just Metrics
If AI design is partly a problem of meaning, then evaluation needs a semantic layer. This does not mean abandoning scoring systems. It means widening them. A useful mental model is to think of AI evaluation in three layers:
- Functional correctness: Does the system do the task?
- Interpretive clarity: Does the user understand what the system means, knows, and does not know?
- Interactional safety: Does the language of the system steer behavior in ways that are responsible, transparent, and contextually appropriate?
Most teams overinvest in the first layer and underinvest in the second and third. That imbalance is why many AI products feel impressive in demos and brittle in reality. They are functionally competent but semantically underdesigned.
For example, a hiring screening tool may rank resumes efficiently. But if it explains its recommendations in templated language that sounds objective without being interpretable, it may encourage overtrust. Users may infer a level of fairness or precision that the system does not actually possess. The output is not just a score. It is an authority claim.
A semantic evaluation process would test not only whether outputs are correct, but whether the system's phrasing makes its limitations visible. It would ask whether the model uses hedging appropriately, whether it distinguishes evidence from inference, whether it adapts to user expertise, and whether it avoids linguistic patterns that impersonate certainty. In that sense, evaluation becomes a discipline of truthfulness.
The best AI system is not the one that sounds most decisive. It is the one whose language best matches the real shape of its knowledge.
Designing for Honest Automation
The phrase “automate the use of AI” sounds efficient, but it should also sound risky. If automation is layered on top of systems that already produce ambiguous or overconfident language, we can end up scaling confusion at machine speed. The goal should not be to automate more speech. It should be to automate more honest speech.
Honest automation has at least four qualities:
First, it distinguishes certainty from plausibility. A system should not collapse probability into declaration. If it is inferring, it should say so. If it is unsure, it should show that uncertainty in the wording, not just in hidden telemetry.
Second, it adapts to stakes. A casual recommendation and a clinical recommendation should not use the same linguistic tone. High-stakes contexts require language that is slower, clearer, and more explicit about limits.
Third, it respects user interpretation. If users are likely to read output as a command, the system must not phrase suggestions in a way that sounds mandatory. If users may misread technical language, the system should translate rather than merely abbreviate.
Fourth, it aligns evaluation with lived use. A concept should not be selected because it wins on an abstract score alone. It should also be tested in interaction, where wording, trust, and misunderstanding emerge.
A practical way to do this is to add a language audit to concept evaluation. Before choosing a design, ask a multidisciplinary team to review sample outputs for hidden assumptions, misleading certainty, tone mismatches, and jargon. Then test those outputs with real users, not just against idealized criteria. This is not extra polish. It is part of correctness.
A system that scores well but communicates badly is not fully designed. It is only partially specified.
Key Takeaways
- Do not treat language as presentation. In AI systems, wording changes meaning, trust, and user behavior.
- Use scores, but do not worship them. Numerical evaluation is useful for comparing concepts, but it often misses interpretive and social effects.
- Add a semantic layer to design reviews. Test whether the system makes uncertainty, limits, and inference visible in its language.
- Match tone to stakes. The same phrasing that works in a low-stakes assistant can be misleading in a high-stakes context.
- Run a language audit before launch. Review outputs for false certainty, jargon, hidden authority, and user misunderstanding.
Conclusion: The Interface Is a Theory of the User
The most important thing an AI system says is not its answer. It is its model of the person reading the answer. Every phrase assumes something about the user's expertise, patience, vulnerability, and willingness to trust. That is why AI design is never just about functionality. It is about constructing a relationship between machine output and human judgment.
Scoring systems help us choose among concepts, but language determines whether the chosen concept becomes intelligible, trustworthy, and safe in the real world. The deepest failure in AI design is not a low score. It is a high score attached to a system that speaks more confidently than it understands.
Once you see this, evaluation changes. You stop asking only which system is best. You start asking which system tells the truth in the most usable way. That is a more demanding standard, but it is also the only one worthy of automated intelligence.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣