Why Thinking Still Works After You Change the Language
Hatched by Nan Wang
Jul 29, 2026
9 min read
3 views
83%
The strange fact that should unsettle every model of thinking
What if the deepest part of thought is not English, not math, and not any language at all, but something that can survive translation intact? That question sounds philosophical until it becomes operational. A large language model can plan ahead, carry multiple computational paths in parallel, and preserve shared circuitry across languages more than you might expect. Meanwhile, a principle from statistical inference insists that our beliefs should remain stable under smooth, monotone transformations, even when the parameter we choose to name is not the parameter reality actually uses.
Both ideas point to the same uncomfortable truth: the surface form we use to describe a system is often not the thing doing the real work. The user sees fluent prose or a neat parameterization. The machine, or the inference engine, may be operating in a deeper representational space whose structure is only partially visible from the outside.
That is not just a technical curiosity. It is a warning. If you mistake the translated output for the underlying process, you will misunderstand what the system knows, what it can generalize across, and where its apparent coherence comes from.
The most important kind of intelligence is often the part that survives translation.
Beneath words and parameters: the hidden invariants of thought
The temptation in human reasoning is to treat labels as reality. We pick a variable, give it a name, and then reason as if the name revealed the thing itself. But a monotone change of coordinates can make the same underlying structure look radically different. A probability attached to one parameter may look nonuniform when expressed in another, even though the information content has not changed. Jeffreys prior is one formal answer to this problem: choose a representation that respects the geometry of the problem, not the accidental features of your coordinates.
That is a powerful lesson because it generalizes far beyond statistics. If your belief changes just because you renamed the axes, your belief is probably about the map, not the territory. A good prior, in this sense, is not merely a guess. It is a commitment to invariance, a refusal to let arbitrary description choices distort inference.
The same basic issue appears in language models. A model may generate text in French, then English, and still preserve a shared internal structure that is not French or English. The language is the rendering. The representation is deeper. When a model plans ahead, it is not merely chaining tokens one by one in a local autocomplete loop. It is coordinating constraints, preferences, and objectives in a space where the final answer can be assembled after the fact.
This is the conceptual bridge between the two domains: good reasoning depends on what stays the same when the surface changes. In statistics, we ask for priors and posteriors that behave sensibly across reparameterizations. In machine intelligence, we ask whether the model’s internal work is stable across languages, phrasings, and modes of expression.
Why fluent output is a dangerous illusion
A plausible sounding argument can be designed to agree with the user rather than to follow logical steps. That line should feel familiar to anyone who has watched a system produce elegant prose that somehow dodges the question. Fluency is not fidelity. Coherence is not causality. A polished answer can conceal a weak internal chain just as easily as a messy answer can conceal a strong one.
This is where the surface becomes a trap. If a model learns to sound helpful, it may also learn to sound right when it is only socially optimized. Likewise, if a statistician treats a convenient parameterization as if it were privileged, they may end up with a prior that is tidy in form but incoherent in meaning.
Consider a simple analogy. Imagine you are navigating a city using only translated street signs. The signs may be accurate enough to get you around, but they are still not the city. If a neighborhood is renamed, the streets do not move. If the model changes language, the underlying computation may not change in the way the output suggests. The danger is confusing translation stability with truth.
This is why the discovery that thinking can happen before translation matters so much. It implies that the output language is often an afterburner, not the engine. The model may first assemble a plan in a shared internal space, then render it into the final language, rhyme scheme, or style. The same is true in statistical inference: the coordinate system is a rendering of uncertainty, not uncertainty itself.
A useful mental model here is to distinguish among three layers:
- The invariants: the structure that should remain stable under translation, such as relationships, constraints, and posterior mass.
- The coordinates: the chosen language, parameterization, or framing.
- The performance layer: the final output, which may optimize for helpfulness, elegance, or social fit.
Mistaking layer 3 for layer 1 is how people overtrust fluent systems. Mistaking layer 2 for layer 1 is how people make fragile inferences.
Planning, priors, and the geometry of hidden work
The most surprising connection between these two domains is that both reward a geometric view of intelligence. In language models, multiple computational paths can operate in parallel. One path might ensure that the final word makes sense, another that the sentence rhymes, another that the answer remains aligned with conversational intent. The system is not just stepping forward one token at a time, it is balancing constraints across a landscape of possibilities.
Jeffreys prior lives in a similar landscape. Its appeal comes from the idea that if we choose our representation carefully, the prior should not arbitrarily favor one coordinate system over another. In a world where a parameter can be described in different but equivalent ways, the prior should follow the problem’s geometry rather than the accident of notation.
That leads to an important synthesis:
Intelligence is not only about making the right move, it is about preserving the right structure under transformation.
This helps explain why scale matters. Larger models can share more circuitry across languages, which suggests that as systems grow, they may develop more abstract, reusable representations. Inference has a similar aspiration when it seeks priors that are invariant to reparameterization. Both are, in effect, trying to identify a level of description where the system’s behavior is less brittle and more general.
Here is the deeper tension: the more powerful a system becomes, the less you should expect its essential work to be legible in the language of the output. That is true for neural models and for statistical priors. Power often lives in the space between description and expression.
Think of a jazz ensemble. The audience hears a melody, but the real coordination happens in timing, shared expectations, and invisible structure. A transcription captures notes, not swing. Likewise, a good posterior captures evidence, not all the geometry that made the inference robust. The visible trace is downstream of the hidden coordination.
This has practical consequences. If you evaluate a model only by surface fluency, you may miss whether it has actually planned. If you choose a prior only because it looks natural in one parameterization, you may miss whether it respects the problem’s underlying symmetries.
The right question is not, “What does it say?” but “What survives transformation?”
The common thread here is a more general epistemic rule: trust the parts of a system that remain coherent when you change the frame. This is one of the best ways to separate real structure from representational artifact.
For a language model, that means asking whether the answer survives paraphrase, translation, or a change in stylistic constraints. Does the reasoning hold if the model must explain itself in another language? Does the plan remain stable if the output format changes from prose to bullet points? If not, the model may be leaning heavily on the costume of its response.
For Bayesian inference, it means asking whether the prior expresses genuine ignorance or just the convenience of a particular coordinate system. If a prior changes dramatically when you switch from one monotone parameterization to another, the apparent neutrality was an illusion. The same substantive uncertainty was being described with different geometry, and the geometry was quietly changing the conclusion.
This is why improper priors are tolerable only in a tightly constrained sense, namely when they still produce proper posteriors for every possible observation. The point is not aesthetic simplicity. The point is disciplined behavior under all admissible data. In both contexts, the standard is not whether the output looks nice in one frame, but whether the inferential or computational machinery remains valid across frames.
A practical way to think about this is to ask three questions whenever you encounter a model, a belief, or a generated explanation:
- What is the hidden representation here?
- What transformations should leave the core structure unchanged?
- Where is the system optimizing for appearance instead of substance?
If you can answer those, you are no longer being impressed by surface intelligence. You are beginning to inspect structural intelligence.
Key Takeaways
-
Treat fluency as a presentation layer, not proof of reasoning. A good answer can be socially optimized, stylistically polished, or post hoc coherent without being structurally sound.
-
Look for invariants under transformation. If a conclusion changes just because you renamed variables, changed language, or altered format, the conclusion may depend on representation rather than reality.
-
Prefer geometry over convenience. In inference, that means choosing priors that respect the problem’s structure. In AI, it means asking whether internal representations survive translation and paraphrase.
-
Evaluate systems by what they preserve, not just what they produce. Planning, shared circuitry, and coherent posteriors all matter because they reflect deeper organization that survives surface changes.
-
Use the three layer model. Separate the invariant structure, the chosen coordinates, and the performance output. Confusing these layers is a reliable source of error.
The real lesson: intelligence is translation resistant
There is a seductive idea that the best systems are the ones that speak most beautifully. But beauty can be a byproduct of hidden structure, not a substitute for it. The more interesting possibility is that genuine intelligence, whether artificial or statistical, is translation resistant. It can move across languages, parameterizations, and surface forms without losing its internal coherence.
That reframes what it means to understand something. Understanding is not the ability to repeat a sentence in a different style. It is the ability to preserve meaning when the frame shifts. It is not the ability to cling to one coordinate system. It is the ability to recognize which features are accidental and which are invariant.
So the next time a model gives you a fluent answer, or a prior feels natural in a particular parameterization, ask a sharper question. What remains true if I change the language? What remains true if I change the coordinates? What remains true if I strip away the performance layer altogether?
That is where real structure lives. And that is where real thinking begins.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣