Why the Future of Intelligence May Look More Like a Crystal Than a Sentence

Alessio Frateily

Hatched by Alessio Frateily

Jul 25, 2026

10 min read

88%

0

The strange coincidence of magnets and language models

What do a room temperature antiferromagnet and a masked language model have in common? At first glance, almost nothing. One lives in the physics of electrons moving through a crystal lattice. The other lives in the statistics of tokens being filled in by a neural network. Yet both point to the same unsettling idea: the behavior you care about may come from structure, not from the mechanism you assumed was essential.

That is the real surprise. A Hall effect in a collinear antiferromagnet looks impossible if you only trust the textbook picture that Hall responses need obvious magnetization. A strong language model built on diffusion rather than next token prediction also looks implausible if you believe intelligence belongs to autoregression. In both cases, the deeper lesson is that what appears to be the “obvious engine” is often just one convenient route through a larger landscape of symmetry, constraint, and distributed coordination.

This matters because we tend to mistake surface form for necessity. We see a ferromagnet and assume net spin is the only way to get transverse transport. We see an autoregressive model and assume left to right prediction is the essence of language intelligence. These are comforting stories, but they are too small.

The deeper pattern is not that one system imitates another. It is that both reveal how rich behavior can emerge when a system is trained to respect the full structure of its world.


When the expected mechanism is not the real cause

The Hall effect begins with a simple geometric intuition. Push electrons one way, and in certain materials they veer sideways. For a long time, that sideways motion was tied to broken time reversal symmetry in familiar magnetic orders: ferromagnets, noncollinear antiferromagnets, skyrmions. The intuition was seductive because it matched the visible picture. If the spins are arranged in a twisting or net polarizing way, then maybe the electrons are being “steered” by that order.

But the discovery of a strong Hall response in a collinear antiferromagnet complicates the story. In such a material, the spins are not swirling in a fancy pattern, and the net magnetization can still vanish. Yet the crystal symmetry itself can generate the effect. The important point is not merely that magnetism matters. It is that the lattice, the symmetry, and the internal spin structure together define what kinds of motion are allowed.

A useful analogy is traffic in a city. You might think congestion depends mostly on the number of cars, but sometimes the decisive factor is the road geometry, one way streets, hidden barriers, and synchronized signals. Two cities can have the same car count and dramatically different flow because the architecture encodes the rules of movement. The Hall effect in these materials is like a city where the layout itself creates a lateral bias, even if no single intersection looks dramatic.

Now translate that to language modeling. The usual story says autoregression is special because language is generated one token at a time. That seems plausible because speaking and writing do unfold sequentially. But the more interesting claim is that the intelligence of large language models may come not from the left to right mechanism itself, but from generative modeling as a principle: approximating the true data distribution by maximum likelihood.

That changes the question. Instead of asking, “Why does autoregression work so well?” we should ask, “What properties does a good model need to capture the structure of language at all?” Once framed that way, masked diffusion becomes not a weird detour but another path through the same terrain. It learns by partial corruption and reconstruction, by repeatedly asking what belongs in the gaps. That is not a lesser form of intelligence. It is a different way of respecting structure.

The hidden symmetry between these two stories is this: a system can produce a directional, useful effect without the mechanism you assumed was doing the directional work. The sideways electron flow is not dependent on a ferromagnet’s obvious net spin. The impressive behavior of the language model is not dependent on autoregressive token chaining.


The real engine is not sequence, it is constraint

This is where the two domains become philosophically interesting. Both suggest that constraint is more fundamental than sequence.

In physics, the electron does not move arbitrarily. Its motion is constrained by symmetry, band structure, and the topology of the material. The fact that a current responds transversely is not a quirk added after the fact. It is a consequence of the allowed states of the system. A material can therefore be engineered to reveal behavior that seems to violate intuition, while in fact obeying a deeper order.

In language, a model does not merely predict the next word. A strong model internalizes the distribution of plausible continuations across many levels of abstraction: syntax, semantics, style, facts, discourse structure. Autoregression is one sampling strategy, but it is not the whole story. Masked diffusion pushes harder on the idea that a model should learn the global compatibility of tokens, not just the next step in a chain.

A good way to see this is through puzzle solving. If you solve a crossword left to right, you are using sequence. If you solve it by filling in intersecting clues, you are using constraint. The second method often feels more intelligent because it exploits the fact that the answer must satisfy many relationships simultaneously. Diffusion style sampling does something similar. It starts with everything hidden, then iteratively reveals tokens that best satisfy the whole pattern.

This is why the diffusion framing is so interesting. It says intelligence may be less like reciting a sentence and more like finding a stable configuration. That is a profoundly different metaphor. One model is a chain. The other is an equilibrium process. One advances by making local commitments. The other improves by globally revising uncertainty until the whole structure becomes coherent.

The crystal Hall effect carries a related lesson. The sideways motion is not a simple “spin causes deflection” story. It is a stability story, a symmetry story, a selection story. The material’s internal order makes some responses possible and others impossible. In both physics and machine learning, the most interesting behavior may emerge not from adding more force, but from changing the constraints under which the system settles.

Sequence is often just the path. Constraint is the landscape.


A framework for thinking about intelligence as symmetry plus inference

If we want a deeper synthesis, we need a better framework than “both are complicated systems.” Here is one:

1. Symmetry tells you what cannot matter directly.

In a crystal, many naive intuitions fail because the symmetry of the lattice forbids certain responses. In a language model, the training objective can make the exact generation order less important than the distributional structure the model learns. What seems central at the surface may be incidental at the level of the governing rules.

2. Structure tells you where the signal can hide.

A collinear antiferromagnet can still produce a large Hall effect if the structure supports it. A diffusion based language model can still be highly capable if the pretraining objective learns the right statistical regularities. The surprising behavior is not magic. It is latent potential made visible by the right probe.

3. Inference is the art of revealing hidden compatibility.

Electrons move through the material in ways that reveal the underlying allowed states. A masked model reveals language by repeatedly inferring missing pieces. In both cases, the system is not merely generating outputs. It is uncovering what configurations are consistent with the underlying order.

This framework helps resolve a common misconception in AI research and, more broadly, in systems thinking. We often overidentify intelligence with a visible procedure. But the procedure is only one way to expose a deeper geometry. Autoregression is like reading a novel one sentence at a time. Diffusion is like restoring a damaged fresco by continually asking what colors and shapes are most compatible with the remaining fragments. The fresco was always there. The process only determines how it comes back into view.

Likewise, an antiferromagnet may look “nonmagnetic” if you focus only on net moment. But the deeper internal arrangement can still encode enough asymmetry to generate a sizable transverse response. The absence of one familiar sign does not imply the absence of structure.

This is a valuable mental shift for anyone designing models or materials. Ask not only what is present, but what is forbidden, implied, or stabilized by the rules. That is often where the real capability hides.


The actionable insight: stop worshipping the obvious path

There is a practical temptation in both fields to overfit to the successful default. In physics, that can mean assuming only ferromagnetic order matters for Hall effects. In AI, it can mean assuming autoregression is the natural and therefore inevitable architecture for language intelligence.

But progress often comes from separating capability from implementation. If the capability is robust, then multiple implementations may exist, each revealing a different facet of the same underlying principle. That matters for science, engineering, and even product design.

For materials research, the lesson is to search beyond the visually dramatic cases. A material does not need obvious net magnetization to host useful transverse transport. Symmetry analysis can identify families of candidates where the effect is hidden in the crystal itself. That is a method, not just a result: look for the governing constraints, then scan for systems where they permit the response you want.

For AI, the lesson is equally important. If intelligence can emerge from a generative objective plus the right optimization and data structure, then architectural dogma should loosen. The question becomes not “Which one mechanism deserves loyalty?” but “Which generative process best captures the full geometry of the problem?” Sometimes left to right generation is the best tool. Sometimes iterative global refinement may better match the task. A mature field should be able to hold both.

This is also a warning against shallow analogies. The point is not that crystals and language models are literally the same. The point is that both show how emergent behavior can be more architecture dependent than mechanism dependent. That is a powerful pattern. It suggests that the frontier is not always in making the same thing bigger, but in discovering a better description of what the system is actually doing.


Key Takeaways

  1. Separate capability from mechanism. Do not assume that the most visible process, whether autoregression or ferromagnetism, is the only path to the result you care about.

  2. Look for constraint before force. Many strong effects come from symmetry, structure, and compatibility rather than from obvious intensity or motion.

  3. Use global consistency as a design principle. Whether solving a puzzle, training a model, or engineering a material, ask what configuration best satisfies the whole system, not just the next step.

  4. Search for hidden degrees of freedom. A system may appear simple on the surface but contain latent channels for dramatic behavior if the internal order is right.

  5. Treat alternative implementations as experiments in meaning. If two different mechanisms can produce the same capability, the comparison can reveal what the capability truly depends on.


The deeper conclusion: intelligence may be a property of allowed states

The most interesting connection between a crystal Hall effect and a diffusion based language model is not novelty for its own sake. It is the shared possibility that powerful behavior arises when a system is trained, shaped, or arranged so that only certain states are easy to reach.

In the crystal, the electrons do not need a dramatic net magnetization to veer sideways. The lattice and spin structure already define a subtle map of allowed motion. In the language model, intelligence may not depend on a token by token march through the sentence. It may come from learning a landscape of mutual constraints so well that the model can reconstruct meaning from partial information.

That reframes how to think about design in general. The goal is not always to choose the most direct path. Sometimes the better strategy is to construct the right geometry, then let the system discover the path that was hidden inside it all along.

If that is true, then the future of intelligence, in machines and in matter, may belong not to the loudest mechanism but to the deepest structure.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣