The Myth of Catastrophe: Why Bigger Systems Sometimes Reveal a Smaller World
Hatched by Rob Russell
May 09, 2026
6 min read
4 views
91%
What if apocalypse is mostly a local event?
We like to tell ourselves that the biggest shocks reshape everything. A supereruption should plunge the planet into darkness. A vastly larger neural network should think in a totally new way. Scale, in our imagination, turns ordinary systems into world-ending or world-expanding events.
But nature keeps offering a quieter, more unsettling lesson: big events do not always produce big uniform consequences. Ash can circle a continent without freezing every ecosystem beneath it. A model can become vastly larger without becoming qualitatively intelligent in the way people expect. The deeper story is not that size is irrelevant. It is that scale does not act like a hammer, it acts like a filter. It amplifies some structures, bypasses others, and exposes the hidden architecture underneath.
That is the real connection between volcanic ash and modern artificial intelligence. Both force us to confront a deceptively simple question: when does more of the same become something different, and when does it merely reveal that the system was already more resilient, more constrained, or more structured than we thought?
The first illusion of scale: more force, same world
The instinctive model of scale is linear. More energy means more damage. More parameters means more intelligence. More ash means more winter. We picture the system as a bucket, and once you pour in enough, it overflows in a predictable way.
Real systems are rarely buckets. They are networks of thresholds, bottlenecks, and buffers.
Consider the volcanic case. A supereruption sounds like the sort of event that ought to flatten ecological difference across vast regions. Yet evidence from East Africa suggests that even after a giant eruption, local climate effects may not have been uniformly catastrophic. Geography, atmospheric circulation, rainfall patterns, and ecological adaptation can absorb a shock that looks globally decisive from a distance. The headline event is real, but its consequences are mediated by the terrain it meets.
This same logic applies to large neural networks. For years, machine learning progress looked like a simple race to larger models. More parameters, more data, more compute, better performance. That part is true, but incomplete. Raw scale does not guarantee useful reasoning. What matters is not just size, but how scale interacts with structure: training signals, token sequences, intermediate representations, and the ability to externalize steps of thought.
Scale is not a substitute for organization. It is a stress test for organization.
That is why chain-of-thought reasoning matters. It does not merely make a model bigger in an abstract sense. It gives the model a way to spread computation across time, turning one opaque leap into several manageable moves. In other words, it changes the medium of intelligence. The question is not whether there is more power. The question is whether the power can be routed through the right channels.
The hidden lesson of both ash and attention: systems survive by partitioning shock
The most important thing about a catastrophe is often not the catastrophe itself, but the partitioning it reveals.
A volcanic eruption deposits ash broadly, but the ecological impact is partitioned by local conditions. Some areas lose light, others lose crops, still others barely notice. A global event becomes a mosaic of local realities. Similarly, a large model receives a flood of information, but only part of it becomes relevant in any given step. The rest is filtered, ignored, compressed, or deferred.
This is where the two stories become deeply instructive together. We often think intelligence is about brute force accumulation: more facts, more weights, more memory. But the real challenge is routing attention under constraint. The system must know what to ignore, what to preserve, and when to spend computation.
Think of a city during a blackout. The city does not need every street to remain equally lit. It needs hospitals, transit hubs, and emergency services to stay functional. Resilience comes from selective prioritization, not uniform strength. A model that reasons well is similar. It does not treat every token with equal importance. It creates a temporary hierarchy of relevance, then uses that hierarchy to move toward an answer.
This is one reason chain-of-thought is so powerful. It turns reasoning into a visible sequence of checkpoints. Each step is like a watershed, shaping what flows downstream. Without those checkpoints, the model may still contain the necessary information, but it lacks a clean way to serialize complexity. The result is not ignorance, exactly. It is congestion.
Volcanic systems, climate systems, and cognitive systems all share this feature: they are not judged only by what enters them, but by how that input gets partitioned across subsystems. A shock is never just a shock. It is a routing problem.
Why bigger can look weaker before it looks smarter
One of the most confusing features of scale is that it often produces a phase where progress is visible but not yet legible. A larger model may seem to memorize more, imitate better, or answer more fluently, while still failing at tasks that require disciplined multi-step inference. A massive eruption may leave evidence everywhere, yet not produce the expected climate collapse. In both cases, our intuitive theory of causation is wrong because we expect magnitude to map neatly onto outcome.
But systems often cross thresholds in strange ways. Before chain-of-thought style reasoning became prominent, a model might have possessed enough latent capacity to solve a problem but not enough internal scaffolding to reveal that capacity. The model was not empty. It was compressed beyond its own interpretability.
This is a useful mental model: capacity is not capability until it becomes addressable.
A reservoir can hold enormous water, but if the pipes are too narrow, the city still suffers a drought. Likewise, a neural network can store astonishing amounts of distributed knowledge, but if the task demands structured intermediate steps, the model may fail unless it can externalize them. Chain-of-thought functions like widening the pipes. It does not create intelligence from nothing. It lets latent competence move.
The same principle may explain why some disasters do not behave as expected. A supereruption is not a single lever that controls the planet. The climate system has feedback loops, regional variations, and adaptive responses. The eruption is real, but its effect is mediated by the system’s own architecture. That architecture can blunt, redirect, or localize the shock.
So the surprise is not that large events are weaker than expected. The surprise is that systems are often stronger, and less uniform, than our simplifications allow.
The deeper thesis: scale reveals the grammar of a system
Here is the synthesis that matters most: scale does not merely increase intensity, it reveals grammar.
By grammar, I mean the rules that determine how parts combine, which channels matter, and what kinds of intermediate states are permitted. A giant volcanic eruption reveals the grammar of climate and ecology by showing which regions are coupled and which are buffered. A giant neural network reveals the grammar of cognition by showing which forms of decomposition, abstraction, and stepwise reasoning the architecture can support.
That is why these stories belong together. Each one exposes the limits of a naïve, monocausal worldview. We want to say,
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣