Why the Best Prediction Models Sometimes Look Disappointingly Average
Hatched by Emil Funk Vangsgaard
Jul 30, 2026
9 min read
4 views
73%
The Strange Beauty of Near Zero Difference
What do a blood pressure drug comparison and a metabolic simulation of bacteria have in common? More than it first appears. In both cases, the deepest lesson is not that models are useless, but that reality is stubbornly harder to compress than we want it to be.
One result shows two treatments with nearly identical admission rates across all causes, cardiovascular disease, and respiratory disease. Another points to the ambition of predicting phenotype from genotype using kinetic models that integrate many kinds of biological data. Put those side by side and a provocative idea emerges: the best test of a model is not whether it produces an impressive story, but whether it can survive contact with the boring, near equal, ambiguous world.
That matters far beyond medicine and systems biology. We live in an era addicted to prediction. We want forecasts for markets, elections, customers, diseases, and cells. Yet the more complex the system, the more likely the signal is diluted by hidden variables, measurement noise, and the sheer entanglement of causes. The real challenge is not making a model that can explain everything after the fact. It is making one that can remain useful when the outcome is almost the same across competing explanations.
The hardest truth in prediction is that many important differences are small, local, and conditional. Grand theories often fail not because they are wrong, but because they are too smooth for a messy world.
When the Outcome Refuses to Cooperate
The most tempting interpretation of a near tie is that nothing interesting happened. But that is exactly backward. A near tie is where scientific humility begins.
In clinical decision making, a comparison that yields nearly identical admission rates suggests that the visible distinction between two interventions may be less important than the structure of the patients receiving them, the timing of care, adherence, dose, coexisting illness, and countless other factors. In other words, the apparent variable is not always the causal lever. The result invites a more mature question: what kinds of differences actually matter, and at what scale?
This is where mechanistic modeling enters the picture. A kinetic model of metabolism does not merely correlate input and output. It tries to encode how enzyme concentrations, reaction rates, and metabolic effectors combine to produce phenotype. That is a powerful ambition because it treats biology as an interacting system rather than a black box. But it also reveals the central tension of all high dimensional prediction: when you try to represent every relevant pathway, the model becomes fragile, and when you simplify too much, it becomes shallow.
Think of weather forecasting. A seven day forecast is not useless because it misses the exact time of rain in one neighborhood. It is valuable because it captures the larger structure of the atmosphere. Yet the closer you zoom in, the more the forecast must contend with tiny disturbances that can redirect the whole outcome. Biology behaves similarly. A drug effect, like a storm, is often the product of many local interactions, some visible, some hidden, some changing over time.
This is why comparisons that look statistically flat are intellectually rich. They force us to distinguish between three different things that are often confused:
- Prediction, what happens next.
- Explanation, why it happens.
- Intervention, what we can actually change.
A model can be strong in one and weak in the others. A treatment comparison can show no obvious difference in a population while still being clinically meaningful in a subset. A metabolic model can explain pathways beautifully while still failing to predict phenotype robustly under novel conditions. The challenge is not to choose between empirical data and mechanism, but to know when each is doing honest work.
The Real Problem Is Not Complexity, It Is Compression
Every model is a compression scheme. It takes a wildly detailed world and turns it into a smaller representation that we can think with, test, and use. The question is never whether compression will happen. The question is what gets preserved and what gets discarded.
This is why kinetic models are so compelling. They compress biology without abandoning causality. Instead of saying, “these variables are associated,” they say, “these reactions, concentrations, and rates interact in structured ways.” That makes them more interpretable than many purely statistical approaches. But it also means their success depends on whether the compression matches the scale of the problem.
A useful mental model is to imagine three layers of biological reality:
- Layer 1: Fine grain chemistry, where reaction rates and concentrations matter.
- Layer 2: Network behavior, where pathways and feedback loops dominate.
- Layer 3: Population outcome, where admissions, symptoms, and clinical events are observed.
A model that works brilliantly at Layer 1 may still be a weak predictor at Layer 3, because the path from molecule to outcome passes through many filters. Likewise, a treatment comparison at Layer 3 may reveal almost nothing about what happened at Layer 1. The mismatch is not failure, it is scale.
This helps explain why the most sophisticated models do not automatically outperform simpler ones. If the question is too coarse, elegance at the mechanistic level may not improve prediction. If the question is too specific, broad statistical regularities may miss the crucial causal lever. The art is not building the most detailed model possible. It is building the right compression for the decision at hand.
Good models do not merely fit the world. They fit the scale at which action is possible.
Consider a doctor choosing between medications. The clinically relevant question is rarely, “Which molecule is more interesting?” It is, “Which choice changes outcomes for this patient, given their history, comorbidities, and context?” A model that gives beautiful biochemical explanations but cannot separate meaningful risk from background noise has missed the practical scale. The same logic applies to bacteria in a bioreactor. If the goal is to engineer metabolism, the model must be detailed enough to capture control points, but not so brittle that it breaks under ordinary variation.
Mechanism Without Calibration Becomes Storytelling
There is a seductive failure mode in quantitative science: the more mechanistic a model sounds, the more people trust it. But mechanism alone is not enough. A model can tell a coherent story and still be wrong in ways that matter.
That is why benchmarking is so important. It acts as an antidote to theory inflation. Instead of asking whether a model feels explanatory, benchmarking asks whether it generalizes, whether it reproduces known behavior, and whether it stays stable when the data become unfamiliar. In biology, this is especially important because systems are not built like clean machine parts. They are adaptive, redundant, and context dependent.
A practical analogy is GPS navigation. A map can be extremely detailed, but if the road is closed, the map becomes less useful than a live traffic update. In the same way, a metabolic model that encodes detailed reactions must still be validated against observed phenotypes. Otherwise it is a beautifully drawn map of a place that no longer exists.
This leads to a deeper framework: mechanism is a hypothesis about structure, calibration is a hypothesis about relevance. Mechanism tells us what could matter. Calibration tells us what does matter under actual conditions. When the two align, prediction becomes powerful. When they diverge, even the most elegant model can mislead.
The same principle explains why real world clinical comparisons can look underwhelming. If two interventions produce nearly identical population level outcomes, it may not mean the science is trivial. It may mean the intervention sits inside a larger causal web, where other factors dominate. In that case, the key task is not to insist on a difference where none appears. It is to ask where the hidden heterogeneity lives.
That is the frontier where mechanistic modeling and clinical evidence can genuinely enrich each other. Population outcomes tell you whether a difference survives the real world. Kinetic and systems models tell you where such a difference might be created, amplified, or erased. Together they form a loop: observe, hypothesize, simulate, test, refine.
The Most Useful Models Are Honest About Their Blind Spots
The temptation in prediction is to promise certainty. The better ambition is to become precise about uncertainty.
In practice, this means asking three questions before trusting any model:
- What scale is this model actually designed for?
- What sources of variation are being ignored or averaged out?
- What decision will change if the model is wrong?
These questions matter because many errors are not numerical, they are category errors. A model meant to guide biochemical understanding is not automatically suitable for bedside prescribing. A study that compares two treatments in a broad population is not automatically informative about a specific subgroup with unusual physiology. And a dataset rich enough to fit a kinetic model may still be incomplete in the places that matter most.
The best thinkers in quantitative fields often share a habit that is easy to miss: they treat mismatch as information. If a model fails, they do not just ask how to improve the fit. They ask what the failure reveals about hidden variables, wrong assumptions, or mismatched scales. In that sense, failure is not the opposite of understanding. It is the route to it.
This is also why near equal outcomes are so intellectually valuable. They prevent us from overinterpreting noise as signal. They force us to admit that many interventions are not magic bullets, many causal stories are incomplete, and many differences wash out when exposed to the full complexity of real systems. That realization is frustrating, but it is also liberating. It means we can stop mistaking elegance for efficacy.
If you work in science, medicine, analytics, or product strategy, the same lesson applies. A model should not be judged only by whether it explains the past. It should be judged by whether it changes the next decision in a way that matters.
Key Takeaways
- Treat near equal outcomes as a signal, not a disappointment. They often reveal that the main driver is elsewhere, or that the effect exists only in specific contexts.
- Match the model to the decision scale. A mechanistic model can be brilliant at explaining pathways and still be the wrong tool for a population level question.
- Separate explanation from prediction. A model can make biological sense and still fail to forecast outcomes, especially when hidden variables dominate.
- Benchmark relentlessly. Elegant structure is not enough. Compare models against real behavior under varied conditions, not just familiar cases.
- Look for heterogeneity before declaring equivalence. Flat average effects may hide meaningful subgroup differences, timing effects, or context dependence.
The Reframing: Average Is Not Boring, It Is Diagnostic
The deepest connection between clinical equivalence and kinetic modeling is this: both teach us to respect the limits of average behavior. A population average can conceal the very forces that make a system interesting. But it can also protect us from the fantasy that every apparent difference is meaningful.
That is the paradox of modern prediction. The more detailed our models become, the more we must learn to live with the possibility that the world will answer, “not much difference.” Far from being a failure, that answer is often the first honest clue. It tells us our next job is not to add more complexity for its own sake. It is to find the missing structure, the hidden subgroup, the overlooked rate limit, the scale where causality actually lives.
In the end, the most powerful models are not the ones that make reality look tidy. They are the ones that help us understand why reality so often refuses to be tidy, and what to do next when it does.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣