Why the Best AI Systems Need a Better Theory of Foreknowledge

Xuan Qin

Hatched by Xuan Qin

Jul 18, 2026

10 min read

87%

0

The Strange Question Hiding in Plain Sight

What does it really mean for a system to learn from experience? At first glance, the answer seems obvious: it looks at the past, finds patterns, and uses them to predict the future. But that simple story breaks down quickly when we try to build intelligent machines, or even when we try to measure whether one variable truly influences another.

The deeper problem is this: prediction is not the same as explanation, and precedence is not the same as causation. A model can seem brilliant because it forecasts well, while remaining blind to the actual structure of the world. A dataset can seem powerful because it produces impressive behavior, while quietly depending on hidden assumptions about the kinds of patterns it contains. The real challenge is not just making systems that respond, but making systems that respond for the right reasons.

That tension connects two ideas that look unrelated on the surface. One is about open instruction tuning, the idea that a human generated dataset can teach a large language model to behave in a more conversational, helpful way. The other is about Granger causality, a statistical method that asks whether one time series improves prediction of another, without claiming true causation. Together, they point to a deeper lesson: modern intelligence, whether human, machine, or statistical, often begins with forecasting, but it only becomes trustworthy when we stop confusing foresight with truth.


The Lure of Prediction

Prediction has a seductive quality because it is measurable. If a model predicts better after adding a variable, we feel we have discovered something real. If a language model answers more helpfully after instruction tuning, we feel we have improved intelligence itself. In both cases, the gain is visible, immediate, and easy to celebrate.

That is why prediction becomes the default proxy for understanding. It gives us a concrete score, a benchmark, a before and after. In time series analysis, the Granger framework turns the vague word “influence” into something testable: does the past of X help predict the future of Y beyond Y’s own history? In language models, instruction tuning turns raw text completion into something more interactive: does training on human instructions make the model more useful in dialogue?

But there is a trap here. Prediction can improve for many reasons that have little to do with the thing we think we have captured. A weather app may become better because it now sees correlated signals from pressure and humidity, not because it understands storms. A chatbot may become more fluent because it has learned surface patterns of helpfulness, not because it has acquired judgment. The map can get sharper while the territory remains mysterious.

This is why the phrase “Granger causality” is so revealing. It is often better described as precedence: X comes before Y in a way that improves forecast accuracy. That is enormously useful, but it is not the same as saying X causes Y. It identifies statistical advantage, not metaphysical truth.

A better forecast is evidence of informational value, not proof of causal power.

That distinction matters more than it first appears, because it mirrors what happens in machine learning at large. We increasingly build systems by asking what improves output quality, not what model of the world is actually being learned. The result is a kind of operational intelligence, powerful but potentially brittle.


Instruction Tuning as a Kind of Foreknowledge

Instruction tuning is often described as a way of making a model follow commands. But that description undersells what is really happening. A human generated instruction dataset does more than teach syntax or style. It teaches a model a relationship between context and response: when someone asks something in this way, here is the kind of answer that tends to satisfy the hidden human goal.

In other words, instruction tuning does not merely teach the model to speak. It teaches the model to forecast what usefulness looks like. The model sees many examples of prompt and response pairs, and it learns a regularity: certain kinds of replies tend to be judged as helpful, clear, safe, or complete. This is not unlike a Granger test in spirit. The model is not discovering ultimate meaning. It is learning which past signals are informative for producing a better next token sequence.

Think of a good restaurant server. They do not know your life story, but they infer a surprising amount from your first request. Are you in a hurry? Do you want recommendations? Are you uncertain or decisive? Their skill lies in reading weak signals and anticipating what comes next. Instruction tuning does something analogous at scale. It teaches the model to treat the user’s wording as an evolving time series of intent.

That analogy is useful because it exposes a hidden truth: helpfulness is a forecasting problem. To answer well, a system must predict what kind of response will best fit the user’s actual need, not merely their literal words. In that sense, the magic of instruction tuning is not just compliance. It is anticipatory alignment.

But anticipatory alignment has limits. A model can learn that certain kinds of responses tend to be rewarded without learning why they are rewarded. It can optimize the surface of interaction while missing deeper context. This is precisely the kind of limitation that Granger analysis warns us about in statistics. Past correlation is informative, but it is not enough to reveal hidden mechanisms, interventions, or counterfactuals.


When Good Forecasts Fool Us

The danger of all powerful predictive systems is that they tempt us into epistemic overconfidence. If a model is helpful, we start to assume it understands. If a variable forecasts another variable, we start to assume it drives it. But prediction can be a downstream artifact of many hidden structures: shared causes, selection effects, noise, or training bias.

Consider a medical example. Suppose hospital visits for asthma rise shortly after pollen counts rise. Pollen may Granger cause asthma admissions in a forecasting sense, because its past values improve prediction. That is useful for planning. But if we mistake that for the full causal story, we may miss other factors such as air pollution, medication access, or housing conditions. Forecasting tells us where to look, not what to believe.

Now consider a large language model tuned on human instructions. It may become excellent at producing confident, coherent answers. Yet coherence itself can be misleading. A model may learn that a polished tone is often rewarded, which means it can forecast the shape of a good answer, even when the underlying reasoning is shallow. The better the forecast, the more plausible the illusion.

This is where the analogy deepens. Granger causality works because the past of one series can contain information useful for predicting another. Instruction tuning works because human examples contain information useful for predicting desirable responses. In both cases, the method is a powerful compression of experience into usable signal. In both cases, the danger is treating compressed signal as complete understanding.

Intelligence often begins as pattern recognition, but wisdom begins when pattern recognition admits its own blind spots.

The most important question is not whether a system predicts well. It is whether its predictions are robust under change. If you alter the environment, the wording, the distribution, or the incentives, does the learned behavior still hold? A model trained only on helpful-looking answers may fail when asked to reason, to disclose uncertainty, or to refuse a harmful request. A Granger relationship may vanish when the regime changes, because the informational link was never fundamental in the first place.


A Better Mental Model: From Causality to Forecast Usefulness to Truth

To connect these ideas productively, it helps to think in three layers.

1. Forecast usefulness

This is the most basic layer. Does past information from X improve prediction of Y? Does a set of human instructions help a model produce a more useful response? This layer is about operational value. It answers: does the signal help?

2. Mechanistic explanation

This layer asks why the predictive relationship exists. Is X genuinely upstream of Y, or do both reflect a hidden driver? Is the instruction data teaching a model a general conversational capability, or just a narrow set of response patterns? This layer is about structure. It answers: what is actually happening?

3. Counterfactual truth

This is the deepest layer. If we intervene on X, does Y change for the reason we expect? If we alter the training distribution, will the model still behave well? This layer is about causation in the strongest sense. It answers: what would happen if the world were different?

Most modern AI systems are strongest at the first layer and weakest at the third. That is not a flaw unique to AI. Human cognition works similarly. We often start with useful heuristics, then mistake them for truths, and only later discover the limits of our model. The point is not to abandon forecast usefulness. It is to place it in its proper hierarchy.

This three layer framework is helpful because it prevents a common mistake: treating all gains in predictive performance as equivalent. A chatbot can become more useful without becoming more truthful. A time series model can become more accurate without revealing causality. Progress at layer one is real, but it should never be confused with progress at layer three.


What Open Datasets and Granger Tests Share

There is a surprising moral symmetry between open instruction datasets and Granger causality tests. Both are attempts to make invisible structure legible. One does it by showing a model many examples of human intent and response. The other does it by testing whether one series carries predictive information about another. Both transform opaque behavior into a usable signal.

That is why openness matters. A human generated instruction dataset can be inspected, debated, refined, and criticized. Likewise, a forecasting relationship can be checked against assumptions, model specifications, and alternative explanations. In both domains, transparency does not guarantee truth, but it does make error easier to detect.

Closed systems are especially dangerous when they are highly predictive. If a model appears to work but cannot be examined, we may not notice that it is exploiting a brittle shortcut. If a forecasting relationship is strong but untested across conditions, we may not notice that it only holds in one regime. Openness does not solve the problem of inference, but it reduces the cost of being wrong.

This is the real connection between the two ideas: both point to an era where intelligence is increasingly engineered from examples, but the quality of those examples and the humility of the interpretation determine whether the result is useful or merely convincing.

If you train a model on human instructions, you are not just teaching it to answer. You are shaping what it thinks a good answer looks like. If you use Granger tests to analyze a system, you are not just finding links. You are deciding what counts as evidence of useful precedents. In both cases, the method encodes a philosophy of knowledge: learn from what came before, but do not confuse predictive leverage with ultimate explanation.


Key Takeaways

  1. Treat prediction as evidence, not proof. A model or variable that improves forecasting has informational value, but it does not automatically reveal causation.

  2. Separate usefulness from truth. A system can be highly useful at generating answers or forecasts while still being shallow, biased, or brittle.

  3. Ask what survives distribution shift. The best test of a learned relationship is whether it still works when conditions change.

  4. Use open examples to expose hidden assumptions. Whether training a model or analyzing a time series, transparency makes it easier to spot shortcuts and confounders.

  5. Upgrade your mental model from correlation to intervention. Forecasting tells you where to look, but only counterfactual thinking tells you what really matters.


The Real Lesson: Intelligence Is Anticipation, But Truth Is More Than Anticipation

The deepest connection between open instruction tuning and Granger causality is not technical. It is philosophical. Both remind us that intelligence, in practice, often starts as a machine for anticipation. We learn what tends to come next. We compress experience into patterns that improve response. We call this being smart because it works.

But a world built only on anticipation would be a world of elegant illusions. The model would answer smoothly. The statistic would forecast accurately. The system would look intelligent from the outside and remain only partially understood from within. That is why the next frontier is not merely better prediction, but better discrimination between what predicts, what explains, and what truly changes reality.

The temptation is to worship the forecast. The wiser move is to respect it without mistaking it for the whole story. In that sense, the most important question in AI is not whether a model can learn from examples. It already can. The real question is whether we can teach ourselves, and our machines, to recognize when examples are only pointing toward the truth, not containing it entirely.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣