When the Outcome Arrives Late, Your Data Must Learn to Speak in Proxies

Nan Wang

Hatched by Nan Wang

Apr 17, 2026

11 min read

86%

0

The uncomfortable question behind every delayed result

What do you do when the thing you actually care about arrives too late to use?

That question shows up everywhere. A school district wants to know whether a reading intervention will improve adult earnings, but earnings will not be visible for years. A public health team wants to know whether a program changes long run mortality, but mortality is too slow and too rare. A company wants to know whether a new onboarding flow increases retention six months later, but it can only observe early engagement now. In all of these cases, the temptation is the same: grab the earliest visible signal and treat it as a stand in for the truth.

But proxies are dangerous. A proxy can be informative without being faithful. It can move in the right direction while still missing the mechanism. It can predict the outcome in one setting and fail in another. The central tension is not whether we can use short term signals, because we often must. The real question is: when does a proxy become evidence, and when is it just a convenient story?

That tension sits at the intersection of three problems that usually get discussed separately: causal inference, prediction, and finite sample reality. Together, they point to a deeper principle: the best proxy is not a shortcut around uncertainty, but a disciplined way of paying down it.


Why a good proxy is not just a smaller version of the truth

A lot of people think of surrogates as simple stand ins. If short term employment predicts long term employment, then maybe we can use the early measure and be done. That intuition is understandable, but incomplete. A useful surrogate is not just correlated with the outcome. It has to do something much stronger: it must carry the causal information that matters.

That is why the ideal surrogate condition is so demanding. In plain language, once you know the surrogate, the treatment should not tell you anything more about the long term outcome. If treatment still matters after conditioning on the surrogate, then the surrogate is missing part of the causal pathway. It is like using the temperature outside to predict whether a lake will freeze. Temperature helps, but if you ignore wind, depth, and salt content, you may still be surprised.

This leads to a useful mental model: a surrogate is not a mirror, it is a lens. A mirror reflects the outcome; a lens compresses many causal pathways into a smaller signal. If the lens is well designed, it preserves the structure that matters. If it is poorly designed, it distorts the picture while still looking scientific.

That distinction matters because proxy based inference often becomes attractive precisely when direct evidence is unavailable. Long term outcomes are delayed. Sample sizes are small. Costs are high. The practical impulse is to move fast. Yet the more fragile the evidence environment, the more tempting it is to mistake an informative shortcut for an identification strategy. The discipline is to ask not only, “Does this proxy predict?” but also, “What causal channels does it preserve, and what channels might it erase?”

A proxy is useful only when it is not merely correlated with the outcome, but structurally aligned with the causal story that generates the outcome.


The three hidden tests: can you trust the signal, can you transport it, can you afford the error?

The most powerful way to think about surrogate based inference is as a three part test. Each part answers a different skepticism.

1. Can you trust the treatment comparison?

If treatment assignment is not effectively random, then early differences may reflect selection rather than impact. In an experiment, this problem is often solved by design. In observational data, it must be addressed by adjustment. Without that, the surrogate could faithfully summarize the wrong thing.

This is the first layer of trust: the treatment effect on the proxy must be interpretable as causal. If the comparison is just between two naturally different groups, the estimated effect is not evidence of treatment impact, only group difference. In an A/B test, this sounds trivial. In the wild, it is everything.

2. Can you transport the proxy relationship from one sample to another?

The interesting move in surrogate methods is that the outcome and the treatment are often not observed in the same place. One sample gives you treatment and surrogate outcomes. Another sample gives you surrogates and the long term outcome. The bridge between them is a conditional relationship: how the long term outcome depends on the surrogate and pre treatment variables.

This is where comparability enters. It is not enough that the surrogate predicts the outcome in one dataset. The mapping from surrogate to outcome must be stable enough to move across samples. Otherwise the surrogate index is a local artifact, not a general tool.

A good analogy is translation. A phrase can mean one thing in one dialect and something slightly different in another. If you build a translation model from one community and deploy it in another, you are assuming not just shared words but shared semantics. Comparability is the semantic assumption of causal inference.

3. If the proxy is imperfect, how wrong can it make you?

Here is where the story becomes practically valuable. Real proxies are never perfect. The key question is not whether there is bias, but how much bias remains after the surrogate captures part of the causal pathway.

The insight is surprisingly sharp: the bias shrinks when the surrogate explains most of the treatment variation or most of the outcome variation. If the treatment is almost fully encoded by the surrogate, then little remains to be missed. If the outcome is almost fully explained by the surrogate, then the surrogate has already done most of the work.

This gives a usable intuition for deciding whether a proxy is worth trusting. A surrogate is safer when it is not just predictive, but absorptive. It absorbs the pathways through which treatment works. The closer it gets to doing that, the less room is left for hidden direct effects.

The quality of a surrogate is not only about predictive accuracy. It is about how much unmodeled causal energy remains after the proxy is built.


Why more data is not always better, and why fewer clusters can change the method

There is a second tension hiding in the background: even if your proxy is conceptually right, your statistical machinery may still be too brittle for the sample you have.

This is where the usual obsession with precision can become misleading. More variables do not always produce more truth. Sometimes they produce more noise, more overfitting, and more fragility. In some settings, adding more surrogate periods can actually reduce precision rather than improve it. That is counterintuitive but important: a proxy is not automatically better because it is richer. If the added pieces contribute little explanatory power, they may only inflate estimation error.

This is one reason variable selection in surrogate problems is not just a prediction exercise. Surrogates play two roles at once. They help predict the outcome, and they help explain treatment. The best selection strategy therefore focuses on intermediate outcomes that are strongly linked to the primary outcome or strongly linked to the treatment. In practice, that often means favoring the smallest set of signals that captures the largest share of the causal and predictive structure.

The same theme appears in inference with clustered data. If there are few clusters, conventional robust methods can become unreliable. Then the right move may not be to push harder with standard asymptotics, but to change the design or use an alternative strategy, such as synthetic control in settings where the number of independent units is too small for comfortable cluster based inference.

This connection matters because both problems are about thin identification. In one case, the signal is thin because the outcome is delayed and only partially observed through surrogates. In the other, the signal is thin because the effective number of independent units is small. In both cases, the question is not, “Can I calculate a standard error?” but, “Do I have enough independent information to support the claim I want to make?”

A powerful way to think about it is this: surrogates are a substitute for time, and clusters are a substitute for replication. When either substitute is scarce, your method has to become more conservative.


The deeper synthesis: causal inference is a theory of compressed evidence

At first glance, proxy outcomes and cluster robust inference seem like different worlds. One is about predicting delayed effects. The other is about getting uncertainty right when observations are not really independent. But they share a hidden structure.

Both are responses to a basic limitation: the world gives you partial information, and you need to make a decision anyway.

That is why the best way to understand surrogate methods is not as a trick for speeding up evaluation, but as a theory of compressed evidence. You begin with a full causal process, then replace it with a smaller representation that is easier to observe earlier or more reliably. The entire enterprise depends on whether that compression preserves the quantities you care about.

This perspective clarifies why the usual binary framing, valid or invalid, is too crude. Surrogacy is not a yes or no property in the practical world. It is a degree of compression. Some proxies keep most of the signal. Others discard critical structure. The useful question is not whether a surrogate is perfect, because almost none are. It is how close the compressed representation comes to being sufficient for the causal target.

That is also why the best diagnostics are not purely in sample. A surrogate that looks promising today may fail tomorrow if the environment shifts, the mechanism changes, or the sample relationship between early and late outcomes weakens. Validation over time, especially by checking whether the surrogate index tracks later experimental outcomes as time passes, turns the proxy into a living object rather than a fixed assumption.

Think of this as causal calibration. The proxy is not declared valid once and for all. It is repeatedly tested against reality. As more of the long term outcome becomes visible, you ask whether the surrogate index still tracks it. If it does, confidence grows. If it does not, the proxy was never carrying the whole load.


A practical framework for deciding when to trust a surrogate

When you are choosing whether to use a short term proxy for a delayed outcome, the right questions come in a sequence.

Step 1: Is the short term signal plausibly on the causal path?

If the surrogate has no connection to the mechanism by which treatment works, it is just noise with better timing.

Step 2: Can the treatment effect on the surrogate be estimated cleanly?

If assignment is randomized, that part is easy. If not, you need adjustment, and if the adjustment is weak, the surrogate may only reproduce selection bias faster.

Step 3: Is the mapping from surrogate to outcome stable across samples?

A surrogate index built in one context must transport into the other. Without this comparability, the bridge collapses.

Step 4: How much of the remaining uncertainty does the surrogate explain?

If the surrogate explains very little of the treatment or the outcome, it will not save you much. If it explains a great deal, it may be a credible compressed summary.

Step 5: What is the cost of being wrong?

Sometimes an approximate answer is sufficient for screening, prioritization, or early warning. Sometimes it is not. The lower the tolerance for error, the more evidence you need before leaning on a surrogate.

This framework is valuable because it shifts the question from “Can I substitute one number for another?” to “What kind of information compression can this application survive?” That is the right question in policy, medicine, education, and product evaluation alike.


Key Takeaways

  1. Do not treat a proxy as a smaller version of the truth. Treat it as a compressed causal summary that must preserve the relevant pathways.
  2. Check three things separately: causal identification, transportability across samples, and sensitivity to misspecification. A proxy can fail at any one of them.
  3. More surrogate information is not always better. Add only the signals that meaningfully explain the treatment or the outcome.
  4. Use validation over time whenever possible. A good surrogate should track later outcomes as they become observable.
  5. When information is thin, change the inferential strategy, not just the standard errors. If the design is weak, methods like synthetic control may be more credible than forcing a cluster based answer.

The real lesson: evidence is often a race against delay

The deepest insight here is not technical. It is philosophical. We often think of measurement as a neutral act, as if the only question is whether we have observed the right thing. In reality, measurement is an act of compression under uncertainty. Long term outcomes arrive late. Some samples are too small. Some comparisons are contaminated. So we substitute, approximate, and borrow strength wherever we can.

That is not a weakness of empirical work. It is the core condition of empirical work.

The trick is to stop asking whether a proxy is perfect and start asking whether it is honest about what it leaves out. The best surrogate does not pretend to be the outcome. It earns its place by carrying enough of the causal structure to make early inference worthwhile, while being transparent about the parts of reality it cannot recover.

In that sense, the goal is not to eliminate delay or uncertainty. The goal is to build a representation of the world that is good enough to act on, without confusing convenience for truth. Once you see that, proxies stop looking like shortcuts and start looking like what they really are: a test of whether your understanding of the mechanism is strong enough to survive compression.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣