The Counterfactual Is Not a Forecast: Why Causal Inference Depends on Rebuilding the World That Never Happened
Hatched by Nan Wang
Aug 23, 2026
12 min read
1 views
94%
What if the most important number in an experiment is not the outcome you observed, but the outcome that became impossible to observe?
A government introduces a congestion charge in one city. Traffic falls by 12 percent. A hospital adopts a new triage protocol, and mortality declines. A school district changes its curriculum, and test scores rise. In each case, the observed result is easy to report. The difficult question is almost invisible: What would have happened without the intervention?
That missing outcome is not a minor detail. It is the entire basis of a causal claim. Without it, we have a before and after, or a treated group and an untreated group, but not necessarily an effect. We may be looking at weather, economic growth, demographic change, a competitor’s failure, or a shock that would have altered the outcome anyway.
The deeper lesson is that causal inference is not primarily about predicting the future. It is about reconstructing a credible version of the past that never occurred.
The hidden object every causal claim requires
For each unit, whether that unit is a person, city, hospital, school, or country, imagine two possible outcomes. One is the outcome under treatment. The other is the outcome without treatment. Only one can ever be observed for the same unit at the same time.
This is the fundamental difficulty of causality. The effect of a policy on a treated city is the difference between its actual outcome after the policy and its potential outcome without the policy. The first quantity can be measured. The second is a counterfactual.
The average treatment effect summarizes the average difference across a population. The average treatment effect on the treated asks a narrower and often more relevant question: how much did the intervention affect the units that actually received it? That distinction matters because the treated units are rarely ordinary. Cities that adopt congestion pricing may have unusually severe traffic. Hospitals that implement a new protocol may be unusually innovative. Countries that impose a policy may be responding to a crisis.
The central problem is therefore not simply missing data. It is structured missingness. The missing outcome is attached to a unit that was selected for treatment, often for reasons related to its trajectory. A random control group solves much of this problem because treatment assignment breaks the connection between selection and potential outcomes. But in many real settings, random assignment is implausible, unethical, or impossible. There may be only one treated city, three treated hospitals, or a handful of regions.
This creates a temptation: find a statistical model, extrapolate the untreated outcome, and subtract it from the observed outcome. But a forecast can be accurate for the wrong reason, and a causal estimate can be precise while resting on an implausible comparison.
The counterfactual is not the most likely number in isolation. It is the outcome generated by a believable version of the same world, with the intervention removed.
That phrase changes how we should evaluate causal methods. The question is not merely, “How well does the model predict?” It is, “Why should this prediction represent what would have happened under non intervention?”
Why local resemblance beats global brilliance
Consider a city that introduces a congestion charge in January 2020. To estimate its effect, we might build a model using national income, fuel prices, employment, weather, and dozens of other variables. The model could fit traffic patterns across an entire country impressively well. Yet it might still be a poor guide to the city’s counterfactual.
The reason is simple: causal comparisons are local. We do not need a model that explains every city in every period. We need a credible untreated version of this city during this particular episode.
This is where a synthetic control becomes conceptually powerful. Rather than choosing one supposedly similar control city, we construct a weighted combination of untreated cities whose pre intervention trajectory resembles the treated city. The synthetic city is not intended to be a real place. It is a carefully assembled comparison that approximates the treated unit before treatment.
Imagine that the treated city’s monthly traffic index before the policy was:
- January: 100
- February: 103
- March: 101
- April: 106
- May: 108
Suppose three untreated cities show different patterns. One has the right average but the wrong trend. Another follows the trend but has a different level. A weighted combination of all three may reproduce the treated city’s path much more closely than any single city can.
If the synthetic city tracks the treated city before January 2020, then the gap that opens afterward becomes evidence about the intervention. It is still not proof by magic. The credibility depends on assumptions. But the design makes those assumptions visible and testable.
This is the first major connection between potential outcomes and synthetic comparison methods: the abstract missing outcome becomes an engineering problem. We are trying to construct a replacement world using observed units, while preserving the features of the treated unit that matter for its trajectory.
The crucial phrase is local fit. A method that fits the pre treatment period near the intervention may be more useful than a method that fits a broad historical record but misses the conditions immediately surrounding the intervention. Global predictive performance can conceal local failure. In causal inference, the region around the treatment date is often where the argument lives or dies.
A useful analogy is architectural restoration. If you are reconstructing a missing wall in an old building, you do not need a material that resembles every brick ever made. You need a material that matches the surviving structure at the point of connection: its texture, weight, aging, and load bearing properties. The counterfactual is the missing wall. Pre treatment fit is evidence that the reconstruction belongs to the same building.
The real assumption is not “the model is right”
Every counterfactual method relies on assumptions, but they are often expressed too vaguely. Analysts may say that the model is well specified or that the control group is comparable. Those phrases are too broad to be useful.
A better approach is to separate causal credibility into four questions.
1. Can the method reproduce the treated unit before treatment?
If a weighted combination of control units cannot track the treated unit during the pre treatment period, why should it represent that unit afterward? Poor pre treatment fit does not automatically invalidate an analysis, but it is a warning that the comparison lacks structural resemblance.
A good fit is not sufficient either. A flexible model can fit almost anything. The fit must be achieved without bizarre weights, implausible extrapolation, or dependence on a single unstable control unit.
2. Are the relevant forces shared?
The treated unit and its controls need not be identical. They need to respond similarly to the forces that would have affected them after treatment. These forces might include national economic conditions, seasonal patterns, technological change, or common shocks.
This is why adding more controls is not always helpful. A large donor pool can provide flexibility, but it can also introduce units whose apparent similarity is accidental. The relevant question is not, “How many controls do we have?” It is, “Do these controls span the untreated dynamics that matter for this unit?”
3. Would the disturbance process remain comparable?
Even a strong pre treatment fit cannot protect against a post treatment shock that affects the treated unit differently from the controls. A pandemic, natural disaster, factory closure, or political crisis may alter the relationship that previously held.
One way to reason about this is through the behavior of residuals, the differences between actual and estimated outcomes. If shocks are stationary and weakly dependent, and if their distribution would not be changed by the intervention, then patterns of pre treatment error can tell us something about post treatment uncertainty. Residuals can be permuted across time to generate placebo or uncertainty distributions.
This does not turn uncertainty into certainty. It creates a disciplined way to ask whether the observed post treatment gap is unusually large relative to the disturbances that occurred before treatment.
4. Is the result stable under reasonable construction choices?
A causal estimate should not depend entirely on one arbitrary modeling decision. Does the estimated effect survive modest changes in the control pool? Does it persist when the pre treatment window changes? Does it appear when treatment is falsely assigned to control units? Does it depend on one especially influential donor?
Stability is not cosmetic. It is evidence that the counterfactual is a property of the data rather than an artifact of the analyst’s choices.
These four questions reveal a broader principle: causal inference is an argument about invariance. We observe that a relationship held before treatment and argue that, absent treatment, it would have continued to hold afterward. Every method differs in how it constructs, probes, and defends that claim.
A concrete example: the policy that looked successful until the comparison changed
Suppose a region introduces a four day workweek for public employees. Productivity appears to rise by 8 percent in the first year. The region’s leaders declare success.
A naive before and after comparison is weak. The first year might coincide with an economic recovery, new software, or a change in how productivity is recorded. A single neighboring region may also be a poor control if it experienced a local strike or population shift.
A more careful analysis assembles a synthetic comparison from several regions. The weights are selected using pre policy productivity, employment, sector composition, public spending, and other predictors. The resulting synthetic region closely follows the treated region for six years before the policy.
After implementation, the treated region moves above its synthetic counterpart. That is more persuasive than the original before and after result, but the analysis is not finished.
The researcher should now ask whether the divergence is exceptional. Apply the same procedure to each untreated region as if it had received the policy. If many placebo regions show gaps as large as the treated region, the apparent effect may be ordinary noise. If the treated region’s gap is unusually large, confidence increases.
Next, examine timing. Did productivity begin rising immediately, or only after a separate investment? If the effect appears months before the policy, the design may be detecting anticipation or model failure. If the effect appears only in sectors directly exposed to the workweek change, the mechanism becomes more plausible.
Then test donor dependence. If removing one control region causes the estimate to vanish, the result is fragile. If a range of reasonable donor pools produces similar effects, the estimate is more stable.
Finally, distinguish the effect from the decision. The policy may have been adopted because leaders expected productivity to improve. The effect on the treated is still estimable, but the result may not generalize to regions with different motives, institutions, or baseline conditions.
This example shows why the best causal analysis behaves less like a single calculation and more like an investigation. The estimate is the conclusion of a chain of counterfactual checks.
The counterfactual budget: a practical framework
Analysts often treat assumptions as a checklist. A more useful mental model is a counterfactual budget. Every causal claim spends credibility in several places, and a weakness in one category must be acknowledged rather than hidden by strength in another.
The budget has four accounts:
- Similarity: How closely can observed controls reproduce the treated unit before treatment?
- Coverage: Do the controls represent the relevant combinations of underlying forces?
- Invariance: Why should the relationship between treated and control outcomes continue after treatment?
- Uncertainty: How much could the estimate change because of shocks, timing, weights, or model choices?
A study with excellent similarity but weak invariance is not automatically credible. A study with many controls but poor pre treatment fit is not rescued by sample size. A study with narrow confidence intervals but unstable weights may be measuring computational precision rather than knowledge.
This framework also explains why settings with few treated units require special care. Standard approaches often depend on plausible random assignment or large samples. When only a few units receive treatment, those assumptions become difficult to defend. Synthetic methods can help because they construct comparisons at the unit level, but they do not eliminate uncertainty. They shift the burden toward pre treatment information, donor quality, and stability.
In such settings, having many control units is useful only if there are enough pre treatment periods to learn the treated unit’s structure. A long historical record can reveal recurring relationships, seasonal variation, and the scale of ordinary shocks. But time alone is not enough. The distribution of disturbances must remain sufficiently comparable after intervention, and the policy itself must not transform the entire data generating process in a way that invalidates the reconstruction.
The practical implication is demanding but liberating: do not ask a method to answer a question that the design cannot support. If the data provide a credible local comparison but not broad generalization, report the local effect. If the pre treatment record is short, treat stability claims cautiously. If the intervention coincides with a unique crisis, say that the counterfactual depends on an assumption that cannot be fully tested.
Key Takeaways
-
Define the missing outcome before choosing a method. State explicitly what would have happened to the treated unit without the intervention. This prevents a generic forecast from being mistaken for a causal counterfactual.
-
Prioritize local pre treatment fit. A comparison that reproduces the treated unit near the intervention is often more valuable than a model with impressive average performance across unrelated settings.
-
Treat residuals as evidence about shocks, not as disposable noise. Their dependence, stability, and distribution help determine whether pre treatment uncertainty can inform post treatment inference.
-
Run stability and placebo checks as part of the main analysis. Vary the donor pool, time window, weights, and false treatment assignments. A result that survives reasonable perturbations is more credible.
-
Report the assumptions in operational language. Replace “parallel trends” or “good fit” with concrete claims about shared forces, unchanged shock behavior, and the conditions under which the comparison should remain valid.
The most important shift is conceptual. Causal inference is often presented as a search for the correct estimator. In practice, it is a search for a defensible alternate world. The estimator matters, but it is only the machinery. The intellectual work lies in deciding which observed units can stand in for the missing outcome, how closely they must resemble the treated unit, and which features of the pre treatment world can reasonably be carried into the post treatment period.
A causal estimate should therefore be read as a conditional achievement. It says: given this reconstruction, given this stability of shocks, given this local resemblance, the intervention appears to have changed the outcome by this amount.
That may sound less absolute than a headline number. It is also more honest and more useful. The goal is not to make the counterfactual disappear behind statistical confidence. The goal is to make its construction visible.
The strongest causal claim is not the one that hides the world that never happened. It is the one that shows us how that world was rebuilt, and lets us decide whether we believe it.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣