The Hidden Art of Estimating What Never Happened
Hatched by Nan Wang
May 17, 2026
11 min read
4 views
86%
The problem behind every counterfactual is not prediction, but membership
What if the hardest part of estimating a treatment effect is not forecasting the outcome, but deciding who belongs in the comparison group?
That is the quiet problem running through very different kinds of causal inference. In one setting, we want the effect of a job training program for the people who would actually comply with it. In another, we want the effect of a policy on a treated unit by comparing it to a counterfactual world that never happened. These sound like different problems, yet both hinge on the same uncomfortable truth: the quantity we want is never directly observed, so we must infer it by reconstructing an invisible population.
This is why the usual language of prediction is slightly misleading. We are not merely asking, “What would have happened?” We are asking, “Which units would have behaved like this under the intervention, and which units would have behaved like that?” In other words, the real object of inference is often latent membership: compliance types, principal strata, or counterfactual analogs. Once you see that, many causal methods stop looking like separate tools and start looking like different answers to the same deeper question.
Causal inference is often not a problem of estimating effects first. It is a problem of discovering the right hidden population first.
Why “good fit” is not enough
A common instinct in both policy evaluation and principal causal effect estimation is to search for the best possible model fit. If the counterfactual outcome can be predicted accurately, or if compliance can be classified well, then the causal estimate should be trustworthy. But this instinct hides a trap: global fit can be less important than local alignment with the right comparison units.
In principal stratification, one may use pretreatment covariates to estimate principal scores, then match, weight, or regress to estimate the complier average causal effect. The logic is intuitive. If certain covariates predict who would comply, then people with similar covariates should be similar in their latent response types. But the crucial assumption is more delicate than it looks: principal stratum membership must be conditionally independent of potential outcomes given observed covariates. If that fails, then even a well calibrated compliance model can produce biased effects.
The same fragility appears in synthetic control and related policy methods. A treated unit is approximated by a weighted combination of controls, but the point is not to achieve perfect prediction everywhere. The point is to match the treated unit closely in the pre intervention period, where the counterfactual is anchored. This is why a method can tolerate a weak global model and still do well, provided it achieves a good local fit where it matters.
The deeper connection is that causal identification often depends on the stability of relationships inside a narrow region of the data, not on the overall excellence of the model. A model that is slightly wrong everywhere can be safer than one that is brilliant on average but wrong about the subgroup that carries the causal estimand.
Think of it like hiring a tutor for a student who struggles with algebra. A beautiful model of general academic performance is less useful than a rough but precise understanding of how this specific student responds to algebra practice. The target is not the average learner. The target is the learner whose hidden response type makes the intervention meaningful.
The shared assumption hiding under different names
At the heart of both domains sits a family of assumptions that look technical but are really statements about transportability across hidden states.
In principal score methods, the key claim is that, conditional on observed covariates, the control potential outcome does not differ across principal strata. That is, among people with the same pretreatment profile, compliers and never takers would have had the same outcome under control. This allows control group members to be scored, matched, or weighted as if their latent compliance type were partially observable from covariates.
In counterfactual policy evaluation, a related assumption appears in a different form: the error process before intervention is informative about the error process after intervention, provided the distribution of shocks is stationary and weakly dependent and remains invariant under the intervention. Here the logic is not about compliance, but about invariance. What happened before can stand in for what would have happened after, because the disturbance structure is believed to persist.
These are not the same assumption, but they are cousins. Both say that some hidden structure survives across regimes. Both give us permission to use observed data from one regime to infer what unobserved data would have looked like in another. Both fail when the intervention changes not just the mean, but the very mechanism that generates variation.
That is the key insight: causal inference is an exercise in testing whether the right invariances survive the intervention.
A useful mental model is to imagine the data as having two layers. The top layer is what we see, outcomes, treatment assignment, compliance, post intervention trajectories. The deeper layer is the stable geometry that may or may not persist, such as compliance propensities or error distributions. Good causal methods are methods for estimating the top layer while borrowing structure from the deeper layer. Bad causal methods assume the deep structure is stable when it has actually shifted.
The paradox of using prediction to estimate what is not observed
Prediction and causation are often treated as opposites. Yet both of these frameworks rely on prediction as a bridge into unobserved worlds. That does not mean they are the same. It means prediction is only useful when it is disciplined by the causal question.
In principal effect estimation, one can estimate principal scores using treatment group data, then project those scores onto the control group. The treatment group serves as a training set for compliance behavior. This is elegant, but it also exposes a subtle asymmetry: because only one side reveals actual treatment receipt, the model is not learning from symmetric data. It is learning from an observed manifestation of an unobserved latent class.
Likewise, in counterfactual policy analysis, one can build synthetic controls, factor models, or matrix completion models to predict the treated unit without intervention. But prediction alone does not prove causal validity. The model may fit the pre period and still fail in the post period if the intervention breaks the error invariance or if the treated unit was never well approximated by the control donor pool.
This produces a paradox: the better your predictive machinery becomes, the easier it is to confuse fit with identification.
That is why the most powerful causal tools are often not the most flexible, but the most disciplined. Matching, weighting, constrained least squares, and constrained Lasso all impose structure. They force the analyst to say, “I am willing to borrow strength only through these channels, and only under these invariances.” This is not a weakness. It is the source of credibility.
Imagine trying to infer the health of a river from sensors upstream and downstream. A model that predicts water temperature perfectly at every known sensor is not enough if a dam changes the flow pattern in the exact stretch you care about. You need a model that respects the physics of the system, not merely its historical correlations. Causal inference is similar. It is prediction under constraint, prediction with a map of what may remain unchanged.
Local truth beats global beauty
One of the most important lessons shared across these methods is that local truth matters more than global beauty.
In policy evaluation with few treated units, assignment mechanisms are often hard to model directly. Random assignment is implausible. The aim then shifts from discovering the whole assignment process to finding a credible local approximation to the treated unit. That is why methods can require a large number of pre treatment periods and emphasize invariance of the error distribution rather than a perfect structural model of the entire panel.
In principal effect estimation, the same instinct appears in the decision to focus on compliers rather than the full population. The causal effect for compliers is meaningful because it refers to the subgroup whose treatment status is actually moved by assignment. Never takers can be weighted to zero when estimating the complier effect because they are not part of the estimand. This is not a flaw. It is a recognition that average effects across irrelevant strata can blur the very contrast we care about.
A powerful way to think about this is to distinguish targeted similarity from overall similarity.
- Overall similarity asks: Are these units generally alike?
- Targeted similarity asks: Are these units alike in the specific latent dimension that governs the causal estimand?
This distinction explains why propensity based stratification, matching, and synthetic control can succeed even when the broader model is imperfect. The method is not trying to recreate the entire data generating process. It is trying to recreate the slice of the process that matters for the counterfactual comparison.
In practical terms, this means that a good causal workflow should ask at least three questions:
- Which latent group is the estimand about?
- Which observable features plausibly proxy membership in that group?
- Which assumption must remain stable when we move from observed to unobserved potential outcomes?
If you cannot answer those three questions, your estimate may be precise but not meaningful.
A unified framework: causal estimation as latent alignment
Here is the deeper synthesis: both principal score methods and counterfactual policy methods can be viewed as latent alignment problems.
In the first case, we align control units to latent compliers. In the second, we align the treated unit to a synthetic or model based control trajectory. In both cases, we are not just predicting outcomes. We are aligning an observed pattern with an unobserved reference class under a stability assumption.
This suggests a useful framework with four steps:
1. Define the hidden object
What is the thing you cannot directly observe but need to approximate? It may be compliance type, a counterfactual trajectory, or the distribution of shocks.
2. Identify the stable bridge
What remains invariant across the observed and unobserved regimes? It may be conditional independence, pre period error structure, or the mapping from covariates to latent membership.
3. Estimate only where the bridge is credible
Use matching, weighting, regression adjustment, constrained fitting, or matrix completion only where the bridge plausibly holds. Do not ask the model to extrapolate beyond its domain of credibility.
4. Treat sensitivity as part of the estimate
If the estimate collapses when the principal ignorability assumption weakens slightly, or when the error invariance is violated, then that fragility is not a footnote. It is the central result.
This framework changes how we think about robustness. Robustness is not just about standard errors or placebo tests. It is about whether the latent alignment survives small perturbations in assumptions. If the answer is yes, the estimate is informative. If not, the estimate is a story about modeling convenience, not causal reality.
The best causal methods do not merely predict the unseen. They identify the narrow conditions under which the unseen can be trusted to resemble the seen.
Key Takeaways
- Stop asking only what the counterfactual outcome is. Ask which hidden subgroup or latent trajectory your estimand depends on.
- Prefer local credibility over global fit. A method that fits the pre period or the compliance model where it matters is often better than a more flexible model that drifts from the estimand.
- Name the invariance explicitly. Whether it is principal ignorability or stationary shocks, causal inference depends on a stability assumption that should be stated and stress tested.
- Use prediction as a bridge, not a destination. Predictive scores, synthetic controls, and constrained models are tools for latent alignment, not substitutes for identification.
- Treat the target group as part of the estimand. Sometimes never takers should be weighted to zero, sometimes treated units should be matched to a donor pool, because the question is about a specific causal world, not the entire population.
The real lesson: causality is a theory of resemblance under intervention
The deepest connection between these ideas is not about propensity scores or synthetic control specifically. It is about how we justify resemblance when the world has been split by an intervention.
In ordinary prediction, resemblance is cheap. Past data resemble future data because the same process generated both. In causal inference, resemblance is expensive. It must be earned by arguing that certain structures remain stable across treatment and control, before and after intervention, compliers and never takers, observed and unobserved worlds.
That is why causal work is always partly philosophical. It asks not just, “Can we fit the data?” but, “What exactly is supposed to stay the same when everything else changes?” The answer to that question is the real engine of identification.
So the next time you see a carefully weighted sample, a matched principal stratum, or a synthetic counterfactual trajectory, do not think first about the machinery. Think about the hidden promise it makes. It is promising that some deep feature of the world, compliance structure, shock distribution, covariate relationship, remains intact enough that the invisible can be approximated by the visible.
That promise is fragile. But when it holds, it turns data into something more than memory. It turns data into a disciplined way of reasoning about worlds that never existed.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣