When You Cannot Randomize, Learn to Borrow Strength Carefully

Nan Wang

Hatched by Nan Wang

May 13, 2026

10 min read

76%

0

The real problem is not missing data. It is missing counterfactuals.

What do a clinical trial with changing treatment assignments and a panel of economic data with gaps in the matrix have in common? At first glance, almost nothing. One is about saving lives under strict ethical and operational constraints. The other is about reconstructing a world of firms, countries, or households from incomplete observations. Yet both face the same uncomfortable truth: you rarely get to observe the world you need in the form you want.

That is the deeper problem. In medicine, you cannot always keep patients on a fixed protocol just to preserve purity of design. In observational panel data, you cannot simply wait for every unit to be observed at every time point. In both cases, the challenge is not merely missing information. The challenge is that the missing information is the very thing needed to answer the causal question.

This creates a powerful tension. The more you try to protect inference by freezing the design, the more you may weaken the study’s practicality, ethics, or efficiency. But the more you adapt to reality, the more you risk distorting the very comparisons that make causality possible. The central question becomes: how do you preserve causal credibility while allowing the data generating process to remain imperfect, adaptive, and incomplete?


The hidden similarity between adaptive trials and matrix completion

A useful way to think about both problems is this: they are not primarily about prediction. They are about recovering a latent structure under structured uncertainty.

In a clinical trial, adaptive design changes how participants are assigned as evidence accumulates. The design may drop inferior arms, enrich promising subgroups, adjust allocation probabilities, or modify sample size. That flexibility is valuable because it can make trials faster, safer, and more efficient. But every adaptation creates a new possibility for bias, because the observed data are now shaped by earlier outcomes.

In causal panel data, matrix completion tries to infer missing entries by exploiting low-rank structure. The idea is that countries, firms, or individuals may follow a few common latent factors, so the observed panel can be used to reconstruct the unobserved parts. This is also a form of borrowing strength: from the patterns you can see, infer the structure you cannot directly observe.

The connection is deeper than a shared taste for statistical ingenuity. Both methods depend on a disciplined hope that the observed world is informative about the unobserved world because the system has hidden regularities. In one case, the hidden regularity is a causal response surface across treatment options and subgroups. In the other, it is a latent factor structure across units and time. The principle is the same: if the world has enough structure, you can tolerate incompleteness without surrendering inference.

The core wager in both settings is not that data will be complete. It is that data will be structured enough to be useful.

That wager is powerful, but it is dangerous. When it works, it lets us do science under real constraints. When it fails, it can produce elegant nonsense.


The paradox of flexibility: better decisions can make inference harder

Adaptive designs seem almost too good to be true. Why keep assigning patients evenly to inferior treatments when early results suggest a better option? Why not learn as you go and improve the trial while it is still running? The answer is that adaptation creates a new statistical debt: the more you let the accumulating data influence future data, the more your observations become entangled with your own decisions.

This is not a flaw unique to medicine. It is the same logic that haunts every attempt to recover a causal effect from incomplete panels. If missingness is random, life is easy. But if the gaps are systematic, the observed matrix is no longer a neutral sample of the whole. The missingness itself may reflect treatment, selection, attrition, or strategic behavior. The data are now a product of the system you are trying to study.

Think of it like trying to reconstruct a melody from notes heard through a wall. If the notes are missing at random, you can infer the tune. But if the missing notes are systematically the loudest ones, or the ones played after a particular chord, then reconstruction becomes much harder. You are no longer hearing the song, only the song filtered through a process that chooses what to hide.

Adaptive trials have a similar complication. Once the design begins reacting to interim evidence, the randomization probabilities, the subgroup composition, and sometimes even the endpoint structure can shift. The trial becomes a living system. That responsiveness is ethically attractive, but analytically it means the trial is no longer a simple static comparison. Inference must now account for the path that generated the data.

This is why good adaptive design is not just a matter of tweaking probabilities. It requires a design philosophy: know in advance what kinds of adaptation are allowed, how they will be monitored, and how they will be analyzed. The same is true in matrix completion for causal panels. You cannot simply throw a machine learning method at missing outcomes and call the result causal. You need assumptions about factor structure, treatment timing, intervention stability, and the plausibility of counterfactual reconstruction.

The best methods in both domains are not the most flexible ones. They are the ones that are flexible in a controlled way.


A new mental model: causal recovery as constrained borrowing of strength

Here is the synthesis: both adaptive clinical design and causal matrix completion are instances of constrained borrowing of strength.

That phrase matters. Borrowing strength means using information from other patients, other time periods, other units, or earlier stages of a trial to infer what is missing now. Constrained means you only do this under rules that preserve the meaning of the comparison.

This mental model helps unify several things that are often treated separately:

  1. Randomization is not the enemy of adaptation. It is the constraint that keeps adaptation honest.
  2. Low-rank structure is not just a computational trick. It is the assumption that makes reconstruction possible without overfitting.
  3. Reporting standards are not bureaucratic overhead. They are the mechanism that lets others see whether the borrowing of strength was disciplined or opportunistic.

Consider a simple analogy. Suppose you are trying to estimate how much a student would have learned from a new tutoring program if they had not received it. If you know the student’s prior scores, attendance, and peer group, you may be able to infer a plausible counterfactual. But only if you believe those features capture the important latent factors. Likewise, if a clinical trial allows adaptive randomization, you can still estimate treatment effects, but only if the adaptation rules are predefined and transparent enough to reconstruct the path of evidence.

Both domains therefore ask the same question in different language: How much structure do we need to impose before incomplete data become interpretable, but not so much that we mistake our assumptions for reality?

That is the sweet spot. Too little structure, and you are left with noise. Too much structure, and you force the world into your model.


Why causal inference lives or dies on design, not just estimation

There is a temptation in modern data science to believe that stronger algorithms can rescue weak designs. Adaptive trials and matrix completion both warn against that illusion.

A sophisticated estimator can recover hidden patterns only if those patterns are actually present and the data collection process has not broken them beyond repair. If a trial’s adaptive rules are opaque, or if the trial drifts without proper control, estimation becomes a post hoc patch on a fundamentally compromised design. If a panel is missing in a way that correlates with the treatment effect itself, matrix completion may reconstruct a beautiful matrix that is causally wrong.

This suggests a broader principle: in causal work, design determines what kind of truth is even available. Estimation can refine, stabilize, and compress uncertainty, but it cannot manufacture identifiability out of thin air.

A helpful way to see this is to imagine two layers:

  • The recovery layer, where the goal is to infer missing outcomes or assign probabilities in the face of incomplete observation.
  • The validity layer, where the goal is to ensure that the recovered quantities mean something causal rather than merely predictive.

Adaptive designs operate across both layers simultaneously. They change the recovery problem because future observations depend on interim data. They also threaten validity unless the adaptation rules are part of the inferential framework from the beginning. Matrix completion for causal panels faces the same two-layer challenge. It must recover the missing outcomes, but also justify why the recovered counterfactuals are meaningful rather than just numerically plausible.

This is why the best practice is not to ask, “Can we estimate it?” but rather, “What structure makes this estimate defensible?”

That shift in question is decisive.


The ethics of incomplete knowledge

There is also an ethical dimension that is easy to miss. Adaptive trials are often justified because they can reduce exposure to inferior treatments and make better use of participants’ contributions. Matrix completion methods, though more abstract, are similarly appealing because they extract more information from each observed unit, reducing waste in data already collected.

In both cases, the ethical logic is the same: when information is scarce or costly, using it efficiently is a moral as well as statistical imperative.

But efficiency is not the highest good. A trial that adapts too aggressively may overreact to early noise. A panel model that fills in too confidently may hide uncertainty and produce false certainty about policy effects. The ethical task is not simply to maximize information extraction. It is to respect the fragility of the conclusions that extracted information supports.

This leads to a less glamorous but more important ideal: epistemic humility by design.

That means pre-specifying adaptation rules, reporting deviations transparently, quantifying uncertainty honestly, and resisting the urge to present imputed counterfactuals as though they were observed facts. It means treating reconstructed data as a hypothesis about the unseen world, not a substitute for direct observation. And it means recognizing that some questions cannot be fully answered, only bounded more or less responsibly.

A medicine that adapts wisely and a causal estimate that borrows strength wisely are both acts of stewardship. They use limited data without pretending limitation is gone.


Key Takeaways

  1. Treat incomplete data as a design problem, not just an estimation problem. Ask what assumptions make the missing pieces recoverable before choosing a method.

  2. Borrow strength only under explicit constraints. Whether you are adapting a trial or completing a matrix, structure is what keeps flexibility from becoming bias.

  3. Separate recovery from validity. A good estimate of a missing value is not automatically a credible causal estimate.

  4. Predefine what can change. If a study adapts, decide in advance what may adapt, when, and according to which rules.

  5. Report uncertainty as part of the result, not a footnote. Reconstructed counterfactuals and adaptive comparisons are inherently conditional on assumptions and design choices.


Conclusion: the future of causal inference is disciplined improvisation

We usually tell ourselves that good science is about control. But these two methods point to a subtler truth: good science is often about disciplined improvisation. The world does not hand us perfect experiments or complete panels. It hands us partial observations, shifting incentives, ethical limits, and missing entries. The challenge is not to eliminate these conditions, but to build systems that can learn within them without lying to themselves.

That is why adaptive trials and matrix completion belong in the same intellectual conversation. Both are responses to the same reality: causality must often be inferred from fragments. Both succeed only when they respect the hidden structure of the system. And both fail when we confuse clever reconstruction with truth.

The deepest lesson is this: incomplete data do not merely create technical problems. They force us to decide what kinds of regularity we are willing to believe in.

Once you see that, causal inference stops being just a set of methods. It becomes a philosophy of knowledge under constraint, where the goal is not to know everything, but to know what can be known, responsibly, from the pieces we have.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣