Why Good Probabilities Need Good Counterfactuals

Nan Wang

Hatched by Nan Wang

Jun 03, 2026

9 min read

85%

0

The hidden problem with being right

What does it actually mean to be correct about a prediction? If you say there is a 70 percent chance of rain and it rains, did you succeed? If you say a policy will increase sales and sales rise, did you identify the cause? Both questions sound simple, yet both hide a deeper trap: observing an outcome is not the same as understanding the probability behind it, and an observed difference is not the same as a causal effect.

That is the common tension linking prediction and causation. One lives in the world of uncertainty, where a forecast can be honest even when it is wrong. The other lives in the world of counterfactuals, where the real question is not what happened, but what would have happened otherwise. Put together, they reveal a more demanding standard for knowledge: not just being accurate after the fact, but being calibrated before the fact, and causal after the fact.

This is why so many smart systems fail in practice. They confuse confidence with truth, correlation with intervention, and favorable outcomes with valid explanations. The result is a world full of models that seem sharp but are poorly calibrated, and analyses that seem persuasive but cannot survive the question, “What if we had done nothing?”


Probability is not a verdict, it is a distance

A probability forecast is often treated like a binary test: right or wrong, winner or loser. But the real question is subtler. A probability is not a claim that something will happen. It is a claim about how much weight to place on different possible worlds. That is why a Brier score is so revealing: it measures the distance in the probability domain.

If you predict 0.9 and the event happens, you are rewarded more than if you predicted 0.6. If you predict 0.9 and it does not happen, you are punished more severely than if you predicted 0.6. The logic is simple, but the implication is profound: good probabilistic thinking is not about being spectacularly right, it is about being appropriately uncertain.

This matters because many people unconsciously evaluate forecasts using only the final outcome. They ask whether the rain came, whether the customer converted, whether the stock rose. But a forecast lives before the outcome. Its quality depends on whether it tracked reality in advance, not whether it enjoyed a lucky coincidence afterward.

Consider two weather apps. App A says 90 percent rain every day. App B says 20 percent rain on most days, 80 percent on genuinely stormy ones. Over time, App B may feel less confident, but it will usually be the better forecaster. Why? Because calibration means your probabilities line up with frequencies. When you say 80 percent, about 8 out of 10 such situations should actually occur.

A probability forecast is not a promise. It is a commitment to honesty about uncertainty.

This is where the Brier score becomes more than a metric. It is a moral instrument for probabilistic thinking. It asks: are you matching your confidence to reality, or merely dressing up guesses in the language of certainty?


The counterfactual leap: from what happened to what would have happened

Prediction alone is not enough once decisions enter the picture. Suppose a company runs a promotion and sees sales rise. That looks like success, but it does not answer the real question: would sales have risen anyway? Causal reasoning begins where simple observation ends, because causality is fundamentally about potential outcomes.

The core idea is disarmingly elegant. For each unit, there is an observed outcome, and there is a potential outcome without treatment. But we never observe both at once. We see the treated world or the untreated world, never the same unit in both states simultaneously. That missing comparison is the heart of causal inference.

This is why questions such as average treatment effect and average treatment effect on the treated matter so much. They are attempts to summarize a world of unobservable alternatives. The treatment itself is never the whole story. The effect is defined against a baseline that stays hidden.

Imagine a tutoring program. A student attends and later improves. Did the tutoring cause the improvement? Maybe, but maybe the student was already on an upward trajectory, or the exam was easier, or peers also improved. To answer causally, you need the counterfactual: what would have happened to that student without tutoring? Since you cannot literally run time twice for the same person, you need design, comparison, and inference to approximate the missing world.

The deeper tension is that causal truth is not directly observed, it is reconstructed. That reconstruction can be elegant, but it can never be naive. It requires asking not just what changed, but what changed relative to what never happened.


The bridge between calibration and causation

At first glance, calibration and causality seem like separate disciplines. One evaluates forecasts, the other evaluates interventions. Yet they are connected by a single philosophical demand: do not confuse observed outcomes with underlying structure.

A calibrated probability says, “When I assign this number, it means something stable about the world.” A causal estimate says, “When I compare treated and untreated, I am approximating the effect of an intervention, not merely narrating a coincidence.” Both are attempts to make beliefs track reality in a disciplined way.

Here is the useful mental model: calibration is about the reliability of your uncertainty; causation is about the reliability of your comparisons.

That distinction clarifies why models can be misleading in two different ways. A predictive model may be highly accurate on average but badly calibrated at the extremes. A causal analysis may find a compelling difference between groups but fail to identify a true intervention effect. In both cases, the surface answer can look useful while the deeper question remains unanswered.

Think of a medical model that predicts a 30 percent chance of relapse. If, among patients assigned 30 percent, roughly 30 percent relapse, the model is calibrated. But suppose a treatment is then given to those high risk patients. If their outcomes improve, we still cannot say the treatment caused the improvement unless we know the counterfactual baseline. The same person can be probabilistically understood and causally misunderstood at the same time.

This intersection creates a powerful insight: decision making requires both probabilistic humility and causal imagination. Humility keeps you from overstating certainty. Imagination keeps you from mistaking correlation for intervention.


A practical framework: three questions before you trust a model

Whenever you face a model, forecast, or policy claim, ask three questions in sequence.

1. How well calibrated is the uncertainty?

If the model says 80 percent, does reality deliver about 80 percent over time? Calibration tests whether the number means what it says. Without this, a prediction is just theater.

A good practical test is to group similar predictions together and compare predicted probabilities to observed frequencies. If your 70 percent bucket behaves like a 70 percent bucket, you are at least speaking the language of probability honestly.

2. What is the counterfactual baseline?

If something improved, what would have happened without the intervention? If a group outperformed another, were the groups comparable? This question exposes the hidden assumptions behind causal claims.

A promotion may increase sales, but compared with what? A product may seem effective, but against which alternative? No causal claim is complete until it names the baseline world being subtracted away.

3. Are you measuring the right thing for the decision?

A well calibrated probability may be enough if you simply need to price risk. A causal estimate may be essential if you need to choose an intervention. Many failures happen because people use the wrong standard. They demand causal certainty from a forecasting tool, or they accept a causal story from a mere association.

This framework helps separate three different kinds of usefulness:

  • Forecasting usefulness: Can I estimate what will happen?
  • Causal usefulness: Can I estimate what would change if I act?
  • Decision usefulness: Can I choose better given costs, uncertainty, and alternatives?

The best analyses do all three, but very few do. Most systems collapse one into another and then wonder why they break.


Why overconfidence is the common enemy

Calibration and causality are united by a shared warning against overconfidence. Overconfidence in prediction produces numbers that are too sharp, too extreme, too certain. Overconfidence in causality produces narratives that are too tidy, too linear, too complete.

The world is messy in exactly the ways our minds dislike. Outcomes are noisy. Treatments are selective. Observed groups differ for reasons we do not fully see. That means a persuasive story can be wrong in both domains: it can be overconfident about probabilities and overconfident about effects.

This is especially dangerous because certainty feels useful. A manager prefers a confident forecast to a nuanced one. A policymaker prefers a clean causal story to a complicated identification problem. But the cost of false certainty accumulates quietly. Bad probabilities distort resource allocation. Bad causal stories distort intervention design.

A calibrated model may sometimes sound timid. A careful causal analysis may sometimes sound inconvenient. That is not a weakness. It is the price of respecting reality.

The world rarely rewards the person who sounds most certain. It rewards the person whose uncertainty is most informative.


What this means in practice

If you build models, run experiments, or make decisions based on data, the lesson is not merely technical. It is intellectual discipline.

Do not ask only whether a model predicts well. Ask whether its probabilities are calibrated, whether its confidence maps to reality, and whether it remains honest under different conditions. Do not ask only whether a treatment coincided with improvement. Ask what the untreated world would have looked like, how the comparison was formed, and whether the effect survives alternative explanations.

In organizations, this can change the way teams talk. Instead of saying, “The campaign worked,” say, “We observed an increase, and our best estimate of the incremental effect is X under these assumptions.” Instead of saying, “The model is 90 percent sure,” say, “Among cases like these, outcomes occur about 90 percent of the time.” These are not semantic tweaks. They are guards against self-deception.

A good rule of thumb: use probabilities to describe uncertainty, and causal estimates to describe change under action. When you blur the two, you get confident narratives with weak foundations.


Key Takeaways

  1. Treat probabilities as calibrated commitments, not verdicts. A forecast should match frequencies over time, not just sound impressive in the moment.
  2. Treat causal claims as counterfactual comparisons, not mere observations. An outcome after treatment is not proof of an effect.
  3. Separate prediction from intervention. A model that forecasts well may still be useless for choosing a policy, and a causal estimate may require a completely different method.
  4. Ask for the missing world. Every causal question depends on what would have happened without the treatment.
  5. Reject overconfidence in both domains. Precision without calibration and explanation without counterfactuals are both forms of intellectual error.

The real standard of understanding

The deepest connection between calibration and causality is that both force you to respect what is not directly visible. Calibration asks you to respect uncertainty as a measurable feature of belief. Causality asks you to respect the unobserved alternative as a necessary feature of explanation.

That means understanding is not just seeing what happened. It is learning to think in distributions and counterfactuals at the same time. The first keeps your confidence honest. The second keeps your stories honest.

And perhaps that is the real challenge of data, in life as in science: not to become more certain, but to become more precise about the kind of uncertainty you have, and the kind of change you are claiming. Once you do that, prediction becomes more trustworthy, causation becomes more disciplined, and decision making becomes less like storytelling and more like contact with reality.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣