How Can Probability Zero Still Mean Possible?

3.3M views
•
April 12, 2020
by
3Blue1Brown
YouTube video player
How Can Probability Zero Still Mean Possible?

TL;DR

A probability of zero does not make an exact outcome impossible when values vary continuously. Each individual value corresponds to an infinitely thin slice with zero area, while ranges can have positive probability measured by area under a probability density function. The full area remains one, preserving a valid distribution across all possible values.

Transcript

Imagine you have a weighted coin, so the probability of flipping heads might not be 50-50 exactly. It could be 20%, or maybe 90%, or 0%, or 31.41592%. The point is that you just don't know. But imagine that you flip this coin 10 different times, and 7 of those times it comes up heads. Do you think that the underlying weight of this coin is such tha... Read More

Key Insights

  • An unknown coin weight h is a real number from 0 to 1, representing outcomes from a coin that always lands tails to one that always lands heads, including every intermediate probability.
  • Assigning positive probability to every exact value in a continuous range creates a problem because there are uncountably infinitely many possible values, while assigning zero to each appears to make their total zero instead of one.
  • Ranges of values are the fundamental objects for continuous probabilities. Instead of asking only whether h equals precisely 0.7, a meaningful practical question asks whether h lies within an interval such as 0.6 to 0.8.
  • Probability is represented by the area of each histogram bar, not its height. As buckets become narrower, their probabilities decrease through shrinking widths while their heights stay roughly stable, preserving and refining the distribution's overall shape.
  • Probability density is probability per unit along the horizontal axis. Because a bar's probability equals width times height, its height expresses density, while areas over selected ranges express actual probabilities.
  • A probability density function assigns density to individual inputs. The probability that a continuous random variable lies between two values equals the area under the density curve between those values, and the total area under the full curve equals one.
  • An exact value in a continuous distribution has probability zero because it corresponds to an infinitely thin slice with zero area. All possible values together still have probability one because the full curve covers an area of one.
  • Measure theory unifies discrete, continuous, and mixed probability settings. It can handle a random number that equals 0 with 50% probability while otherwise following a positive continuous distribution shaped like half of a bell curve.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: Why can a possible outcome have probability zero?

A possible exact outcome can have probability zero when the random variable varies continuously. A single value corresponds to a range with width zero, so the area under the probability density curve over that slice is zero. This does not remove the value from the set of possibilities. The complete range contains all such values, and the total area across that range equals one.

Q: What is a probability density function?

A probability density function, or PDF, is a curve that assigns probability density to the possible values of a continuous random variable. Its height is not itself a probability. Instead, the probability that the variable lies between two values is the area under the curve across that interval. For a valid distribution, the total area under the entire curve must equal one.

Q: Why does probability use area instead of height for continuous values?

Probability uses area because increasingly fine histogram buckets become thinner as their ranges narrow. The probability inside each bucket decreases with its width, while the heights can remain roughly stable and approach a smooth density curve. If height represented probability directly, every height would approach zero and the limiting graph would lose all information. Area preserves the shape and the total probability of one.

Q: What does the height of a probability density curve mean?

The height of a probability density curve represents probability per unit in the horizontal direction, not the probability of an exact value. For a histogram bucket, probability equals width multiplied by height, so the height must describe density. A taller region indicates more probability concentrated per unit of horizontal width, but an interval's actual probability still requires calculating the area across that interval.

Q: Why can probabilities of exact continuous values not be added to get one?

The familiar rule of adding individual outcome probabilities works in finite and countably infinite settings, but it does not extend in the same way to an uncountable continuum. In a continuous distribution, every exact value may have probability zero, yet a range can have positive probability. Ranges are treated as fundamental, and their probabilities are determined by areas under a density curve rather than sums over individual points.

Q: How should you ask about an unknown coin's heads probability?

For a coin whose unknown heads probability is h, the useful question is not the probability that h equals one exact value such as 0.7. Instead, find the probability density function describing h after observing toss outcomes, then ask for the probability that h falls within a range. For example, the density can determine the probability that h lies between 0.6 and 0.8.

Q: How do histograms become probability density curves?

Start with buckets representing ranges of possible values, and let each bucket's area represent its probability. As the buckets become finer and narrower, the probability in each bucket approaches zero because its width shrinks. The heights remain roughly stable, so the histogram approaches a smooth curve. This limiting curve preserves the distribution's shape, and areas beneath it continue to represent probabilities.

Q: How does measure theory connect discrete and continuous probability?

Measure theory provides a rigorous framework for assigning probability to subsets of possible outcomes in a way that combines and distributes consistently. It unites finite, countably infinite, continuous, and mixed settings. For example, it smoothly handles a random number that equals 0 with 50% probability and otherwise takes a positive value according to a distribution shaped like half of a bell curve.

Summary & Key Takeaways

  • For a weighted coin with unknown heads probability h, asking for the probability that h equals exactly 0.7 creates an apparent paradox. Uncountably many individual values cannot all receive positive probabilities, yet assigning zero to each seems unable to produce the required total probability of one across every possible value.

  • The paradox is resolved by treating ranges, rather than individual values, as the fundamental objects carrying probability. As increasingly narrow buckets approach a smooth curve, each bucket's probability shrinks with its width while its height approaches probability density. Probability is therefore represented by area, and the total area must equal one.

  • A probability density function answers continuous probability questions through areas under its curve. An exact value has probability zero because its slice has zero width, while an interval can have positive probability. Measure theory provides a rigorous framework that unifies discrete, continuous, and mixed distributions under compatible rules for assigning probability to sets.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from 3Blue1Brown 📚