How Does Cross-Entropy Connect LLMs and Compression?

773.9K views
•
July 16, 2026
by
3Blue1Brown
YouTube video player
How Does Cross-Entropy Connect LLMs and Compression?

TL;DR

Cross-entropy measures the average number of bits needed when data follows one probability distribution but is encoded using probabilities optimized for another. The same weighted negative-log formula appears in compression and language-model training, allowing pre-training to be reframed as learning a probability model that compresses language effectively rather than merely predicting the next token.

Transcript

There's a kind of wild paper from 2002 called Language Trees and Zipping, which shows how you can find structure between languages just by using very general file compression. For example, let's say I give you a bunch of text documents in many different languages, and your goal is to automatically cluster them by language. And actually, you can aim... Read More

Key Insights

  • Cross-entropy is the expected coding cost when symbols follow an actual distribution P but the encoding lengths are determined by another distribution Q. Its formula weights each negative base-2 logarithm of Q by the corresponding probability from P.
  • Information content is the negative base-2 logarithm of an event’s probability. It can be understood as the number of halvings needed to reach that probability, and it approximates the number of bits an optimal encoding should allocate to the event.
  • Message information is additive across individual symbols. Although a single symbol’s information content is often fractional and cannot always correspond to a simple fixed bit string, the summed information of a long message approximates the length of its optimal encoding.
  • A mismatched code can be measurably inefficient. In the robot example, using the original code after the instruction probabilities change produces an average length of 2.625 bits per instruction, which is the cross-entropy for that coding situation.
  • Cross-entropy is order-sensitive because the actual distribution supplies the weights while the coding distribution supplies the negative-log costs. Swapping their roles can change the result, as demonstrated by reversing even and 90-10 binary distributions.
  • An even coding distribution assigns one bit of information to each of two possible events. Because both costs are identical, applying that code to a 90-10 actual distribution still produces a cross-entropy of one bit per symbol.
  • A skewed coding distribution performs poorly when reality is even. When a code based on a 90-10 split is applied to equally likely events, the cross-entropy rises to about 1.74 bits because the supposedly rare event occurs much more often.
  • Compression can expose relationships between documents by measuring how well a snippet from one compresses after a compressor has processed another. Smaller added size suggests more similar patterns, enabling language clustering, recovery of shared lineage, and authorship identification.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What is cross-entropy in information theory?

Cross-entropy is the average coding cost produced when data follows one probability distribution but is encoded according to another. For every possible symbol, take its negative base-2 log probability under the coding distribution, then weight that value by how often the symbol actually occurs. The resulting sum represents the average number of bits used per symbol in the mismatched setting.

Q: How does cross-entropy relate to data compression?

Cross-entropy describes how efficiently a compression scheme optimized for one context performs in another context. A compressor effectively assigns shorter descriptions to patterns it considers likely and longer descriptions to patterns it considers unlikely. When the real data differs from those expectations, weighting the assigned description lengths by the real frequencies reveals the average number of bits required.

Q: How is information content calculated for an event?

The information content of an event is calculated as the negative base-2 logarithm of its probability. A highly probable event has low information content and should receive a short representation in an efficient code. A rare event has greater information content and requires a longer representation. The expression can be viewed as counting how many halvings reach the event’s probability.

Q: Why can information content be a fractional number of bits?

Probabilities do not usually produce whole numbers when their negative base-2 logarithms are calculated. That means an optimal encoding cannot always assign each individual symbol a simple bit string whose length exactly equals its information content. Instead, the information values add across a full message, and the total approximates the number of bits required by an efficient encoding of that message.

Q: Why does the order of distributions matter in cross-entropy?

The two distributions play different roles in the calculation. The actual distribution determines how frequently each symbol contributes to the average, while the coding distribution determines the negative-log cost assigned to each symbol. Swapping them changes both roles and can change the answer. In the binary examples, one ordering gives one bit, while the reverse gives about 1.74 bits.

Q: How can file compression identify related languages?

A compressor can process document A, then compress a small snippet of document B appended to it. The increase in compressed size indicates how well patterns from B fit a compressor largely optimized for A. Similar linguistic patterns tend to produce a smaller increase. A distance based on this co-compression behavior was sufficient to recover a tree representing shared language lineage.

Q: What does the robot instruction example demonstrate?

The robot example shows how symbol probabilities determine efficient code lengths and how changed probabilities create inefficiency. The original code gives one bit to up, two bits to down, and three bits each to left and right. After the distribution shifts toward right while the code remains fixed, the weighted average becomes 2.625 bits per instruction, illustrating cross-entropy concretely.

Q: How does cross-entropy connect LLM training with compression?

Cross-entropy appears both when measuring the cost of encoding data with mismatched probabilities and when training language models. This shared formula supports a compression-based interpretation of pre-training: the model learns probability assignments for language, and those probabilities determine potential coding lengths. Training can therefore be viewed as improving the model’s ability to compress language, not only to predict the next token.

Summary & Key Takeaways

  • Cross-entropy arises naturally when a code designed for one probability distribution is used on symbols generated by another. Each symbol receives a cost equal to the negative base-2 logarithm of its probability under the coding distribution. Weighting those costs by the actual distribution gives the average number of bits required per symbol.

  • Compression can reveal meaningful structure without built-in linguistic knowledge. By appending part of one document to another and measuring the compressed-size increase, researchers estimated how well patterns learned from the first document matched the second. This general co-compression method recovered language relationships and could also help identify a document’s author.

  • The same mathematical form used to evaluate mismatched compression appears as the loss function in language-model pre-training and distillation. Understanding its origin in coding theory changes the interpretation of training: a model is not only learning to predict the next token, but also learning probabilities that can encode language more efficiently.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from 3Blue1Brown 📚