A bit is one good question
entropy, cross-entropy, and KL divergence, worked out at a café counter
00What you'll have at the end
By the end of this page you will know what a bit actually measures, why the standard classification loss in PyTorch and TensorFlow is called cross-entropy, and what the KL divergence is the price of. You will be able to look at a loss value and say which part of it is unavoidable and which part is the model being wrong.
We are going to do all of it at a café counter, with a barista who guesses orders by asking yes-or-no questions. No information-theory words until the counting has earned them.
01 new here · a yes-or-no questionFour drinks, two questions
The café sells four drinks. The regulars walk in mid-phone-call: they will nod or shake their head at the barista, but they will not speak. So ordering is a game. The barista asks yes-or-no questions, the customer nods or shakes, until the drink is pinned down. Start with the simplest case, where every drink is equally likely.
One question can split four into two and two: "caffeine?" A nod leaves espresso and tea, a shake leaves chocolate and juice. A second question splits the remaining two. Every order takes exactly two questions, no matter which drink it is.
Count the halvings. Four possibilities, halve once to get two, halve again to get one. Eight drinks would take three halvings, sixteen would take four.
A question that cuts the possibilities in half is worth one bit. Four equally likely drinks is two bits of uncertainty, because it takes two halvings to pin the order down. Bit is short for binary digit: a nod or a shake, written as 1 or 0, is the smallest answer there is, and a run of them is a binary string. That is all the binary format is.
The logarithm base two is just "how many times do I halve this to get to one". That is the only thing it ever means on this page.
02 new here · unequal sharesAsk about the likely drink first
Real customers are not equally likely. At this café, half of all orders are espresso, a quarter are tea, and chocolate and juice are one in eight each.
A smarter barista asks "espresso?" first. Half the time the answer is yes and the game is over after one question. Otherwise "tea?", then "chocolate?", and a no on that last one means juice.
Espresso happens one time in two: one halving. Tea, one in four: two halvings. Chocolate, one in eight: three halvings. The number of questions a drink needs is how many times you halve to get down to its share.
The number of questions a particular order takes is that order's surprisal, or information. A drink ordered half the time carries one bit: you were not very surprised. A drink ordered one time in eight carries three bits: you had to work to find it.
\(1/p_i\) is "one share in how many". The minus sign is not a convention; it is \(\log(1/x) = -\log x\), and \(1/p\) is always at least one, so the question count is never negative.
Suppose everyone ordered espresso, every single time. Its share is 1, one share in one, and the barista never needs to ask: zero questions, zero surprise. That is the top of the scale, and it is the other way to see why the count is never negative. A drink ordered every other time is the espresso line from the table: one question. And a drink ordered once in 1024 orders takes 10 questions. Drag that number. Each doubling of the orders between sightings adds one question, and the orders can keep doubling, so there is no ceiling: the rarer the drink, the more surprised you are, without limit. A drink that is never ordered at all has no share and no question count; it is not on the menu.
03 new here · the average over a dayHow many questions, on average?
Over a long day, how many questions does the smart barista ask per order? Weight each drink's question count by how often it happens, and add.
The smart scheme is a quarter of a question cheaper per order than asking two every time. And no scheme can beat 1.75 for this café: the question counts already match the shares exactly, one halving per factor of two.
The average number of questions with the best possible question order is the café's entropy. It measures how unpredictable the orders are. A café where everyone orders espresso has entropy zero: no questions needed. Four equally likely drinks has entropy two, which is step 01 again.
Average surprisal: each drink's question count, weighted by how often you meet it. The unit is bits because every question was one halving.
"On average" needs a day to average over. Here are the first 60 orders of one day, with the questions each one took, and the running average settling toward 1.75. Drag the number to watch a longer or shorter day.
04 new here · the answers written downThe ticket is a transcript
The barista stops asking out loud and just writes the answers on the ticket: 1 for yes, 0 for no. Espresso's ticket reads 1. Tea reads 01. Chocolate 001, juice 000. Ticket length equals question count. The list of four, one ticket per drink, is this café's shorthand.
A whole morning fits on one strip with no gaps, and it reads back only one way. Try it:
A shorthand like this is a code: one ticket per drink, and a strip of tickets reads back only one way. The whole game is information theory, Shannon's 1948 subject: a shorthand is an encoding, and the entropy of step 03 is the shortest average ticket any encoding can reach, which is why no scheme beat 1.75. When every share is a power of two, the question tree reaches it exactly.
05 new here · a ticket length read as a betRead the shorthand backwards
Now read the shorthand backwards. A shorthand is a question order written down, so a short ticket for a drink means you planned to ask about it first. Giving espresso a one-letter ticket only makes sense if espresso is about half the orders. Giving juice three letters is a bet that juice is about one in eight. The shorthand is a set of bets about the customers.
Every shorthand implies a guess about how often each drink is ordered: a ticket of length \(\ell\) is a bet on a share of \(1/2^{\ell}\). Short tickets for the drinks you expect, long tickets for the ones you do not.
\(q_i\) is the share the code is betting on. For this shorthand the bets add up to one and match the café, which is why its average length equals the entropy.
06 new here · a shorthand from another caféThe wrong shorthand
A new barista arrives from a juice bar across town, where juice is half the orders and espresso is one in eight, and brings that shorthand along. Juice is 1, chocolate 01, tea 001, espresso 000. Use it here, where the customers have not changed, and count the letters per order.
The average ticket length when the shorthand was written for one café and the orders come from another is the cross-entropy between them: the cost of describing the customers you have using the bets you made. When the bets match the customers, it equals the entropy. It can never be less.
Same weighted sum as the entropy, with one change: the lengths come from the shorthand \(q\), the weights from the customers \(p\). The first line is the juice bar's shorthand from this step: 2.625 letters per order, against the 1.75 the café's own costs. Pick another shorthand above and its line appears beneath. Set \(q = p\) and the sum collapses to the entropy.
07 new here · the gap between two shorthandsThe surcharge
With the juice bar's shorthand the café pays 2.625 letters per order. With its own, 1.75. The difference, 0.875 of a letter on every single ticket, is pure waste: the price of believing the wrong café. See where it comes from, drink by drink.
Notice that two of the four terms are negative. The wrong shorthand does give juice a shorter ticket than it deserves. But juice is rare, so that saving is small, and espresso is common, so its two wasted letters cost a full letter per order on average. The total is positive, and it always will be: bets that match the customers are the only bets with no surcharge.
The surcharge, cross-entropy minus entropy, is the KL divergence from the customers \(p\) to the bets \(q\). It is zero only when the bets are exactly right, never negative.
Line by line. The surcharge is the wrong shorthand's average minus the right one's (steps 06 and 03). Both averages weight drink \(i\) by the same \(p_i\), so the two sums can be paired term by term under one sum. Each bracket is then a drink's ticket length in the wrong shorthand minus its length in the right one: the extra letters in the figure above. A difference of two logarithms is the logarithm of the ratio, \(\log a - \log b = \log\frac{a}{b}\), and \(\frac{1/q}{1/p}\) is \(\frac{p}{q}\). The formula at the bottom and the counting at the top are the same four numbers.
Each term is a drink's share times the extra letters its ticket carries, and extra can be negative for a drink the wrong shorthand happens to favour. The sum cannot be: it is zero when \(q = p\) and positive otherwise, which is the claim in the box above, not something these six lines prove.
Aside: a shorthand with no ticket for juice
Suppose the visiting shorthand simply has no ticket for juice, a bet that juice never happens. The first juice order cannot be written down at all: the surcharge is infinite. Betting zero on something that can happen is the mistake with an infinite price, which is also why training pushes a model to spread some probability over everything it has ever seen.
08 new here · which café holds the shorthandThe other direction
The surcharge has a direction: who holds the shorthand, and whose customers walk in.
The three shorthands so far are three points. A model is free to bet any shares it likes, so here is the surcharge over every bet at once. Across: the share bet on espresso. Down: the share bet on tea. The remaining share is split evenly between chocolate and juice. Hover a cell to read that shorthand's bets and its surcharge; the dot marks this café's own bets, where the surcharge is zero.
KL divergence is not symmetric. The surcharge this café pays for the juice bar's shorthand is a different question from the surcharge the juice bar would pay for this café's, and in general they are different numbers.
In the numbers above the two cafés happen to be mirror images, so both directions come out at 0.875; with any other pair of cafés they differ.
09 new here · a customer who is certainA customer with no surprises
another example · One regular orders espresso every single time. For her, the café's "shares" are one for espresso and zero for everything else. How many questions does the smart barista need? None. Her entropy is zero.
Now hand her order to a barista with some shorthand. The ticket length is whatever that shorthand assigned to espresso, and since nothing about her is uncertain, every letter on it is surcharge.
When the customers are certain, cross-entropy and KL divergence coincide, and both reduce to one number: how many letters the shorthand spent on the thing that actually happened. A one-letter ticket means the shorthand bet half on it. A three-letter ticket means the shorthand bet one in eight. Shorter is better, and the only way to shorten it is to bet more on the right answer.
With one share equal to one and the rest zero, the weighted sum has a single term.
10 new here · the nouns onlyThe same game, told again
second story · same pictureanother example · The pager goes off at three in the morning. The incident log says what it usually is: 🚀 a bad deploy half the time, 🔌 a dependency down one time in four, ⚙️ a config change one in eight, 💾 a full disk and 🔒 an expired certificate one in sixteen each. The on-call engineer plays the barista's game: yes-or-no checks, likeliest cause first.
New nouns, same picture: a cause's check count is how many halvings reach its share, and the average is the on-call rota's entropy, 1.875 checks per page. "Look at the deploy first" is not folklore; it is the one-letter ticket. An expired certificate is a four-check night, which is why it feels like one. Nothing about coffee was ever the point.
11 new here · the PyTorch namesThe same counter, back in machine learning
Everything above was tickets. Here is the dictionary.
- The model is the shorthand. A classifier's softmax output is a set of bets \(q\) over the classes. The label is a certain customer, \(p\) one-hot. The loss is the length of the ticket the model assigned to the true class. \[ \mathcal{L} = -\log q_{y} = H(p, q) = D_{\mathrm{KL}}(p\,\|\,q) \quad\text{when } p \text{ is one-hot} \]
- Minimizing cross-entropy is minimizing KL. \(H(p,q) = H(p) + D_{\mathrm{KL}}(p\|q)\), and \(H(p)\) does not depend on the model. So training can only shrink the surcharge. With one-hot labels the floor is zero. With soft labels, as in distillation, or with real next-token data, the floor is the entropy of the data itself, and no model goes below it.
- Bits versus nats. PyTorch and NumPy take the natural log by default, so losses come out in nats, not bits. Divide by \(\ln 2\) to get bits. A language model's reported loss is cross-entropy in nats per token; its perplexity is \(e^{\text{loss}}\), the effective number of equally likely tokens, which is step 01 run backwards. \[ \text{bits} = \frac{\text{nats}}{\ln 2}, \qquad \text{perplexity} = e^{H(p,q)} \]
- Why the loss takes logits. Cross-entropy needs \(\log q\), and computing \(\log\) of a softmax that has already underflowed to zero gives minus infinity. So the function takes the pre-softmax scores and does \(\log\text{-softmax}\) in one stable step.
# = -log_softmax(logits)[y], mean over the batch loss = F.cross_entropy(logits, y) # D_KL(p || q); note the argument order: log q first, then p kl = F.kl_div(log_q, p, reduction="batchmean")
- KL as a ruler. "How much did this change move the model's predictions?" is a KL divergence between the before and after output distributions. It ignores the data's own entropy and asks only about the surcharge, which is what you want when you ablate a component, compare a compressed model with the original, or score how well a reconstruction preserves behaviour.
- The direction. Training minimizes \(D_{\mathrm{KL}}(p\|q)\), customers first. That penalizes betting zero on anything the data does, so the model spreads mass over everything it has seen. The reverse direction, bets first, penalizes betting on anything the data never does, and tends to pick one mode and commit. Variational methods use the reverse; the choice shapes what the fitted distribution looks like.
| Quantity | In symbols | At the café | The question it answers |
|---|---|---|---|
| Surprisal | \(-\log_2 p_i\) | questions for one drink | How unexpected was this one order? |
| Entropy | \(H(p) = -\sum_i p_i\log_2 p_i\) | average questions, best shorthand | How unpredictable are the customers? The floor. |
| Cross-entropy | \(H(p,q) = -\sum_i p_i\log_2 q_i\) | average letters, some shorthand | How many letters per order does shorthand \(q\) cost at a café whose customers are \(p\)? |
| KL divergence | \(D_{\mathrm{KL}}(p\|q) = H(p,q) - H(p)\) | the surcharge | How much of that cost is \(q\) being wrong about \(p\)? |
Loop closed. A bit is one good question. Entropy is how many you need on average. Cross-entropy is how many letters you actually spend with the shorthand you have. KL is the gap between the two: the letters wasted because the shorthand bet on the wrong customers. The floor belongs to the customers and no shorthand can lower it; the waste belongs to the shorthand, and rewriting the shorthand is all training ever does. The next time a loss curve flattens, ask whether it has hit the floor or whether the shorthand still has the wrong bets in it.
— end —
One of a series of short lessons by Shimin Zhang, written with an AI co-author. One idea each, derived before it is named, with room to breathe. Corrections and arguments welcome: [email protected].