Reading uncertainty · one café ☕, one shorthand 🎟️

A bit is one good question

entropy, cross-entropy, and KL divergence, worked out at a café counter

00What you'll have at the end

By the end of this page you will know what a bit actually measures, why the standard classification loss in PyTorch and TensorFlow is called cross-entropy, and what the KL divergence is the price of. You will be able to look at a loss value and say which part of it is unavoidable and which part is the model being wrong.

We are going to do all of it at a café counter, with a barista who guesses orders by asking yes-or-no questions. No information-theory words until the counting has earned them.

01 new here · a yes-or-no questionFour drinks, two questions

The café sells four drinks. The regulars walk in mid-phone-call: they will nod or shake their head at the barista, but they will not speak. So ordering is a game. The barista asks yes-or-no questions, the customer nods or shakes, until the drink is pinned down. Start with the simplest case, where every drink is equally likely.

One question can split four into two and two: "caffeine?" A nod leaves espresso and tea, a shake leaves chocolate and juice. A second question splits the remaining two. Every order takes exactly two questions, no matter which drink it is.

Each question halves what is left.

Count the halvings. Four possibilities, halve once to get two, halve again to get one. Eight drinks would take three halvings, sixteen would take four.

A question that cuts the possibilities in half is worth one bit. Four equally likely drinks is two bits of uncertainty, because it takes two halvings to pin the order down. Bit is short for binary digit: a nod or a shake, written as 1 or 0, is the smallest answer there is, and a run of them is a binary string. That is all the binary format is.

in symbols
\[ \text{questions} = \log_2(\text{number of equally likely options}) \qquad\qquad \log_2 4 = 2, \quad \log_2 8 = 3 \]

The logarithm base two is just "how many times do I halve this to get to one". That is the only thing it ever means on this page.

02 new here · unequal sharesAsk about the likely drink first

Real customers are not equally likely. At this café, half of all orders are espresso, a quarter are tea, and chocolate and juice are one in eight each.

A smarter barista asks "espresso?" first. Half the time the answer is yes and the game is over after one question. Otherwise "tea?", then "chocolate?", and a no on that last one means juice.

The smart question order. Espresso takes one question, tea two, chocolate and juice three each.

Espresso happens one time in two: one halving. Tea, one in four: two halvings. Chocolate, one in eight: three halvings. The number of questions a drink needs is how many times you halve to get down to its share.

The number of questions a particular order takes is that order's surprisal, or information. A drink ordered half the time carries one bit: you were not very surprised. A drink ordered one time in eight carries three bits: you had to work to find it.

in symbols
\[ \text{questions for drink } i = \log_2\!\Big(\frac{1}{p_i}\Big) = -\log_2 p_i \qquad\qquad \log_2\tfrac{1}{1/8} = \log_2 8 = 3 \]

\(1/p_i\) is "one share in how many". The minus sign is not a convention; it is \(\log(1/x) = -\log x\), and \(1/p\) is always at least one, so the question count is never negative.

the same formula at the ends of the menu

Suppose everyone ordered espresso, every single time. Its share is 1, one share in one, and the barista never needs to ask: zero questions, zero surprise. That is the top of the scale, and it is the other way to see why the count is never negative. A drink ordered every other time is the espresso line from the table: one question. And a drink ordered once in 1024 orders takes 10 questions. Drag that number. Each doubling of the orders between sightings adds one question, and the orders can keep doubling, so there is no ceiling: the rarer the drink, the more surprised you are, without limit. A drink that is never ordered at all has no share and no question count; it is not on the menu.

03 new here · the average over a dayHow many questions, on average?

Over a long day, how many questions does the smart barista ask per order? Weight each drink's question count by how often it happens, and add.

Each drink's share of the average. The smart questions average 1.75 per order; two-questions-every-time averages 2.

The smart scheme is a quarter of a question cheaper per order than asking two every time. And no scheme can beat 1.75 for this café: the question counts already match the shares exactly, one halving per factor of two.

The average number of questions with the best possible question order is the café's entropy. It measures how unpredictable the orders are. A café where everyone orders espresso has entropy zero: no questions needed. Four equally likely drinks has entropy two, which is step 01 again.

in symbols
\[ H(p) = \sum_i p_i \log_2\!\frac{1}{p_i} = -\sum_i p_i \log_2 p_i \] \[ H = \tfrac12\cdot 1 + \tfrac14\cdot 2 + \tfrac18\cdot 3 + \tfrac18\cdot 3 = 1.75 \text{ bits} \]

Average surprisal: each drink's question count, weighted by how often you meet it. The unit is bits because every question was one halving.

"On average" needs a day to average over. Here are the first 60 orders of one day, with the questions each one took, and the running average settling toward 1.75. Drag the number to watch a longer or shorter day.

One dot per order (1, 2, or 3 questions). The line is the running average. The dashed rule is the entropy. A short day wanders; a long day settles.
predict, then check
If a fifth drink were ordered one time in sixteen, how many questions would it take?

04 new here · the answers written downThe ticket is a transcript

The barista stops asking out loud and just writes the answers on the ticket: 1 for yes, 0 for no. Espresso's ticket reads 1. Tea reads 01. Chocolate 001, juice 000. Ticket length equals question count. The list of four, one ticket per drink, is this café's shorthand.

A whole morning fits on one strip with no gaps, and it reads back only one way. Try it:

Decoding a strip. No ticket is the start of another, so each one ends where the next begins.

A shorthand like this is a code: one ticket per drink, and a strip of tickets reads back only one way. The whole game is information theory, Shannon's 1948 subject: a shorthand is an encoding, and the entropy of step 03 is the shortest average ticket any encoding can reach, which is why no scheme beat 1.75. When every share is a power of two, the question tree reaches it exactly.

05 new here · a ticket length read as a betRead the shorthand backwards

Now read the shorthand backwards. A shorthand is a question order written down, so a short ticket for a drink means you planned to ask about it first. Giving espresso a one-letter ticket only makes sense if espresso is about half the orders. Giving juice three letters is a bet that juice is about one in eight. The shorthand is a set of bets about the customers.

This café's shorthand read backwards: each ticket's length, and the share it is betting on. The bets add up to one and match the customers.

Every shorthand implies a guess about how often each drink is ordered: a ticket of length \(\ell\) is a bet on a share of \(1/2^{\ell}\). Short tickets for the drinks you expect, long tickets for the ones you do not.

in symbols
\[ \ell_i = -\log_2 q_i \quad\Longleftrightarrow\quad q_i = 2^{-\ell_i} \qquad\qquad \tfrac12 + \tfrac14 + \tfrac18 + \tfrac18 = 1 \]

\(q_i\) is the share the code is betting on. For this shorthand the bets add up to one and match the café, which is why its average length equals the entropy.

06 new here · a shorthand from another caféThe wrong shorthand

A new barista arrives from a juice bar across town, where juice is half the orders and espresso is one in eight, and brings that shorthand along. Juice is 1, chocolate 01, tea 001, espresso 000. Use it here, where the customers have not changed, and count the letters per order.

Letters per order at this café, for each shorthand, counted drink by drink the same way as the entropy in step 03. Pick a shorthand above.

The average ticket length when the shorthand was written for one café and the orders come from another is the cross-entropy between them: the cost of describing the customers you have using the bets you made. When the bets match the customers, it equals the entropy. It can never be less.

in symbols
\[ H(p, q) = \sum_i p_i \log_2\!\frac{1}{q_i} = \sum_i p_i\, \ell_i^{(q)} \]

Same weighted sum as the entropy, with one change: the lengths come from the shorthand \(q\), the weights from the customers \(p\). The first line is the juice bar's shorthand from this step: 2.625 letters per order, against the 1.75 the café's own costs. Pick another shorthand above and its line appears beneath. Set \(q = p\) and the sum collapses to the entropy.

predict, then check
One more shorthand, from a tea house where tea is half the orders and espresso is the rarest: tea 1, chocolate 01, juice 001, espresso 000. Before counting anything, used at this café: against the two-letters-each shorthand's 2 letters per order, does it cost more or less? Against the juice bar's 2.625? Then the number:

07 new here · the gap between two shorthandsThe surcharge

With the juice bar's shorthand the café pays 2.625 letters per order. With its own, 1.75. The difference, 0.875 of a letter on every single ticket, is pure waste: the price of believing the wrong café. See where it comes from, drink by drink.

Extra letters per order, by drink. Espresso and tea get tickets that are too long; chocolate and juice get tickets that are too short. The savings on the rare drinks never make up for the waste on the common ones.

Notice that two of the four terms are negative. The wrong shorthand does give juice a shorter ticket than it deserves. But juice is rare, so that saving is small, and espresso is common, so its two wasted letters cost a full letter per order on average. The total is positive, and it always will be: bets that match the customers are the only bets with no surcharge.

The surcharge, cross-entropy minus entropy, is the KL divergence from the customers \(p\) to the bets \(q\). It is zero only when the bets are exactly right, never negative.

in symbols · from the surcharge to the formula
\[ \begin{aligned} D_{\mathrm{KL}}(p \,\|\, q) &= H(p,q) - H(p) \\[2pt] &= \underbrace{\sum_i p_i \log_2\frac{1}{q_i}}_{H(p,q),\ \text{step 06}} \;-\; \underbrace{\sum_i p_i \log_2\frac{1}{p_i}}_{H(p),\ \text{step 03}} \\[6pt] &= \sum_i p_i \Big( \log_2\frac{1}{q_i} - \log_2\frac{1}{p_i} \Big) \\[2pt] &= \sum_i p_i \big( \ell_i^{(q)} - \ell_i^{(p)} \big) \\[2pt] &= \sum_i p_i \log_2\frac{1/q_i}{1/p_i} \\[2pt] &= \sum_i p_i \log_2\frac{p_i}{q_i} \end{aligned} \]

Line by line. The surcharge is the wrong shorthand's average minus the right one's (steps 06 and 03). Both averages weight drink \(i\) by the same \(p_i\), so the two sums can be paired term by term under one sum. Each bracket is then a drink's ticket length in the wrong shorthand minus its length in the right one: the extra letters in the figure above. A difference of two logarithms is the logarithm of the ratio, \(\log a - \log b = \log\frac{a}{b}\), and \(\frac{1/q}{1/p}\) is \(\frac{p}{q}\). The formula at the bottom and the counting at the top are the same four numbers.

the instance, both ways

Each term is a drink's share times the extra letters its ticket carries, and extra can be negative for a drink the wrong shorthand happens to favour. The sum cannot be: it is zero when \(q = p\) and positive otherwise, which is the claim in the box above, not something these six lines prove.

predict, then check
The surcharge for the two-letters-each shorthand, in bits:
Aside: a shorthand with no ticket for juice

Suppose the visiting shorthand simply has no ticket for juice, a bet that juice never happens. The first juice order cannot be written down at all: the surcharge is infinite. Betting zero on something that can happen is the mistake with an infinite price, which is also why training pushes a model to spread some probability over everything it has ever seen.

08 new here · which café holds the shorthandThe other direction

The surcharge has a direction: who holds the shorthand, and whose customers walk in.

The three shorthands so far are three points. A model is free to bet any shares it likes, so here is the surcharge over every bet at once. Across: the share bet on espresso. Down: the share bet on tea. The remaining share is split evenly between chocolate and juice. Hover a cell to read that shorthand's bets and its surcharge; the dot marks this café's own bets, where the surcharge is zero.

Left: the surcharge the café pays for each possible shorthand, \(D(p\,\|\,q)\). Right: the surcharge a café with those bets as its true frequencies would pay for this café's shorthand, \(D(q\,\|\,p)\). Same colour scale. They are not the same picture: the left map climbs steeply toward the edges where a bet on a common drink goes to zero, the right one does not. That difference is the asymmetry.

KL divergence is not symmetric. The surcharge this café pays for the juice bar's shorthand is a different question from the surcharge the juice bar would pay for this café's, and in general they are different numbers.

In the numbers above the two cafés happen to be mirror images, so both directions come out at 0.875; with any other pair of cafés they differ.

09 new here · a customer who is certainA customer with no surprises

another example · One regular orders espresso every single time. For her, the café's "shares" are one for espresso and zero for everything else. How many questions does the smart barista need? None. Her entropy is zero.

Now hand her order to a barista with some shorthand. The ticket length is whatever that shorthand assigned to espresso, and since nothing about her is uncertain, every letter on it is surcharge.

One certain customer, three shorthands. Entropy is zero, so cross-entropy and surcharge are the same number: the length of espresso's ticket.

When the customers are certain, cross-entropy and KL divergence coincide, and both reduce to one number: how many letters the shorthand spent on the thing that actually happened. A one-letter ticket means the shorthand bet half on it. A three-letter ticket means the shorthand bet one in eight. Shorter is better, and the only way to shorten it is to bet more on the right answer.

in symbols
\[ p = (1, 0, 0, 0): \qquad H(p) = 0 \] \[ H(p, q) = D_{\mathrm{KL}}(p\,\|\,q) = -\log_2 q_{\text{espresso}} \]

With one share equal to one and the rest zero, the weighted sum has a single term.

10 new here · the nouns onlyThe same game, told again

second story · same picture

another example · The pager goes off at three in the morning. The incident log says what it usually is: 🚀 a bad deploy half the time, 🔌 a dependency down one time in four, ⚙️ a config change one in eight, 💾 a full disk and 🔒 an expired certificate one in sixteen each. The on-call engineer plays the barista's game: yes-or-no checks, likeliest cause first.

predict, then check
With the best order of checks, how many checks per incident on average? (Count halvings per cause, weight by its share, add.)
The same question tree, drawn by the same function, over causes instead of drinks. Each cause's check count is the number of halvings down to its share.

New nouns, same picture: a cause's check count is how many halvings reach its share, and the average is the on-call rota's entropy, 1.875 checks per page. "Look at the deploy first" is not folklore; it is the one-letter ticket. An expired certificate is a four-check night, which is why it feels like one. Nothing about coffee was ever the point.

11 new here · the PyTorch namesThe same counter, back in machine learning

Everything above was tickets. Here is the dictionary.

🏷️ label
the customers · p · usually certain
→
🤖 model
the shorthand · q · softmax output
→
📏 loss
letters spent · cross-entropy
QuantityIn symbolsAt the caféThe question it answers
Surprisal\(-\log_2 p_i\)questions for one drinkHow unexpected was this one order?
Entropy\(H(p) = -\sum_i p_i\log_2 p_i\)average questions, best shorthandHow unpredictable are the customers? The floor.
Cross-entropy\(H(p,q) = -\sum_i p_i\log_2 q_i\)average letters, some shorthandHow many letters per order does shorthand \(q\) cost at a café whose customers are \(p\)?
KL divergence\(D_{\mathrm{KL}}(p\|q) = H(p,q) - H(p)\)the surchargeHow much of that cost is \(q\) being wrong about \(p\)?

Loop closed. A bit is one good question. Entropy is how many you need on average. Cross-entropy is how many letters you actually spend with the shorthand you have. KL is the gap between the two: the letters wasted because the shorthand bet on the wrong customers. The floor belongs to the customers and no shorthand can lower it; the waste belongs to the shorthand, and rewriting the shorthand is all training ever does. The next time a loss curve flattens, ask whether it has hit the floor or whether the shorthand still has the wrong bets in it.

— end —

One of a series of short lessons by Shimin Zhang, written with an AI co-author. One idea each, derived before it is named, with room to breathe. Corrections and arguments welcome: [email protected].