A matrix is a mix of its columns
and five other honest ways to read a matrix multiplication
00What you'll have at the end
By the end of this page you will be able to read any matrix multiplication in whichever way makes the problem obvious. There are several, they are all the same arithmetic, and each one answers a different question.
You will also see why the rows and the columns of a matrix index two different kinds of thing, and why that is the first question to ask whenever a shape looks wrong.
We are going to do all of it at a fruit stand. No linear algebra words until the arithmetic has earned them.
01 new here Β· one basket, scaledThe stand sells baskets
A fruit stand sells baskets. Every basket of a given kind always has the same fruit in it.
Buy two of basket A. Count what you take home.
Doubling a basket doubles every number in it.
A number times a column scales every entry. Columns are written tall on purpose: a basket is a list that goes down.
02 new here Β· a second basket, added inTwo kinds of basket
The stand adds a second kind.
Buy 2 of basket A and 1 of basket B. Scale each basket by how many you bought, then add the baskets together. The numbers in that sentence are live: drag them, click them, or use the arrow keys.
You took home . Each basket contributed a scaled copy of itself, and the contributions were added.
Now write the two baskets side by side as columns, and the amounts you bought as a short list. What we just did, scaling each column by its amount and adding, is called multiplying the matrix by the vector. The amounts are the mix weights. The result is a mix of the columns. This is the column view.
Two pictures of the same purchase. On the left, each fruit's pile, with the part that came from each basket stacked in its own colour. On the right, the baskets drawn as arrows on a map where east is apples and north is pears: scale each arrow by its amount, lay them tip to tail, and the mix lands on the total.
The numbers follow the sentence above. On the right, the general statement: \(M_{:,j}\) is column \(j\), and the vector's entries \(v_j\) are the weights of the mix.
03 new here Β· the same sum, grouped by fruitCount it fruit by fruit instead
in other words Β· Same purchase. This time do not think in baskets at all. Walk the stand one fruit at a time and ask how many of that fruit you end up with.
Same . The arithmetic was regrouped, not changed.
Each answer is one row of the matrix multiplied entry by entry against the amounts, then summed. This is the row view, also written as a dot product: output i equals row i dotted with the vector.
Column view asks
What did each basket contribute? Scale the columns, add them up.
Row view asks
How many of each fruit did I end up with? One row, one answer.
Two questions, one multiplication.
The same sum as before, now indexed by the output \(i\) instead of grouped by the input \(j\). \(M_{i,:}\) is row \(i\).
04 new here Β· what the two directions indexRows and columns are different kinds of thing
Look at the shape of the matrix. Across the top are kinds of basket: what you buy. Down the side are kinds of fruit: what you get. In general words, columns are the inputs and rows are the outputs.
Nothing about apples lines up with basket A, and nothing should. A row index and a column index answer different questions. The number of rows is how many outputs there are. The number of columns is how many inputs.
Shapes are written rows \(\times\) columns, which reads outputs \(\times\) inputs. The inner numbers must agree: a \(3\times 2\) matrix takes a 2-vector and returns a 3-vector. The 2 is not a rule, it is the number of baskets you can choose from. You cannot walk up and ask for fruit directly; there is no 3-vector to hand over. The only thing you can choose is how many of each basket, and there are two baskets, so what you hand over is always a list of two numbers. When a shape error says the inner dimensions do not match, it is saying you handed over a list of the wrong length for the baskets on offer.
Run the stand backwards, asking which amounts would produce a given pile of fruit, and the column view is the one that makes the question natural: which mix of the columns lands there? That question gets its own page.
05 new here Β· several shopping lists at onceThree customers at once
If you learned to multiply matrices as "row times column", that rule is coming in step 08, and it will turn out to be one cell of everything built here.
A third kind of basket joins the stand.
π© Ana buys 2 A, 1 B, 0 C. π¨ Bo buys one of each. π§ Cy buys 0 A, 2 B, 1 C. Write each order as a column, side by side.
Do Ana's order the way you already know how. Then Bo's. Then Cy's. Put the three receipts side by side.
That is a matrix times a matrix. Column by column: each column of the result is the left matrix applied to one column of the right. Nothing new happened. You did three shopping trips and lined up the receipts.
Column \(j\) of the product is \(M\) applied to column \(j\) of \(N\).
06 new here Β· one fruit counted across everyoneRead the receipts row by row
Read the other way, each row of the result is one fruit counted across all customers. The apples row of the result is the apples recipe applied to the whole order table: a mix of the order table's rows, weighted by how many apples each basket holds. Rows of the left mix the rows of the right, just as columns of the right mix the columns of the left.
Column by column asks
What does one customer get? Bo's order mixes the baskets' columns. One column of the result per customer.
Row by row asks
How is one fruit spread across everyone? The apples recipe mixes the order table's rows. One row of the result per fruit.
Same product. Both ways fill the same table of receipts, cell for cell. They differ only in which direction you walk it.
Row \(i\) of the product is row \(i\) of \(M\) applied to all of \(N\). Beside step 05's formula, it is the same product indexed from the other side.
07 new here Β· the same table, grouped by basketOne basket's share of everything
another example Β· Here is a third way to count the same table. Pick basket A. Ana bought 2, Bo bought 1, Cy bought 0. Ask what basket A alone delivered to every customer.
Do the same for B and for C. Add the three tables.
Each table is one column of the left matrix against one row of the right. A table built that way is called an outer product, and a matrix product is the sum of outer products, one per basket type. This view answers a question the others cannot: how much of the whole result came from one source?
One outer product per basket \(k\): a column of the left against a row of the right. The second formula is the cell rule: entry \(i, j\) of the table is the \(i\)-th number of the column times the \(j\)-th number of the row, nothing summed. For basket A, the apples-for-Ana cell is \(2 \times 2 = 4\): apples per basket A, times baskets A that Ana bought. The \(\top\) lays the second list flat so a tall column times a flat row makes a table. Each outer product is a full table built from two lists, so it has rank one, and the product is a sum of rank-one tables.
08 new here Β· one row against one columnOne cell
in other words Β· Finally, the definition itself. How many π pears does π¨ Bo get? Take the pears row and Bo's column, multiply them entry by entry, add.
One cell of the result is one row of the left dotted with one column of the right. It is correct. What it does not show is how a whole row or a whole column of either matrix shapes the output, which is what the other views were for.
The definition. Every view on this page is this one sum with the terms grouped differently: by \(j\) for the column view, by \(i\) for the row view, by \(k\) for the outer products.
09 new here Β· the nouns onlyThe same stand, with a different menu
second story Β· same pictureanother example Β· Swap the nouns and nothing else. A bakery sells two dishes, and each dish always uses the same ingredients: a π cake takes 3 π₯ eggs, 1 π§ butter and 2 πΎ flour; a stack of π₯ pancakes takes 1 egg, 1 butter and 1 flour. Bake 2 cakes and 1 stack of pancakes.
Dishes are the columns, because a dish is what you choose; ingredients are the rows, because ingredients are what you end up using. The stand is gone and the picture did not notice.
10 new here Β· the PyTorch namesThe same stand, back in PyTorch
Everything above was fruit. Here is where you will meet each view next.
- The Linear layer.
nn.Linearstores its weight as[out, in], rows are outputs and columns are inputs exactly as in step 04, so row i is output neuron i's recipe across all inputs. That is the row view chosen as the storage layout, and it is why the forward pass isx @ W.T. If that transpose has ever felt backwards, this is why: PyTorch stored the row view and you were reading it in the column view. \[ \mathbf{y} = W\mathbf{x} + \mathbf{b}, \quad W \in \mathbb{R}^{\text{out}\times\text{in}} \qquad\qquad Y = X W^{\top} + \mathbf{b} \;\;\text{for a batch } X \in \mathbb{R}^{B\times\text{in}} \] - Any linear solve. Solving is the stand run backwards. So far you handed over basket counts and got a pile of fruit. Now a customer hands you the pile, say 5 apples, 5 pears, 4 grapes, and asks which basket counts produce it. The matrix is unchanged: its rows are still the fruits, one equation each, and its columns are still the baskets, one unknown each. The answer here is 2 of A, 1 of B, 1 of C, and the solver finds it:
amounts = torch.linalg.solve(M, pile) # M: fruits Γ baskets # pile: 3 fruit counts in, amounts: 3 basket counts out
Least squares and changes of basis are the same shape of question. Before you call a solver, say which list is the outputs you were given and which is the inputs you want back. - Attention. A head's output at a position is a mix of value vectors, weighted by the attention pattern. Column view. The pattern is the shopping list, and because the weights are non-negative and sum to one, the output always lands inside the hull of the values. \[ \mathbf{o}_i = \sum_{j} a_{ij}\,\mathbf{v}_j, \qquad a_{ij}\ge 0,\quad \sum_j a_{ij} = 1 \]
- Logits. The residual stream is a sum of pieces, one from every head and every MLP, and the unembedding is linear. So the logits are a sum of per-piece tables. Outer-product view. That is direct logit attribution, and it is why logits add and probabilities do not. \[ \boldsymbol{\ell} = W_U^{\top}\,\mathbf{r}, \qquad \mathbf{r} = \sum_{c} \mathbf{r}_c \;\;\Longrightarrow\;\; \boldsymbol{\ell} = \sum_{c} W_U^{\top}\,\mathbf{r}_c \]
| View | In symbols | The question it answers | Reach for it when |
|---|---|---|---|
| Column | \(M\mathbf{v}=\sum_j v_j M_{:,j}\) | What did each input contribute? | You want to see an output as a mix, or you are solving for the mix. |
| Row | \((M\mathbf{v})_i = M_{i,:}\cdot\mathbf{v}\) | How much of each output did I get? | You care about one output coordinate, or you are reading a stored weight matrix. |
| Column by column | \((MN)_{:,j} = M N_{:,j}\) | Many inputs at once? | Batches. Each column of the result is one example. |
| Row by row | \((MN)_{i,:} = M_{i,:} N\) | How is one output spread across all the examples? | Reading one output, or one feature, across a whole batch at once. |
| Sum of outer products | \(MN=\sum_k M_{:,k} N_{k,:}\) | How much came from one source? | Attribution. Decomposing a result by head, neuron, or feature. |
| One cell | \((MN)_{ij}=\sum_k M_{ik}N_{kj}\) | What is this single number? | Checking one entry by hand. |
The one-cell rule is the view that turns straight into code. Pick a row, pick a column, sum over the inner index: three nested loops, and the k loop is the dot product.
for i in range(rows): # one output row at a time
for j in range(cols): # one output column at a time
for k in range(inner): # the dot product of row i and column j
P[i][j] += M[i][k] * N[k][j]
In einsum notation every view of a matrix-vector product is the same string, "ij,j->i", and every view of a matrix-matrix product is "ik,kj->ij". The string names the index that gets summed away. The views are not different computations. They are different ways of reading which index is grouped with which.
Loop closed. You can read a product six ways, and you know that rows and columns of the same matrix index different things. The next time a shape looks wrong, ask which list is the outputs and which is the inputs before you touch a transpose.
— end —
One of a series of short lessons by Shimin Zhang, written with an AI co-author. One idea each, derived before it is named, with room to breathe. Corrections and arguments welcome: [email protected].