← First Pair Library

3 Matrices, pictured

3.1 A matrix times a vector

Let nn count document rows and pp count their coordinates. Stacking them gives an n×pn\times p matrix XX: the shape always lists rows before columns. Individual observation vectors will be written as columns when used in projection or outer-product formulas; their transposes supply the rows of a data matrix. A transpose exchanges rows and columns and is marked ⊤\top. Thus a column xx has row form x⊤x^{\top}.

For Figure 2, let AA be a 4×34\times3 example matrix and vv a 3×13\times1 column of weights, with entries v1v_1 through v3v_3. Their product AvAv is a 4×14\times1 column. Adjacent mathematical symbols mean multiplication when their shapes permit it.

The actual inputs and result are

A=(201130012111),v=(12−1),Av=(1702).A=\begin{pmatrix}2&0&1\\1&3&0\\0&1&2\\1&1&1\end{pmatrix},\qquad v=\begin{pmatrix}1\\2\\-1\end{pmatrix},\qquad Av=\begin{pmatrix}1\\7\\0\\2\end{pmatrix}.

The second output entry is 1×1+3×2+0×(−1)=71\times1+3\times2+0\times(-1)=7. Each row of AA has three entries and meets all three entries of vv. This is why the number of columns in the matrix must equal the number of entries in the input column. There are four rows to evaluate, so the output has four entries. A negative weight is allowed: it subtracts a contribution rather than adding one.

A matrix times a vector is a weighted sum of the matrix’s columns.

The definition of AvAv is “dot each row of AA with vv”, but the picture to keep is the other one (Figure 2): AvAv is a weighted sum of the columns of AA, with the entries of vv as the weights. If the columns of AA are directions, AvAv is the point you reach by going v1v_1 of the way along the first, v2v_2 along the second, and so on. That is what “coordinates” will mean in §7: the coordinates of a story are the weights that best rebuild it from the basis directions.

In this example, take the first column once, the second column twice, and subtract the third column once. The second entries combine as 1+2×3−0=71+2\times3-0=7, just as in the row calculation. Both descriptions perform the same multiplications; they group them differently. We will use row products to measure coordinates and column combinations to rebuild observations.

3.2 A matrix times a matrix

For Figure 3, take a new 3×43\times4 example AA, let BB be 4×24\times2, and call their product C=ABC=AB. Indices ii and jj select a row and column: entry CijC_{ij} is the dot product of row ii of AA with column jj of BB. Shapes must agree in the middle: (3×4)(4×2)(3\times4)(4\times2) gives 3×23\times2. Each column of BB is a separate input vector; multiplying by a matrix carries out both matrix–vector calculations together.

Here are the inputs, with the second row and second column providing a short calculation:

A=(120101122010),B=(10210311).A=\begin{pmatrix}1&2&0&1\\0&1&1&2\\2&0&1&0\end{pmatrix},\qquad B=\begin{pmatrix}1&0\\2&1\\0&3\\1&1\end{pmatrix}.

The entry C22C_{22} is 0×0+1×1+1×3+2×1=60\times0+1\times1+1\times3+2\times1=6. The four intermediate products correspond to the common inner dimension, 4. The two outer dimensions say that there will be three output rows and two output columns. Matrix multiplication is not entry-by-entry multiplication, and reversing the order need not give the same shape or result.

Two special cases recur. If XX has nn rows and pp columns, X⊤XX^{\top}X has shape p×pp\times p. A row of X⊤X^{\top} is a column of XX, so each output entry compares two columns of XX by their dot product. This is called a Gram matrix, and it is the raw material of covariance. In contrast, XX⊤XX^{\top} compares pairs of observation rows and has shape n×nn\times n.

For the other special case, let kk count selected, mutually perpendicular unit directions and let VV contain them as p×1p\times1 columns. Then XVXV has shape n×kn\times k. Entry (i,j)(i,j) asks how much observation ii points along direction jj. This table contains every document’s coordinates on those directions: the archive projected onto them. We will derive the reconstruction corresponding to those coordinates rather than assuming that taking a projection preserves everything.

Matrix multiplication: each entry of the product is a row of the left factor dotted with a column of the right one.

3.3 Transpose, symmetry, orthogonality

A⊤A^{\top} flips rows and columns. A square matrix has the same number of rows and columns. Call such a matrix SS; it is symmetric when S=S⊤S=S^{\top}. Real symmetric matrices, including covariances, admit a complete set of real, mutually perpendicular eigenvector directions, defined in the next chapter.

3.3.1 Perpendicular directions and what their products say

Two vectors are perpendicular when their dot product is zero. Unit-length perpendicular columns are orthonormal: “ortho” refers to perpendicularity and “normal” to unit length. The identity matrix II has ones on its diagonal and zeros elsewhere; multiplying by it changes nothing. A p×kp\times k matrix VV with orthonormal columns satisfies V⊤V=IV^{\top}V=I, with a k×kk\times k identity. Each diagonal entry in this product is a column’s squared length, so it is 1. Each off-diagonal entry compares two different columns, so it is 0.

Directions are independent when none can be formed as a weighted combination of the others. A basis is an independent collection spanning the space under discussion: its combinations reach every point in that space. The dimension counts the number of basis directions. A matrix’s rank counts the independent directions among its columns, equivalently among its rows. Listing a direction twice does not give us an extra independent direction.

3.3.2 A rotation keeps a whole space; a projection keeps part

If k=pk=p, the matrix VV is square and orthogonal. It has enough orthonormal columns to describe the whole space. Multiplication by VV or V⊤V^{\top} rotates or reflects coordinates and preserves lengths and angles. It also satisfies VV⊤=IVV^{\top}=I, so projecting and rebuilding recovers every input exactly.

If k<pk<p, the map V⊤xV^{\top}x keeps only selected coordinates of an observation column xx. Rebuilding gives VV⊤xVV^{\top}x, its orthogonal projection onto the retained subspace. Now VV⊤VV^{\top} need not be the identity: the missing directions cannot be recovered merely by changing back to the original number of coordinates.

For the notebook’s example, take x=(1,2,3)⊤x=(1,2,3)^{\top} and let VV contain the first two coordinate directions in three dimensions. The superscript ⊤\top makes the displayed list a column. Projection gives (1,2)⊤(1,2)^{\top} and rebuilding gives (1,2,0)⊤(1,2,0)^{\top}. The squared lengths are 1+4+9=141+4+9=14 before projection and 1+4=51+4=5 afterward. Nine units of squared length lie in the discarded coordinate. The rectangular matrix obeys V⊤V=IV^{\top}V=I, but the reconstruction operator is the diagonal matrix with entries 1, 1, 0, not a three-dimensional identity.

A quarter-turn in the retained plane maps (1,2)⊤(1,2)^{\top} to (−2,1)⊤(-2,1)^{\top}. The squared length remains 4+1=54+1=5. This separates two operations that the later pipeline uses for different purposes: retaining a subspace loses information; rotating coordinates within it does not.