Let count document rows and count their coordinates. Stacking them gives an matrix : the shape always lists rows before columns. Individual observation vectors will be written as columns when used in projection or outer-product formulas; their transposes supply the rows of a data matrix. A transpose exchanges rows and columns and is marked . Thus a column has row form .
For Figure 2, let be a example matrix and a column of weights, with entries through . Their product is a column. Adjacent mathematical symbols mean multiplication when their shapes permit it.
The actual inputs and result are
The second output entry is . Each row of has three entries and meets all three entries of . This is why the number of columns in the matrix must equal the number of entries in the input column. There are four rows to evaluate, so the output has four entries. A negative weight is allowed: it subtracts a contribution rather than adding one.
The definition of is “dot each row of with ”, but the picture to keep is the other one (Figure 2): is a weighted sum of the columns of , with the entries of as the weights. If the columns of are directions, is the point you reach by going of the way along the first, along the second, and so on. That is what “coordinates” will mean in §7: the coordinates of a story are the weights that best rebuild it from the basis directions.
In this example, take the first column once, the second column twice, and subtract the third column once. The second entries combine as , just as in the row calculation. Both descriptions perform the same multiplications; they group them differently. We will use row products to measure coordinates and column combinations to rebuild observations.
For Figure 3, take a new example , let be , and call their product . Indices and select a row and column: entry is the dot product of row of with column of . Shapes must agree in the middle: gives . Each column of is a separate input vector; multiplying by a matrix carries out both matrix–vector calculations together.
Here are the inputs, with the second row and second column providing a short calculation:
The entry is . The four intermediate products correspond to the common inner dimension, 4. The two outer dimensions say that there will be three output rows and two output columns. Matrix multiplication is not entry-by-entry multiplication, and reversing the order need not give the same shape or result.
Two special cases recur. If has rows and columns, has shape . A row of is a column of , so each output entry compares two columns of by their dot product. This is called a Gram matrix, and it is the raw material of covariance. In contrast, compares pairs of observation rows and has shape .
For the other special case, let count selected, mutually perpendicular unit directions and let contain them as columns. Then has shape . Entry asks how much observation points along direction . This table contains every document’s coordinates on those directions: the archive projected onto them. We will derive the reconstruction corresponding to those coordinates rather than assuming that taking a projection preserves everything.
flips rows and columns. A square matrix has the same number of rows and columns. Call such a matrix ; it is symmetric when . Real symmetric matrices, including covariances, admit a complete set of real, mutually perpendicular eigenvector directions, defined in the next chapter.
Two vectors are perpendicular when their dot product is zero. Unit-length perpendicular columns are orthonormal: “ortho” refers to perpendicularity and “normal” to unit length. The identity matrix has ones on its diagonal and zeros elsewhere; multiplying by it changes nothing. A matrix with orthonormal columns satisfies , with a identity. Each diagonal entry in this product is a column’s squared length, so it is 1. Each off-diagonal entry compares two different columns, so it is 0.
Directions are independent when none can be formed as a weighted combination of the others. A basis is an independent collection spanning the space under discussion: its combinations reach every point in that space. The dimension counts the number of basis directions. A matrix’s rank counts the independent directions among its columns, equivalently among its rows. Listing a direction twice does not give us an extra independent direction.
If , the matrix is square and orthogonal. It has enough orthonormal columns to describe the whole space. Multiplication by or rotates or reflects coordinates and preserves lengths and angles. It also satisfies , so projecting and rebuilding recovers every input exactly.
If , the map keeps only selected coordinates of an observation column . Rebuilding gives , its orthogonal projection onto the retained subspace. Now need not be the identity: the missing directions cannot be recovered merely by changing back to the original number of coordinates.
For the notebook’s example, take and let contain the first two coordinate directions in three dimensions. The superscript makes the displayed list a column. Projection gives and rebuilding gives . The squared lengths are before projection and afterward. Nine units of squared length lie in the discarded coordinate. The rectangular matrix obeys , but the reconstruction operator is the diagonal matrix with entries 1, 1, 0, not a three-dimensional identity.
A quarter-turn in the retained plane maps to . The squared length remains . This separates two operations that the later pipeline uses for different purposes: retaining a subspace loses information; rotating coordinates within it does not.