A scalar is a single number. A vector is an ordered list of numbers called coordinates. Our observation vectors are rows; when a matrix direction or another object is a column, we will say so. Let denote this example row:
It has three coordinates. The order matters: its first coordinate is 2, its second is 1, and its third is -1. In the running example the coordinates belong to three fixed news axes, called N1, N2, and N3. These names identify toy directions, not actual Anthropology topics.
Let be a second row of the same length. The index selects a coordinate: and are the numbers in position of the two rows. The summation sign means to add over all those positions, from 1 to 3 here. The dot product, written with a centered dot, multiplies corresponding coordinates and adds:
For these rows, the result is . An omitted multiplication sign between scalar entries also means multiply.
The Euclidean length, or norm, of is written . It is the square root of the sum of squared coordinates:
Dividing a nonzero vector by its length gives a unit vector. It keeps the direction while removing the size. The zero vector has no direction and cannot be normalized this way.
Single bars around a number mean its absolute value, or size without its sign. Thus . Single bars around a set will instead count its members; we will introduce those sets before using them. The symbol means approximately equal, as when values have been rounded.
Why square before adding? Simply adding the entries of gives zero even though the row is not zero. Squaring counts both directions as positive contributions. For a separate synthetic row , the squares are 9 and 16, their sum is 25, and the length is . This extends the right-triangle rule: perpendicular coordinate contributions add as squared lengths.
For the earlier row , the squared contributions are 4, 1, and 1. Its squared length is 6 and its length is approximately 2.44949. Dividing each coordinate by that same number gives approximately . Square these normalized coordinates and add: . That is the unit-length check; it does not require every individual coordinate to equal one.
Multiplying by 2 gives and doubles its length from 5 to 10. Normalizing either gives . Multiplication by -2 instead reverses the direction, giving unit row . Positive rescaling preserves direction; a negative rescaling reverses it. Neither normalization establishes that one observation has more supporting evidence than another.
Try it. Change the row to , then to . The first has unit direction . The second has length zero, and division by its length is undefined. A missing or directionless profile must not silently become an average profile.
A matrix is a rectangular table. Let count its rows and its columns; its shape is written . Put one three-coordinate person profile in each of four rows and the result is . A matrix entry uses its row index first and column index second. Diagonal entries have equal row and column indices; they run from the top left toward the bottom right.
Multiplication describes a weighted combination. Its numerical weights are coefficients; coefficients expressing directions in input coordinates are also called loadings. If is a row and is a loading matrix, then is a row. Its first number is the dot product of with the first column of . Its second number uses the second column. The adjacent dimensions must agree:
The transpose, written , exchanges rows and columns. It therefore has shape . Remembering shapes will prevent more mistakes than memorizing a complicated formula.
Use a separate synthetic transformation , with three input coordinates and two output coordinates:
For , the first output is . The second is . Thus . We multiplied a row by a table, producing a row. Each output column says how much of each input enters that output.
Transposing gives . Multiplying the returned row gives , which differs from . This deliberately arbitrary table does not have perpendicular unit columns, and its transpose is not an inverse. Shapes allow a multiplication; they do not promise what that multiplication will recover.
Suppose three measurements are 2, 3, and 7. Their mean is 4. Centering subtracts that mean, giving -2, -1, and 3. A negative centered value means below the chosen mean; it does not mean a negative amount of the original quantity.
For vector rows we average each column separately and subtract the resulting mean row. We will center twice for different purposes: first against a news background within a source-and-decade cell, then against the mean profile of the fitted population. These are two different comparisons, not a redundant repetition.
Variance measures spread by averaging squared deviations from a mean. A standard deviation is the nonnegative square root of a variance; it is zero when all values agree. Covariance averages the product of the centered values of two coordinates: a positive result means they tend to move together, a negative result means they tend to move oppositely. A covariance matrix collects those comparisons for every pair of coordinates. Its diagonal contains their individual variances. The fitted news model supplies its covariance convention; for the people calculation below we explicitly divide by the number of people.
The mean of 2, 3, and 7 is . The centered numbers are -2, -1, and 3; their sum is zero. Their squared deviations are 4, 1, and 9. For this descriptive example we divide by the observation count, three: variance is , approximately 4.6667. The standard deviation is , approximately 2.1602. Variance has squared coordinate units; standard deviation returns to the original units.
Pair these measurements with a second coordinate, 6, 4, and 2, in matching row order. The second mean is 4. Its deviations are 2, 0, and -2. The table keeps the pairings visible:
| Row | First deviation | Second deviation | First squared | Second squared | Product |
|---|---|---|---|---|---|
| 1 | -2 | 2 | 4 | 4 | -4 |
| 2 | -1 | 0 | 1 | 0 | 0 |
| 3 | 3 | -2 | 9 | 4 | -6 |
| Sum | 0 | 0 | 14 | 8 | -10 |
Divide the last three column totals by three. The resulting covariance matrix is
The repeated off-diagonal entry is the average paired product. Reversing the order of the two factors leaves each product unchanged, so the covariance matrix is symmetric. The negative sign records that a positive deviation in one coordinate tends to accompany a negative deviation in the other. It is not a causal claim.
Dividing centered columns by their standard deviations yields unit variances. Their covariance becomes the correlation , approximately -0.9449; standardization has not removed the association. Some estimators divide by one less than the count to estimate a population variance from a sample. Here the count denominator defines the spread of these declared rows; the people model below uses that same convention. Changing a denominator must be an explicit modeling choice.
Check the mechanism. If the second coordinate were constant, every second deviation would be zero. Its variance and every covariance involving it would be zero, and division by its standard deviation would be unavailable.
Perpendicular directions have dot product zero. Perpendicular unit directions are orthonormal. Let denote an identity matrix: it has ones on its diagonal and zeros elsewhere, so multiplying by it changes nothing. Its size matches the matrix product beside it. Let be a square matrix whose columns form a complete orthonormal system. Such a matrix is called orthogonal and satisfies . Multiplying by rotates or reflects coordinates and preserves all lengths and angles.
A tall matrix can also have orthonormal columns: . But when it has fewer columns than rows, multiplying by keeps only some directions. The length can decrease. The statement that an orthogonal change of coordinates preserves lengths applies to a complete rotation; a truncated projection preserves only the retained component. This distinction will explain the navigation duality.
Directions are linearly independent when none is a weighted combination of the others. A basis is a set of independent directions used to express coordinates. The subspace spanned by a collection of directions contains all their weighted combinations. The rank of a matrix counts its independent directions; keeping fewer directions is a low-rank representation.
The Eigen Times matrix pictures provide a longer visual introduction. When comparing a formula written with column vectors to this book, transpose it consistently rather than changing only one factor.
Check 2. A row has 60 news coordinates and its loading matrix has shape . What are the shapes of the forward result, the transpose, the combined matrix , and the returned row?
Take the synthetic row and retain only the horizontal unit direction . The dot product is . Rebuilding three units of that direction gives . The residual is what remains when we subtract the reconstruction from the original: .
The residual is a row, not initially a scalar error score. It tells us where the missing part points. Its squared length is . The reconstruction’s squared length is 9; the original’s is 25. Thus . Their dot product is zero, so the retained and residual parts are perpendicular. A later chapter applies this same calculation to the fitted people plane.
For comparison, multiplying by the complete orthogonal matrix gives . Its squared length is still 25. Multiplication by the transpose recovers exactly. A complete change of orientation has no discarded part; keeping only the horizontal coordinate did.