The mean is an average: add the observations and divide by their count. For the scalar values 1, 2, and 3, the mean is . Their deviations are the observed values minus that mean: −1, 0, and 1. Subtracting a common reference mean is called centering. The centered values say where each observation lies relative to that reference rather than relative to zero on the original scale.
Adding signed deviations gives zero here. That does not mean there is no spread; the negative and positive deviations cancel. Squaring avoids this cancellation. The squared deviations are 1, 0, and 1, whose mean is . This average squared deviation is the variance under the total-count convention used for the fitted covariance. Its nonnegative square root, , is the standard deviation. If observations are measured in some unit, variance is measured in that unit squared; standard deviation returns to the original unit.
Some statistical estimates divide the sum of squared deviations by one less than the count, to correct a bias when estimating a population variance from a sample. That is a different convention. The fitted covariance below divides by total weight. We will identify the reference and divisor when using a standard deviation elsewhere.
For more than one coordinate, calculate a mean for each coordinate. Consider the two observation rows and . Their mean row is , and their centered rows are and . The first-coordinate squared deviations average to 1; the second-coordinate squared deviations also average to 1. The products between the first and second deviations are and , so their average is also 1.
That last average is the covariance between the two coordinates. It is positive when deviations tend to have matching signs, negative when they tend to have opposite signs, and zero when their average product is zero. Statistical independence means that knowing one variable supplies no information about the other. Zero covariance does not generally establish independence: two variables may still have a relationship that this signed product does not detect.
Arrange these averages in a table. Diagonal entries compare a coordinate with itself and are variances; off-diagonal entries are covariances. For the two-row example the table is
To compute every pairwise product at once, place a centered observation in a column and multiply it by its transpose. This outer product produces a square table. Its entry in row , column is the product of centered coordinates and . The two indices select coordinates, not articles. Averaging these tables produces the covariance matrix.
Let index fitted articles, from 1 to , and let count embedding coordinates. Each is a column; row of the data matrix is . Let be its nonnegative fit weight. A weighted mean gives each observation the stated weight, adds the weighted contributions, then divides by the total weight. Let be that total and the weighted mean. A sum over includes every fitted article, and means :
For example, give the two observation rows above weights 1 and 3. The total is 4, the first mean coordinate becomes , and the second becomes 3.5. Centering must use this new weighted mean. The centered rows are and . The weighted mean of either squared coordinate is . Every entry of this example’s weighted covariance is 0.75.
Let name the weighted covariance. Apply precisely the same weights to the outer products of deviations from the weighted mean:
The first edition uses uniform article weights; the later archive fit uses inverse daily article counts, explained in §11. These weights are distinct from the story energy used to rank news. In Eigen Times , so is a matrix—147,456 numbers summarising the fitted article cloud. Figure 4 pictures the operations in their working order: center each observation, take its outer product, and average the products.
A practical point matters for updates. Let be the weighted coordinate sum and the weighted outer-product sum:
The mean is immediately . To recover covariance, first expand a single centered outer product:
Now take the weighted average of all four terms. The first becomes . In the second and third terms, the weighted average of is , so each becomes . The last term is the same for every observation and averages to . Two subtractions and one addition leave one subtraction:
The three running sums therefore retain everything needed for these mean and covariance calculations. They do not retain the full article collection. New observations with fixed weights can be added without revisiting old observations. If a day’s final article count changes its existing weights, the old contributions must also be corrected. Although the identity is exact, subtracting large nearly equal floating-point numbers can lose accuracy; the streaming chapter distinguishes algebraic identities from numerical reliability.
For compact matrix notation, let be the diagonal weight matrix: puts the listed values on the diagonal and zeros elsewhere. Let be a column of ones. Multiplying repeats the mean row times. The centered data matrix is therefore . Multiplying scales each centered row by its weight; multiplying on the left by then sums the weighted products. This gives the same covariance as .
A cloud of points in dimensions has a shape. It can be broad in one direction and narrow in another, even when neither direction follows an original coordinate axis. We want to measure that shape using perpendicular directions. The special directions of the covariance matrix will supply them.
A nonzero column is an eigenvector of with eigenvalue when . The left side multiplies the vector by a matrix; the right side multiplies it by a scalar. Equality says that this input stays on its own line under the matrix operation. For a positive eigenvalue it points the same way and changes length by that factor; a zero eigenvalue sends it to zero.
Consider the separate exact example
Multiplying by gives . Multiplying by gives . We have found two eigendirections and eigenvalues 3 and 1. Their dot product is zero. Their lengths are both , so divide each by to obtain unit columns, called and . Scaling an eigenvector to unit length does not change its eigenvalue.
This is a way to verify an answer, not an assumption that every covariance’s eigenvectors are easy to guess. Numerical algorithms find them for larger matrices. The purpose of the example is to make the equation an operation we can inspect.
For a covariance, eigenvalues are nonnegative. Here is why. Let be any unit column direction and let be observation ’s scalar projection score along it. Its weighted mean is zero because the observations were centered. Its weighted variance is the average squared score. Substituting the covariance definition gives
The left side is an average of nonnegative squares. If is a unit eigenvector with eigenvalue , the right side becomes . Thus the eigenvalue is exactly the variance along that unit direction and cannot be negative.
Choose unit, mutually perpendicular eigenvectors , ordered by decreasing eigenvalue , with . Real symmetric matrices permit such a choice, although repeated eigenvalues need not specify unique individual directions. Let contain those columns and the diagonal matrix of their eigenvalues, both . They diagonalise the covariance, meaning
where and means greater than or equal to. The diagonal representation has zero covariances between different eigenvector coordinates. The principal components are these directions, ranked by variance; using them to summarize the cloud is principal component analysis (PCA).
Read the factorization from right to left when applying it to a column: expresses the input in eigenvector coordinates, scales each coordinate separately, and expresses the result back in the original coordinates. The covariance is simple in this basis because no off-diagonal terms mix the coordinates.
Why does the first direction capture the most variance? Express an arbitrary unit direction as a combination , where each scalar is its coordinate in the full orthonormal eigenbasis. Its unit length means . Its projected variance becomes
The squared coordinates are nonnegative and add to one. The variance is therefore a weighted average of the eigenvalues and cannot exceed the largest, . Choosing reaches that largest value. Requiring the next direction to be perpendicular to sets its first coordinate to zero; the same argument then selects . This explains the ordering without assuming calculus.
Figure 5 shows 300 points in two dimensions: the long axis has eigenvalue 4.06, the short one 0.44, so about 90% of the displayed variation lies along one direction. The synthetic distribution has known mean zero. This illustration forms second moments about that origin, rather than subtracting the finite sample’s empirical mean; the displayed eigenvalues estimate population variances. The notebook computes both choices so that centering is an explicit decision.
This is the content of “the eigen‑news are the directions that
explain the most variance”. Variance describes spread in the chosen
numerical representation; the argument alone does not establish
significance, truth, or editorial importance. Computing eigenvectors for
a
symmetric matrix is a standard numerical operation (eigh in
common linear-algebra libraries; faer in the Rust
implementation) and takes milliseconds in the reported fit.
Let count retained directions; means less than or equal to. The explained-variance fraction is , provided total variance is positive. It reports how much of the fitted cloud’s variation those directions retain. For the Eigen Times embedding basis, 60 components carry 55% of the variance, and the curve is flat beyond; a small eigenvalue alone does not prove that a direction is meaningless.
In the exact two-direction example, total variance is . Keeping the first direction retains and discards . Keeping both retains all the variance. With a longer list of eigenvalues, make the same running calculation: add the retained values, divide by the sum of all values, and inspect what each extra direction contributes. A flat part of that curve means that each new direction adds little to the numerator. It does not mean that the discarded articles or subjects are unimportant.
The Marchenko–Pastur edge asks a different question: how large might covariance eigenvalues be in a specified model of featureless variation? Here noise means variation generated by that reference model, not a judgment about an article. The spectrum of a covariance is its collection of eigenvalues. In this illustration, let count independent observations in coordinates with equal noise variance , unrelated to the singular-value notation used later. For a concrete reference, suppose all entries are independent draws from the same zero-mean Gaussian distribution, a symmetric bell-shaped probability model, with variance . Independence applies across both observations and coordinates; repeated draws use the same distribution. These assumptions specify an illustrative noise model, not a description of news embeddings. Let denote the approximate upper edge of the model’s covariance spectrum. Under that model, in the large-sample limit,
Here raising to means taking a square root. First form the dimension-to-observation ratio , take its square root, add one, square, and multiply by the stipulated variance. The notebook uses , , and , giving approximately 0.010396. That is slightly larger than 0.01: finite sampling can spread estimated variances even when the model gives every coordinate the same underlying variance. This teaching input is not an estimate fitted to news.
Eigenvalues near or below such a reference may be noise; this is not a universal significance test for correlated, weighted news data. The median eigenvalue supplies a rough noise-variance estimate in the reported comparison. With and , the ratio is tiny and the edge is close to that estimate. The deployed bases use fixed , with the variance curve as a check. From here onward usually means only the retained columns and their positive eigenvalue matrix; full decompositions will be identified explicitly.