Figure 19 brings the stages together. Shapes are a useful final check because they force us to say what each row and column represents. They can detect incompatible operations, although matching shapes alone do not prove that row identities, units, weights, or meanings agree.
Here counts articles and is their embedding matrix. This use of names embeddings, distinct from the local covariance-perturbation symbol in the stability section. Each row is an article; each column is an embedding coordinate. Let be the column of ones and the embedding mean. The product repeats that mean in every article row. Subtracting it centres all rows at once.
The retained basis has 384 rows and 60 columns. Define the raw coordinate table by
The multiplication has shape , producing . Each article receives 60 coordinates. To reconstruct the retained part in the original embedding coordinates, multiply by , giving shape , then add the repeated mean. The difference from gives one residual row per article.
The term-space mean remains . The table denotes the standardized raw score rows used for lexical naming; it has the same article ordering as the term matrix. Its cross-product with centred terms has shape . The rotation changes the coordinates within the retained space without increasing their number. The subscript “lsa” below identifies a basis fitted in term space.
The numerical decompositions are consequently small—a covariance, a basis, a loading table, and a rotation—even though some article and story tables remain large. The preceding chapter identifies the objects that still determine memory requirements; small final matrices do not imply that every intermediate calculation is small.
| stage | object | shape | operation |
|---|---|---|---|
| tokenise | TF‑IDF matrix | , sparse | randomized SVD → , |
| embed | streaming | ||
| covariance | eigh →
|
||
| name | centred term–score product | varimax → | |
| measure | , , per row | ||
| stories | centroids | profiles, precedents | |
| days | spectra | era statistics, z‑scores |
The numbers in the table are those of the first edition on 28 August 2026; each grows by a day’s worth every night. The second edition, fitted on 3 September 2026 over the New York Times archive as well, has articles (455 million non-zeros in ), 13,421,090 stories and 63,270 day spectra; the small matrices—, , , —have exactly the same shapes, which is the point of the chapter Streaming the computation.