← First Pair Library

12 The pipeline as matrices

Figure 19 brings the stages together. Shapes are a useful final check because they force us to say what each row and column represents. They can detect incompatible operations, although matching shapes alone do not prove that row identities, units, weights, or meanings agree.

Here NN counts articles and EE is their N×384N\times384 embedding matrix. This use of EE names embeddings, distinct from the local covariance-perturbation symbol in the stability section. Each row is an article; each column is an embedding coordinate. Let 𝟏\mathbf1 be the N×1N\times1 column of ones and μ\mu the 384×1384\times1 embedding mean. The product 𝟏μ⊤\mathbf1\mu^{\top} repeats that mean in every article row. Subtracting it centres all rows at once.

The retained basis VV has 384 rows and 60 columns. Define the raw coordinate table CC by

C=(E−𝟏μ⊤)V.C=(E-\mathbf1\mu^{\top})V.

The multiplication has shape (N×384)(384×60)(N\times384)(384\times60), producing N×60N\times60. Each article receives 60 coordinates. To reconstruct the retained part in the original embedding coordinates, multiply by V⊤V^{\top}, giving shape (N×60)(60×384)=N×384(N\times60)(60\times384)=N\times384, then add the repeated mean. The difference from EE gives one residual row per article.

The term-space mean remains μT\mu_T. The table SS denotes the standardized raw score rows used for lexical naming; it has the same article ordering as the term matrix. Its cross-product with centred terms has shape (50,000×N)(N×60)=50,000×60(50{,}000\times N)(N\times60)=50{,}000\times60. The 60×6060\times60 rotation changes the coordinates within the retained space without increasing their number. The subscript “lsa” below identifies a basis fitted in term space.

The numerical decompositions are consequently small—a 384×384384\times384 covariance, a 384×60384\times60 basis, a 50,000×6050{,}000\times60 loading table, and a 60×6060\times60 rotation—even though some article and story tables remain large. The preceding chapter identifies the objects that still determine memory requirements; small final matrices do not imply that every intermediate calculation is small.

The pipeline, stage by stage, as matrices.
stage object shape operation
tokenise TF‑IDF matrix TT N×50,000N \times 50{,}000, sparse randomized SVD → VlsaV_{\text{lsa}}, UU
embed EE N×384N \times 384 streaming Z,m,MZ, m, M
covariance Σ\Sigma 384×384384 \times 384 eigh → V,ΛV, \Lambda
name centred term–score product β\beta 50,000×6050{,}000 \times 60 varimax → RR
measure C=(E−𝟏μ⊤)VC = (E - \mathbf{1}\mu^{\top})V N×60N \times 60 T2T^2, QQ, ν\nu per row
stories centroids 842,278×60842{,}278 \times 60 profiles, precedents
days spectra 10,102×6010{,}102 \times 60 era statistics, z‑scores

The numbers in the table are those of the first edition on 28 August 2026; each grows by a day’s worth every night. The second edition, fitted on 3 September 2026 over the New York Times archive as well, has N=16,500,781N = 16{,}500{,}781 articles (455 million non-zeros in TT), 13,421,090 stories and 63,270 day spectra; the small matrices—Σ\Sigma, VV, β\beta, RR—have exactly the same shapes, which is the point of the chapter Streaming the computation.