← First Pair Library

Reading and notation guide

0.1 How to work through the book

Chapters 1–3 build the numerical language. Chapters 4–6 find recurring directions and give them interpretable names. Chapter 7 measures a new article or story. Chapters 8–10 explain grouping, unusual days, and how directions retain their identity over time. Chapters 11–12 show how the same calculations can be accumulated in pieces rather than loading an archive into memory.

In a notebook, run the setup cell first and then run cells in reading order. A cell may use a result computed earlier. An assertion is an executable check: it stops the lesson if a mathematical identity or reference result fails. If you edit a teaching input, a reference answer may intentionally fail while an identity should still hold. Restarting the kernel—the running language process—and running all cells checks that the lesson does not depend on hidden state. Python and OCaml are alternative implementations of the same examples; neither language is a prerequisite for reading the prose.

The later volume Eigen Times Math History places these methods in a historical sequence. Its companion adds techniques beyond the present book; its complete teaching text is available in the History Math notebooks. The present chapters are the direct prerequisites for the calculations here; the historical volume is a route to the foundational literature, not a substitute for a missing teaching step.

0.2 Reading symbols before calculating

A scalar is a single number. A vector is an ordered list of numbers. A matrix is a rectangular table. An entry is one number inside a vector or matrix; a coordinate describes an amount along a chosen direction. Coordinates depend on which directions we choose.

Lowercase subscripts identify entries: xix_i is observation number ii, while xijx_{ij} is its coordinate number jj. Context distinguishes an observation index from a coordinate index. An index is a label, not a quantity to multiply. Superscripts usually indicate powers, so a2=aaa^2=a\,a, except that ⊤\top means transpose: exchange rows and columns. An accent is also part of the name: x̂\hat{x} means an estimate or reconstruction of xx, not a new multiplication.

For numbers a1,…,ana_1,\ldots,a_n, the dots mean the intervening entries. The expression ∑i=1nai\sum_{i=1}^{n}a_i means a1+⋯+ana_1+\cdots+a_n. The symbol ∏\prod means multiplication over an indicated set of entries. Parentheses group operations: first calculate what is inside. The square root a\sqrt{a} is the nonnegative number whose square is aa, for a≥0a\geq0. The relations ≤\leq and ≥\geq allow equality; << and >> do not. The relation ≈\approx means approximate equality, usually because a displayed decimal is rounded.

The real numbers, denoted ℝ\mathbb{R}, include positive and negative numbers and fractions. Writing x∈ℝdx\in\mathbb{R}^{d} says that xx has dd real entries. A matrix with nn rows and dd columns has shape n×dn\times d, also written ℝn×d\mathbb{R}^{n\times d}. Shapes count entries; they do not indicate physical units. The zero vector has every entry zero. The identity matrix II is square, with ones on its main diagonal and zeros elsewhere; multiplying by it leaves a compatible vector unchanged.

Single bars |a||a| mean the absolute value of a scalar, its size without its sign. Double bars ‖x‖\lVert x\rVert mean the Euclidean length of a vector: square its entries, add, and take a square root. The dot product multiplies corresponding entries of two equally long vectors and adds the products. Both operations receive explicit calculations in Chapter 2. Adjacent matrices or vectors mean a matrix product only when their inner dimensions agree; Chapter 3 works through that rule.

0.3 The book’s main objects

This is a map of symbols, not a formula to memorize. The chapters repeat the definitions at the point of use.

Symbol Meaning and shape in this book
nn or NN Number of observations in the calculation; NN often denotes an archive total.
pp, dd, kk Number of vocabulary terms, embedding coordinates, and retained directions, respectively. An embedding is a learned numerical representation of text.
xix_i, XX One observation as a d×1d\times1 column; a data table with one transposed observation per row. For lexical data the input width is pp.
μ\mu, Σ\Sigma Mean column and covariance matrix: the average location and the table describing variation around it.
wiw_i, ZZ Nonnegative fitting weight of observation ii and sum of all fitting weights. The sum must be positive.
VV, Λ\Lambda Columns of retained unit directions, and a diagonal table of their variances. Their shapes are d×kd\times k and k×kk\times k.
cc, x̂\hat{x}, rr Coordinates along retained directions, reconstructed observation, and residual—the difference left after reconstruction. Their lengths are kk, dd, and dd.
RR, yy, Γ\Gamma A k×kk\times k naming rotation, the kk named coordinates, and their k×kk\times k covariance.
T2T^2, QQ, ν\nu Variance-scaled squared distance within the retained space; squared residual length outside it; fraction of total centred squared length left outside.
TT, SS, β\beta Term table, table of variance-scaled raw scores, and table associating terms with directions. These are not the scalar statistic T2T^2.
UU, Σsvd\Sigma_{\mathrm{svd}}, VV The three matrix factors in a singular value decomposition; their local shapes are introduced in Chapter 5.
σj\sigma_j, λj\lambda_j Singular value number jj and covariance eigenvalue number jj. A bare σ\sigma in threshold discussions instead denotes a standard deviation.

The symbol VV appears in both the eigenvector and singular value chapters because the directions are closely related; the surrounding data table determines its shape. The symbol Σ\Sigma always means covariance here, except that an original SVD diagram uses it for the singular-value matrix and its caption explains that local convention. This prose writes that matrix as Σsvd\Sigma_{\mathrm{svd}}.

Across the series, the mathematical objects stay the same while storage conventions can differ. Here observations are columns in projection formulas and their transposes form data-table rows; in row-based formulas the same multiplication is transposed. Here Σ\Sigma means covariance and ZZ means total fitting weight. The History companion uses Γ\Gamma for covariance and ZZ for a standardized news table in some lessons. Read the local definition before transferring a formula. The present book reserves Γ\Gamma for the covariance after a naming rotation and keeps these conventions consistent throughout.