← First Pair Library

The Mathematics of Anthropology

A Step-by-Step Companion to The Dual Geometry of People and News

Alexy Khrabrov

4 October 2026 | Edition 1.1.0-9ce09c2a

Preface

Anthropology lets a reader start with a person, find the kinds of news associated with that person, follow those news patterns to other people, and return to the articles. Its research paper, The Dual Geometry of People and News, describes the system and its evidence. This companion explains the mathematics one operation at a time.

The starting point is The Mathematics of Eigen Times. That book explains how articles become vectors, how covariance reveals recurring news directions, and how projections describe a new story. Here we take the next step. A collection of articles associated with a person becomes a profile in news space. A population of profiles produces another set of directions: patterns of how people are covered. One loading matrix connects the two spaces.

You need arithmetic and a willingness to read a small table. The next two chapters introduce the vector and matrix ideas needed to continue. When a prerequisite deserves a longer explanation, a link takes you to the relevant Eigen Times section. We do not repeat its derivations of text weighting, efficient matrix decomposition, or clustering thresholds. We spend that space on the new questions: how to combine a person’s articles, what to compare that coverage with, what a people pattern means, why navigation works in both directions, and how shared ontology concepts could connect different news corpora.

One fictional example runs through the book. It has four people, three news coordinates, two source-and-decade cells, and 45 distinct articles. The running example and plotted numerical matrices are generated by the accompanying script. Short additional arithmetic illustrations are written out in the text. Figures use blue for positive values and rust for negative values; color describes a sign, never a judgment about a person. Displayed values are rounded, while computations use their full precision.

The empirical examples come from the Anthropology paper’s 2 October 2026 edition. Its people models retain the 1 October snapshot. The catalog has subsequently grown to 59,399 people and organizations, but the fitted models still contain 217 Hacker News profiles and 61 general-news profiles. A catalog identity and a supported mathematical profile are different objects. None of the fictional numbers below is a measurement of a real person.

Version 1.1.0. This teaching expansion works through lengths, residuals, covariance, eigenvectors, SVD, profile aggregation, neighbor geometry, attention, and proposed ontology models in smaller steps. The complete Python and native OCaml notebooks compute the new synthetic examples as well as the original fixtures. Original numerical evidence, figure inputs, and the dated Jeff Dean journey are unchanged.

We will also keep implemented operations separate from research proposals. The published system computes the profiles, people basis, reciprocal indexes, and dated news overlay described here. Incremental profile refresh, learned ontology calibration, and a validated alignment between the two corpora are extensions to evaluate. Explaining their mathematics does not make them deployed features.

Where to revisit a prerequisite

If you want more on Read in the Eigen Times companion
Turning words into numbers Text as vectors
Dot products and angles Cosine similarity
Rows, columns, multiplication, transpose Matrices pictured
Means, covariance, and eigenvectors The axes of a point cloud
Singular value decomposition (SVD) and low-rank approximation The singular value decomposition
Naming axes by rotation Varimax
Projection and missing information Projection, reconstruction, residual
Updating sums without rereading everything Sums over blocks

Reading and notation guide

A formula is a compact record of operations. First identify the inputs, then the operation, then what its output measures. Each chapter follows that order. The original four-person example remains one connected calculation; smaller examples are explicitly synthetic and set aside that fitted model temporarily. A successful arithmetic check does not validate the archive’s identities, attribution, or social claims.

In either notebook, run the setup and then the cells in reading order. A kernel is the running language process. Restarting it and running every cell checks that no hidden earlier calculation is needed. An assertion is a check that stops execution when a stated identity or reference answer fails. Edit the declared inputs to experiment; when an input changes, a fixed reference answer can fail deliberately while an identity should still hold. The Python and OCaml versions compute independently and need no private archive.

Read a symbol as a named object

A scalar is one number, a vector is an ordered list, and a matrix is a rectangular table. Here an observation vector is a row. A direction used for projection is a column. A shape such as 4×34\times3 means four rows and three columns. The real numbers, written ℝ\mathbb R, include negative values, zero, and fractions; ℝ4×3\mathbb R^{4\times3} means a table of that shape with real entries.

An index selects an entry: vjv_j is coordinate jj of row vv, and WjℓW_{j\ell} is row jj, column ℓ\ell of matrix WW. The index is not a multiplier. The notation ∑j=13vj\sum_{j=1}^{3}v_j means v1+v2+v3v_1+v_2+v_3. Adjacent scalar symbols mean multiplication; adjacent matrices mean the row-by-column product taught below. Parentheses group operations, so work inside them first. The symbol ∈\in means membership, and |𝒜||\mathcal A| counts the members of a set 𝒜\mathcal A.

A square a2a^2 means aa multiplied by itself. For nonnegative aa, a\sqrt a is the nonnegative number whose square is aa. Single bars |a||a| give a scalar’s absolute value, removing its sign; double bars ‖v‖\lVert v\rVert give a vector’s length. A bar above a symbol denotes an average, and a hat denotes a reconstruction or estimate. The superscripts 𝖳\mathsf T and ⊤\top both transpose a matrix, exchanging rows and columns. The symbol II denotes a square identity matrix, with diagonal ones and other entries zero. A superscript −1-1 on a square matrix means an inverse when it exists. The symbol ≈\approx warns of rounded or approximate equality.

Keep the measuring objects distinct

This table is a reading map. The chapters introduce each object again with its dimensions before using it in a calculation.

Object Symbol and meaning
Article row zaz_a: article aa measured in standardized news coordinates.
Comparison within a cell qgq_g: unique-article background; mpgm_{pg}: person pp’s mean in cell gg.
Person before and after normalization rpr_p: balanced contrast; bpb_p: its unit direction.
Population reference b‾\bar b: average fitted person profile; hp=bp−b‾h_p=b_p-\bar b: centered row.
Learned people directions WW: loadings mapping news coordinates to retained people coordinates upu_p.
Dimension counts dd: embedding width; KK: news width; rr: retained people width. The scalar rr differs from row rpr_p.
Sample counts nn: fitted people; NN: reviewed training articles in the proposed concept model.
Concept model AA: slopes; β\beta: intercepts; ss: predicted score row.
Local reuse of a letter QQ is an SVD direction matrix; QpQ_p is a scalar residual statistic; QalignQ_{\mathrm{align}} is an alignment matrix.

Normalization makes one row have length one. Standardization divides each coordinate by its reference standard deviation. Centering subtracts a reference mean. These answer different questions and cannot be substituted for each other. A positive coordinate names one pole of an axis; its sign does not express approval, goodness, truth, or confidence.

Across volumes, equivalent roles sometimes use different symbols. The History Math companion’s unified notation uses LL for the retained people dimension count, WW for the direction matrix, SS for the score table, Θ\Theta for concept coefficients, and CC for the concept count. This established Anthropology edition writes those roles as rr, the same WW, rows upu_p stacked together, AA, and mm, respectively. History counts people with mm and writes their covariance as ΓB\Gamma_B; this book uses nn and CPC_P for those roles. Do not substitute letters without also matching their shapes and centering conventions. Here CPC_P remains the people covariance, not the concept count. Here QQ names the SVD’s right-direction matrix, while QpQ_p is a person’s scalar discarded squared length. These distinct objects never share a formula merely because their letters resemble one another.

The final notation table collects shapes. The glossary defines terms in words, and the subject index links to worked explanations. The Eigen Times companion sometimes writes observations as columns; transpose its complete formula when translating to this book’s row convention.

1 What it means to make a person a news source

Imagine collecting articles that mention a database founder. Some concern database systems, some concern enterprise sales, and some concern unrelated activities. Each article can be measured on the same fixed news directions. The collection then has a mathematical shape: it spends more of its coverage in some directions than in others.

That is the sense in which a person becomes a source of news. It is a distribution of associated reporting. It does not mean the person wrote, approved, or caused the articles. Nor is the resulting vector a psychological portrait. Change the corpus (the article collection), the dates, or the attribution rules and the profile can change.

Three objects must remain separate throughout the calculation. An identity says which person a record denotes. An article association says that a particular article is a candidate match to that person. A relationship assertion says something specific, such as that one person reported to another during a stated period, with a source for the claim. A reliable identity does not automatically make every name match reliable. Two people appearing in an article does not establish a relationship between them.

The vector model starts from qualified candidate article associations. Anthropology’s graph stores separately sourced assertions and explicitly marked co-mentions. The same interface can display both, but one is not evidence for the other. A short distance in the vector view means similar coverage under that model. A line labeled employment in the graph needs employment evidence.

There will also be two meanings of people vector. A person profile describes one individual. A people pattern is a direction learned from the variation among many profiles. Think of the difference between a city’s weather measurements and a recurring weather pattern. A city has coordinates on the pattern; it is not itself the pattern.

The diagram is a roadmap of operations explained below. Project means measure along selected directions; rotate means change their orientation; standardize means divide by reference standard deviations. Center means subtract a comparison mean, and normalize means divide a nonzero row by its length. Principal component analysis (PCA) finds directions of greatest variation among centered observations. These names label the stages, not additional assumptions about a person’s identity.

The mathematical path from articles to person profiles and people patterns. The evidence graph remains a separate route back to sourced claims.

The book follows this path in the order shown before explaining how the middle can be explored in both directions. At each step we will ask what the numbers mean and what information the operation discards.

Check 1. An article names two people. How many unique articles should it contribute to a background average, and how many person-article associations might it contribute? Keep your answer for the solutions at the end.

2 The small amount of linear algebra we need

2.1 A vector is a measured row

A scalar is a single number. A vector is an ordered list of numbers called coordinates. Our observation vectors are rows; when a matrix direction or another object is a column, we will say so. Let vv denote this example row:

v=(2,1,−1)v=(2,1,-1)

It has three coordinates. The order matters: its first coordinate is 2, its second is 1, and its third is -1. In the running example the coordinates belong to three fixed news axes, called N1, N2, and N3. These names identify toy directions, not actual Anthropology topics.

Let w=(1,0,1)w=(1,0,1) be a second row of the same length. The index jj selects a coordinate: vjv_j and wjw_j are the numbers in position jj of the two rows. The summation sign ∑j\sum_j means to add over all those positions, from 1 to 3 here. The dot product, written with a centered dot, multiplies corresponding coordinates and adds:

v⋅w=∑jvjwj.v\cdot w=\sum_j v_jw_j.

For these rows, the result is 2×1+1×0+(−1)×1=12\times1+1\times0+(-1)\times1=1. An omitted multiplication sign between scalar entries also means multiply.

The Euclidean length, or norm, of vv is written ‖v‖\lVert v\rVert. It is the square root of the sum of squared coordinates:

‖v‖=v⋅v=22+12+(−1)2=6.\lVert v\rVert=\sqrt{v\cdot v}=\sqrt{2^2+1^2+(-1)^2}=\sqrt6.

Dividing a nonzero vector by its length gives a unit vector. It keeps the direction while removing the size. The zero vector has no direction and cannot be normalized this way.

Single bars around a number mean its absolute value, or size without its sign. Thus |−2|=2|-2|=2. Single bars around a set will instead count its members; we will introduce those sets before using them. The symbol ≈\approx means approximately equal, as when values have been rounded.

2.1.1 Length is built from squares

Why square before adding? Simply adding the entries of (3,−3)(3,-3) gives zero even though the row is not zero. Squaring counts both directions as positive contributions. For a separate synthetic row (3,4)(3,4), the squares are 9 and 16, their sum is 25, and the length is 25=5\sqrt{25}=5. This extends the right-triangle rule: perpendicular coordinate contributions add as squared lengths.

For the earlier row v=(2,1,−1)v=(2,1,-1), the squared contributions are 4, 1, and 1. Its squared length is 6 and its length is approximately 2.44949. Dividing each coordinate by that same number gives approximately (0.81650,0.40825,−0.40825)(0.81650,0.40825,-0.40825). Square these normalized coordinates and add: 4/6+1/6+1/6=14/6+1/6+1/6=1. That is the unit-length check; it does not require every individual coordinate to equal one.

Multiplying (3,4)(3,4) by 2 gives (6,8)(6,8) and doubles its length from 5 to 10. Normalizing either gives (0.6,0.8)(0.6,0.8). Multiplication by -2 instead reverses the direction, giving unit row (−0.6,−0.8)(-0.6,-0.8). Positive rescaling preserves direction; a negative rescaling reverses it. Neither normalization establishes that one observation has more supporting evidence than another.

Try it. Change the row to (0,4)(0,4), then to (0,0)(0,0). The first has unit direction (0,1)(0,1). The second has length zero, and division by its length is undefined. A missing or directionless profile must not silently become an average profile.

2.2 A matrix connects two lists of coordinates

A matrix is a rectangular table. Let nn count its rows and KK its columns; its shape is written n×Kn\times K. Put one three-coordinate person profile in each of four rows and the result is 4×34\times3. A matrix entry uses its row index first and column index second. Diagonal entries have equal row and column indices; they run from the top left toward the bottom right.

Multiplication describes a weighted combination. Its numerical weights are coefficients; coefficients expressing directions in input coordinates are also called loadings. If hh is a 1×31\times3 row and WW is a 3×23\times2 loading matrix, then hWhW is a 1×21\times2 row. Its first number is the dot product of hh with the first column of WW. Its second number uses the second column. The adjacent dimensions must agree:

(1×3)(3×2)=(1×2). (1\times3)(3\times2)=(1\times2).

The transpose, written W𝖳W^{\mathsf T}, exchanges rows and columns. It therefore has shape 2×32\times3. Remembering shapes will prevent more mistakes than memorizing a complicated formula.

2.2.1 Multiply one output coordinate at a time

Use a separate synthetic transformation TT, with three input coordinates and two output coordinates:

T=(100211).T=\begin{pmatrix}1&0\\0&2\\1&1\end{pmatrix}.

For v=(2,1,−1)v=(2,1,-1), the first output is 2(1)+1(0)+(−1)(1)=12(1)+1(0)+(-1)(1)=1. The second is 2(0)+1(2)+(−1)(1)=12(0)+1(2)+(-1)(1)=1. Thus vT=(1,1)vT=(1,1). We multiplied a 1×31\times3 row by a 3×23\times2 table, producing a 1×21\times2 row. Each output column says how much of each input enters that output.

Transposing gives T𝖳=(101021)T^{\mathsf T}=\begin{pmatrix}1&0&1\\0&2&1\end{pmatrix}. Multiplying the returned row gives (1,1)T𝖳=(1,2,2)(1,1)T^{\mathsf T}=(1,2,2), which differs from vv. This deliberately arbitrary table does not have perpendicular unit columns, and its transpose is not an inverse. Shapes allow a multiplication; they do not promise what that multiplication will recover.

2.3 Centering gives a comparison point

Suppose three measurements are 2, 3, and 7. Their mean is 4. Centering subtracts that mean, giving -2, -1, and 3. A negative centered value means below the chosen mean; it does not mean a negative amount of the original quantity.

For vector rows we average each column separately and subtract the resulting mean row. We will center twice for different purposes: first against a news background within a source-and-decade cell, then against the mean profile of the fitted population. These are two different comparisons, not a redundant repetition.

Variance measures spread by averaging squared deviations from a mean. A standard deviation is the nonnegative square root of a variance; it is zero when all values agree. Covariance averages the product of the centered values of two coordinates: a positive result means they tend to move together, a negative result means they tend to move oppositely. A covariance matrix collects those comparisons for every pair of coordinates. Its diagonal contains their individual variances. The fitted news model supplies its covariance convention; for the people calculation below we explicitly divide by the number of people.

2.3.1 A full variance and covariance calculation

The mean of 2, 3, and 7 is (2+3+7)/3=12/3=4(2+3+7)/3=12/3=4. The centered numbers are -2, -1, and 3; their sum is zero. Their squared deviations are 4, 1, and 9. For this descriptive example we divide by the observation count, three: variance is (4+1+9)/3=14/3(4+1+9)/3=14/3, approximately 4.6667. The standard deviation is 14/3\sqrt{14/3}, approximately 2.1602. Variance has squared coordinate units; standard deviation returns to the original units.

Pair these measurements with a second coordinate, 6, 4, and 2, in matching row order. The second mean is 4. Its deviations are 2, 0, and -2. The table keeps the pairings visible:

Row First deviation Second deviation First squared Second squared Product
1 -2 2 4 4 -4
2 -1 0 1 0 0
3 3 -2 9 4 -6
Sum 0 0 14 8 -10

Divide the last three column totals by three. The resulting covariance matrix is

(14/3−10/3−10/38/3).\begin{pmatrix}14/3&-10/3\\-10/3&8/3\end{pmatrix}.

The repeated off-diagonal entry is the average paired product. Reversing the order of the two factors leaves each product unchanged, so the covariance matrix is symmetric. The negative sign records that a positive deviation in one coordinate tends to accompany a negative deviation in the other. It is not a causal claim.

Dividing centered columns by their standard deviations yields unit variances. Their covariance becomes the correlation (−10/3)/(14/38/3)=−10/112(-10/3)/(\sqrt{14/3}\sqrt{8/3})=-10/\sqrt{112}, approximately -0.9449; standardization has not removed the association. Some estimators divide by one less than the count to estimate a population variance from a sample. Here the count denominator defines the spread of these declared rows; the people model below uses that same convention. Changing a denominator must be an explicit modeling choice.

Check the mechanism. If the second coordinate were constant, every second deviation would be zero. Its variance and every covariance involving it would be zero, and division by its standard deviation would be unavailable.

2.4 Rotation and projection are different operations

Perpendicular directions have dot product zero. Perpendicular unit directions are orthonormal. Let II denote an identity matrix: it has ones on its diagonal and zeros elsewhere, so multiplying by it changes nothing. Its size matches the matrix product beside it. Let RR be a square matrix whose columns form a complete orthonormal system. Such a matrix is called orthogonal and satisfies R𝖳R=RR𝖳=IR^{\mathsf T}R=RR^{\mathsf T}=I. Multiplying by RR rotates or reflects coordinates and preserves all lengths and angles.

A tall matrix WW can also have orthonormal columns: W𝖳W=IW^{\mathsf T}W=I. But when it has fewer columns than rows, multiplying by WW keeps only some directions. The length can decrease. The statement that an orthogonal change of coordinates preserves lengths applies to a complete rotation; a truncated projection preserves only the retained component. This distinction will explain the navigation duality.

Directions are linearly independent when none is a weighted combination of the others. A basis is a set of independent directions used to express coordinates. The subspace spanned by a collection of directions contains all their weighted combinations. The rank of a matrix counts its independent directions; keeping fewer directions is a low-rank representation.

The Eigen Times matrix pictures provide a longer visual introduction. When comparing a formula written with column vectors to this book, transpose it consistently rather than changing only one factor.

Check 2. A row has 60 news coordinates and its loading matrix WW has shape 60×2460\times24. What are the shapes of the forward result, the transpose, the combined matrix WW𝖳WW^{\mathsf T}, and the returned row?

2.4.1 Meet the residual before measuring it

Take the synthetic row (3,4)(3,4) and retain only the horizontal unit direction (1,0)𝖳(1,0)^{\mathsf T}. The dot product is 3(1)+4(0)=33(1)+4(0)=3. Rebuilding three units of that direction gives (3,0)(3,0). The residual is what remains when we subtract the reconstruction from the original: (3,4)−(3,0)=(0,4)(3,4)-(3,0)=(0,4).

The residual is a row, not initially a scalar error score. It tells us where the missing part points. Its squared length is 02+42=160^2+4^2=16. The reconstruction’s squared length is 9; the original’s is 25. Thus 25=9+1625=9+16. Their dot product is zero, so the retained and residual parts are perpendicular. A later chapter applies this same calculation to the fitted people plane.

For comparison, multiplying (3,4)(3,4) by the complete orthogonal matrix (0−110)\begin{pmatrix}0&-1\\1&0\end{pmatrix} gives (4,−3)(4,-3). Its squared length is still 25. Multiplication by the transpose recovers (3,4)(3,4) exactly. A complete change of orientation has no discarded part; keeping only the horizontal coordinate did.

3 Measuring an article on named news axes

3.1 First inherit a fixed news map

Eigen Times and Eigen Hacks turn an article’s text into an embedding, a list of numbers produced by a text model. Anthropology does not learn a new text embedding for every person. It inherits a versioned news coordinate system and measures each associated article in that system.

Let aa index an article and dd count embedding coordinates. Let xax_a be article aa’s 1×d1\times d embedding row, and μ\mu the 1×d1\times d mean embedding used when the news basis was fitted. Let KK count retained news directions and let VV contain them as orthonormal columns, so VV has shape d×Kd\times K. Denote the resulting 1×K1\times K raw news-coordinate row by cac_a:

ca=(xa−μ)V.c_a=(x_a-\mu)V.

Read this in two steps: subtract the reference mean, then take a dot product with each retained direction. In the audited implementation, d=384d=384 and K=60K=60. This projection has already discarded information before people enter the calculation.

Why these particular directions? They are the directions of greatest fitted variation, found from a covariance matrix. The Eigen Times covariance chapter develops that argument. For the moment, regard VV as a fixed measuring instrument, accompanied by its mean and version.

3.1.1 Read the article pipeline as successive measurements

A synthetic three-coordinate article can make every operation visible. Set its embedding row to xa=(5,2,0)x_a=(5,2,0) and its fitted reference mean to μ=(1,1,1)\mu=(1,1,1). Subtraction produces (4,1,−1)(4,1,-1). Choose all three coordinate directions as V=IV=I and introduce the square naming rotation R=IR=I for this small illustration. The raw and named rows are then both (4,1,−1)(4,1,-1). This identity choice isolates the scaling arithmetic in the next section; it is not a claim that the production embedding or rotation is an identity.

The notebook’s batch fixture uses the same synthetic reference mean and scales for all 45 articles. It constructs the embedding rows from the declared standardized rows and verifies that the forward operations recover those rows. The fitted production mean, directions, and scales instead come from their stated news model; fitting them anew to one person’s articles would change the measuring instrument.

3.2 Naming changes orientation

A statistically useful direction need not have a readable name. Eigen applies a stored K×KK\times K orthogonal rotation RR. Denote the resulting 1×K1\times K named news-coordinate row by yay_a:

ya=caR.y_a=c_aR.

Lexical loadings measure associations between words and numerical directions. Varimax is an orientation rule that makes these loadings more concentrated on individual axes. Eigen uses it to choose the rotation, helping some axes acquire recognizable word lists. The Eigen Times varimax explanation shows why this can improve interpretability.

An ontology is a reusable vocabulary of concepts and their relationships. The word lists are labels to investigate, not ontology definitions. Stemming shortens words by removing endings, so a label such as databas represents a processed word form. A list containing databas, ceo, and interview can describe mixed coverage. An axis has two poles: its positive and negative directions. Positive-loading words may describe one pole better than the other. In the implementation, the lexical naming fit uses raw scores divided by their standard deviations, called standardized raw scores, while the stored rotation acts on unstandardized article scores. Operations commute when changing their order leaves the result unchanged. Scaling and rotation need not commute, so the lexical lists are a naming heuristic rather than exact term contributions to every final coordinate.

3.3 Put each named coordinate in comparable units

Some named directions vary more than others. A coordinate of 2 on one direction may be ordinary while 2 on another is unusual. Anthropology divides each named coordinate by its fitted standard deviation.

Let jj index named news axes, from 1 to KK, and let σj\sigma_j be the positive standard deviation of axis jj. Let DD be the K×KK\times K diagonal matrix containing those scales: its row-jj, column-jj entry is Djj=σjD_{jj}=\sigma_j, and all off-diagonal entries are zero. The superscript −1-1 denotes a matrix inverse, which undoes the corresponding multiplication. Here D−1D^{-1} has reciprocal scales on its diagonal, so multiplying by it divides coordinate jj by σj\sigma_j. Denote the standardized 1×K1\times K article row by zaz_a:

za=yaD−1.z_a=y_aD^{-1}.

For example, with standard deviations (2,1,0.5)(2,1,0.5), named coordinates (4,1,−1)(4,1,-1) become standardized coordinates (2,1,−2)(2,1,-2). When discussing a generic article we omit its subscript and write zz for such a standardized row; zjz_j is its coordinate jj. We will aggregate these rows into people profiles.

An eigenvector of a covariance matrix is a direction that multiplication by that matrix stretches without turning; its eigenvalue is the stretch factor. Let ii index the retained raw axes, from 1 to KK, and let λi\lambda_i be their positive fitted variances, which are eigenvalues of the raw covariance. The coefficient RijR_{ij} is the entry in row ii, column jj of the rotation. Summing over all raw axes gives each named-axis variance:

σj2=∑iλiRij2.\sigma_j^2=\sum_i\lambda_iR_{ij}^{2}.

Each rotated variance is a weighted combination of raw variances. The squared rotation coefficients supply the weights. The mean, directions, rotation, variances, and corpus identity must travel together. Axis N24 in one model is not automatically N24 in another.

3.3.1 Standardize across coordinates, normalize across a row

With named row (4,1,−1)(4,1,-1) and fitted standard deviations (2,1,0.5)(2,1,0.5), divide coordinate by coordinate: 4/2=24/2=2, 1/1=11/1=1, and −1/0.5=−2-1/0.5=-2. The standardized row (2,1,−2)(2,1,-2) has length 4+1+4=3\sqrt{4+1+4}=3, so standardization has not made it a unit row. Normalizing it would be a further operation, producing (2/3,1/3,−2/3)(2/3,1/3,-2/3) and losing its overall magnitude.

Each standard deviation belongs to a column of the fitted reference population. Each row length belongs to an individual observation. The first operation changes units; the second discards size. A zero standard deviation cannot be used as a divisor and requires an explicit modeling policy, not silent division. The formulas in this chapter assume positive fitted scales.

3.4 Standardization is not whitening

Marginal standardization gives each coordinate unit variance separately. Whitening would also make the covariance between different coordinates zero; we will see that these are different operations. In this illustration only, use two raw directions with variances 4 and 1. Rotate them by 45 degrees using the following 2×22\times2 matrix:

R=12(1−111).R=\frac1{\sqrt2}\begin{pmatrix}1&-1\\1&1\end{pmatrix}.

Write Cov⁡(z)\operatorname{Cov}(z) for the feature covariance matrix of the standardized article rows in this illustration. Both rotated coordinates have variance 2.5, but they also have covariance -1.5. Divide both by 2.5\sqrt{2.5} and their variances become 1 while their covariance becomes -0.6:

Cov⁡(z)=(1−0.6−0.61).\operatorname{Cov}(z)=\begin{pmatrix}1&-0.6\\-0.6&1\end{pmatrix}.

They now have the same marginal scale. They still move together. Correlation is covariance divided by the two standard deviations; when both variances are one, the correlation equals the covariance. Full whitening would remove these off-diagonal correlations. This pipeline performs marginal standardization, not full whitening.

A rotation mixes the raw variances. Dividing by the new marginal standard deviations gives unit diagonal entries but leaves correlation.

For the raw row (2,0)(2,0), rotating and then marginally scaling gives approximately (0.894,−0.894)(0.894,-0.894). Scaling the raw coordinates first by (2,1)(2,1) and then rotating gives (0.707,−0.707)(0.707,-0.707). The two operations answer different questions. Their order must be part of the model definition.

This also changes the geometry in which people are compared. Equal lengths in standardized news units need not be equal lengths in the original embedding. A Mahalanobis statistic measures squared distance from a reference mean while accounting for covariance, not just each coordinate’s separate scale. The sum ∑jzj2\sum_jz_j^2 is not generally that original statistic; see the Eigen Times section on residuals and T squared. Outside this two-dimensional illustration, the symbols retain their fitted-model meanings.

Check 3. Does a covariance matrix with ones on its diagonal necessarily describe uncorrelated coordinates? Use the two-dimensional example to answer.

3.4.1 Unroll the rotation arithmetic

Write the raw coordinate covariance as diag⁡(4,1)\operatorname{diag}(4,1), where diag\operatorname{diag} constructs a diagonal matrix from its listed entries. In the displayed rotation, the first named coordinate is the sum of the two raw coordinates divided by 2\sqrt2; the second is their difference, with the first raw coordinate negative, divided by 2\sqrt2.

Because the raw coordinates have zero covariance, the first named variance is 4(1/2)2+1(1/2)2=2+0.5=2.54(1/\sqrt2)^2+1(1/\sqrt2)^2=2+0.5=2.5. The second uses the squared negative coefficient and has the same variance. Their covariance uses paired, unsquared coefficients: 4(1/2)(−1/2)+1(1/2)(1/2)=−2+0.5=−1.54(1/\sqrt2)(-1/\sqrt2)+1(1/\sqrt2)(1/\sqrt2)=-2+0.5=-1.5. Dividing the two named coordinates by 2.5\sqrt{2.5} divides this covariance by 2.52.5=2.5\sqrt{2.5}\sqrt{2.5}=2.5, leaving −1.5/2.5=−0.6-1.5/2.5=-0.6.

The row (2,0)(2,0) first rotates to (2,−2)(\sqrt2,-\sqrt2) and then scales to (2/2.5,−2/2.5)(\sqrt{2/2.5},-\sqrt{2/2.5}). Scaling first instead gives (1,0)(1,0) and rotating gives (1/2,−1/2)(1/\sqrt2,-1/\sqrt2). These routes standardize relative to different coordinate systems, so they use different fitted scale matrices and define different maps. The named-coordinate scale here is a scalar multiple of the identity and actually commutes with the rotation. This example therefore does not test noncommutation of one fixed pair of matrices; each standardizer belongs to the coordinates whose variances it measures.

4 Combining articles into a person profile

4.1 Start with associations and keep the background unique

Our fictional example has people A, B, C, and D. The 45 distinct articles are grouped below only to save space: articles in a batch have identical toy coordinates and associations. They are distinct observations, not duplicate records to count repeatedly. E and L identify two source-and-decade cells. A cell specifies both the source and the decade; the letters do not imply that every early or late article belongs together.

The coordinates in the table are already standardized news coordinates zz. Their reference basis and scales are stipulated for this toy example, not estimated from the 45 rows.

Batch Cell Distinct articles N1 N2 N3 Associated people
E1 E 10 2 0 1 A, B
E2 E 5 1 1 2 A
E3 E 5 0 2 2 C
E4 E 5 -1 0 -2 D
L1 L 5 1 -1 2 A
L2 L 5 2 1 -2 B
L3 L 5 0 3 2 B, C
L4 L 5 -2 0 -1 D

There are 25 distinct articles in E and 20 in L. A has 20 associated articles, B has 20, and C and D have 10 each. Thus there are 60 person-article associations but only 45 unique articles. The ten E1 articles each name two people, as do the five L3 articles. This is why adding person counts is not a way to count the archive.

An incidence matrix records these associations. Its rows are people, its columns are articles, and an entry is 1 when the article is associated with the person and 0 otherwise. It is a bookkeeping device; it does not turn the association into a verified identity match or a social relationship. The figure compresses identical article columns into one column per batch, with its multiplicity shown below.

The toy association table and fixed news coordinates. Shared batches contribute to more than one person but only once to a cell’s background.

4.2 A cell supplies its own comparison

Let pp index a person and gg a source-and-decade cell. Let 𝒜g\mathcal A_g be the set of unique eligible background articles in cell gg, and 𝒜pg\mathcal A_{pg} its subset associated with person pp. For any set 𝒜\mathcal A, |𝒜||\mathcal A| denotes its number of members, called its cardinality. The notation a∈𝒜ga\in\mathcal A_g means that article aa belongs to that set; a sum with that condition includes each member once. Recall that zaz_a is article aa’s standardized news-coordinate row. Let qgq_g be the background mean and mpgm_{pg} the person-cell mean, both 1×K1\times K rows. For nonempty article sets, divide the coordinate sums by their respective article counts:

qg=∑a∈𝒜gza|𝒜g|,mpg=∑a∈𝒜pgza|𝒜pg|.q_g=\frac{\sum_{a\in\mathcal A_g}z_a}{|\mathcal A_g|},\qquad m_{pg}=\frac{\sum_{a\in\mathcal A_{pg}}z_a}{|\mathcal A_{pg}|}.

The background is drawn from qualified candidate-associated articles, not from every article in the news corpus.

Write qE,1q_{E,1} for the first coordinate of cell E’s background row. It is

qE,1=10(2)+5(1)+5(0)+5(−1)25=0.8.q_{E,1}=\frac{10(2)+5(1)+5(0)+5(-1)}{25}=0.8.

Computing the other coordinates in the same way gives

qE=(0.8,0.6,0.8),qL=(0.25,0.75,0.25).q_E=(0.8,0.6,0.8),\qquad q_L=(0.25,0.75,0.25).

Now consider A. Its 15 E articles have mean (1.6667,0.3333,1.3333)(1.6667,0.3333,1.3333); its five L articles have mean (1,−1,2)(1,-1,2). Subtracting each cell’s background gives

mAE−qE=(0.8667,−0.2667,0.5333),m_{AE}-q_E=(0.8667,-0.2667,0.5333), mAL−qL=(0.75,−1.75,1.75).m_{AL}-q_L=(0.75,-1.75,1.75).

The question has changed from “what is A’s coverage like?” to “how does A’s coverage differ from its available comparison coverage in each cell?” A negative coordinate means relatively less in that direction, not an absence of articles or disagreement with a topic.

4.2.1 Keep the numerator and denominator visible

For E, add each batch row multiplied by its distinct-article count. The coordinate sums are (20,15,20)(20,15,20): N2 receives 10(0)+5(1)+5(2)+5(0)=1510(0)+5(1)+5(2)+5(0)=15, while N3 receives 10(1)+5(2)+5(2)+5(−2)=2010(1)+5(2)+5(2)+5(-2)=20. Divide all three sums by 25 to obtain qE=(0.8,0.6,0.8)q_E=(0.8,0.6,0.8). For L the sums are (5,15,5)(5,15,5) and the count is 20, giving qL=(0.25,0.75,0.25)q_L=(0.25,0.75,0.25).

A’s E numerator includes only E1 and E2: 10(2,0,1)+5(1,1,2)=(25,5,20)10(2,0,1)+5(1,1,2)=(25,5,20). Divide by 15 to obtain (5/3,1/3,4/3)(5/3,1/3,4/3). Its L numerator is 5(1,−1,2)=(5,−5,10)5(1,-1,2)=(5,-5,10); division by five returns (1,−1,2)(1,-1,2). Subtract corresponding backgrounds one coordinate at a time. In E the first contrast is 5/3−4/5=25/15−12/15=13/155/3-4/5=25/15-12/15=13/15. The other two are 1/3−3/5=−4/151/3-3/5=-4/15 and 4/3−4/5=8/154/3-4/5=8/15. These exact fractions explain the rounded contrast row above.

Each denominator counts the observations named by that numerator. The background counts unique eligible articles in a cell. The personal mean counts that person’s associated articles within that cell. Neither denominator counts the number of named people in the article.

4.3 Give supported cells equal weight

A has three times as many articles in E as in L. If we pooled the articles, E would dominate. Let 𝒢p\mathcal G_p be the set of cells with enough articles for person pp; |𝒢p||\mathcal G_p| counts retained cells. In production, a cell needs at least five articles associated with that person. For a nonempty retained-cell set, let rpr_p be the 1×K1\times K balanced contrast row. The implemented aggregation computes it by averaging the supported cell contrasts equally:

rp=1|𝒢p|∑g∈𝒢p(mpg−qg).r_p=\frac1{|\mathcal G_p|}\sum_{g\in\mathcal G_p}(m_{pg}-q_g).

In the toy example both cells qualify for every person. A’s average contrast is therefore

rA=(0.8083,−1.0083,1.1417).r_A=(0.8083,-1.0083,1.1417).

For comparison, weighting A’s two contrasts in the article-count ratio 15:5 would give (0.8375,−0.6375,0.8375)(0.8375,-0.6375,0.8375). Neither averaging rule is a law of nature. They encode different questions. Equal supported-cell weighting prevents a densely covered cell from receiving greater weight merely because it contains more articles.

For A, compute a mean in each cell, subtract the cell background, and average the two contrasts. Equal-cell and article-count weighting estimate different quantities.

Equal cell weights are not necessarily equal source weights. A source spanning three supported decades contributes three cells, while another source spanning one contributes one. Nor does background subtraction remove every archive bias. It changes a specified comparison, within the selected coverage that is available.

The production gates require at least five articles in a person-cell and at least ten across the retained cells for a person-window. The contrast must also have nonzero numerical length. A cell with four articles is not quietly padded, and an unsupported profile is missing rather than a zero vector. These are operational support rules, not a proof of statistical reliability.

4.3.1 Average contrasts, then check eligibility

For A, the first balanced coordinate is (13/15+3/4)/2=(52/60+45/60)/2=97/120(13/15+3/4)/2=(52/60+45/60)/2=97/120. The other two are (−4/15−7/4)/2=−121/120(-4/15-7/4)/2=-121/120 and (8/15+7/4)/2=137/120(8/15+7/4)/2=137/120. Thus the exact contrast is (97,−121,137)/120(97,-121,137)/120. A denominator of two gives one vote to each retained cell.

Article weighting gives E the weight 15/20=3/415/20=3/4 and L the weight 5/20=1/45/20=1/4. Its first coordinate is (3/4)(13/15)+(1/4)(3/4)=0.65+0.1875=0.8375(3/4)(13/15)+(1/4)(3/4)=0.65+0.1875=0.8375. The difference comes from weights, not from any changed article coordinate.

The order of the support gates matters. Counts of six and four sum to ten, but the four-article cell is removed first. Only six retained articles remain, so that person-window fails. Counts of fifteen and four leave one retained cell with fifteen articles; its contrast alone supplies the average. Counts of fifteen and five retain both cells and twenty supporting articles. Missing means have no invented zeros. The notebook checks all three cases explicitly.

4.4 Keep direction and report support separately

Let bpb_p denote person pp’s 1×K1\times K unit profile. For a nonzero contrast, normalize it as follows:

bp=rp‖rp‖.b_p=\frac{r_p}{\lVert r_p\rVert}.

For A the result is approximately (0.4688,−0.5847,0.6621)(0.4688,-0.5847,0.6621). Multiplying rAr_A by a positive constant would produce the same bAb_A. This makes directions comparable without turning article volume into profile length. It also discards magnitude. A small, unstable contrast can become a full-length unit vector, which is one reason to display support counts and eventually evaluate uncertainty rather than trusting the norm alone.

Let BB be the matrix formed by stacking those unit profiles. In the toy example its four rows, in person order A, B, C, D, are

B≈(0.4688−0.58470.66210.94840.3161−0.0243−0.21830.75900.6134−0.6882−0.2294−0.6882).B\approx\begin{pmatrix} 0.4688&-0.5847&0.6621\\ 0.9484&0.3161&-0.0243\\ -0.2183&0.7590&0.6134\\ -0.6882&-0.2294&-0.6882 \end{pmatrix}.

Each row describes one person’s direction of relative coverage in the fixed news space. We have not yet learned a people pattern.

Check 4. Why would counting E1 once for A and once for B inside the background change the question? Why is it nevertheless correct for E1 to enter both individual profiles?

4.4.1 Normalize A without hiding the scale

The exact balanced contrast has squared length

‖rA‖2=972+1212+13721202=9409+14641+1876914400=4281914400.\begin{aligned} \lVert r_A\rVert^2&=\frac{97^2+121^2+137^2}{120^2}\\ &=\frac{9409+14641+18769}{14400}\\ &=\frac{42819}{14400}. \end{aligned}

Its length is 42819/120\sqrt{42819}/120, approximately 1.72440. Dividing each entry by that length cancels the common denominator, so bA=(97,−121,137)/42819b_A=(97,-121,137)/\sqrt{42819}. Its squared coordinates add to 42819/42819=142819/42819=1.

This cancellation explains both the convenience and the loss. We can recover the direction from the normalized row, but we cannot determine whether the contrast originally had length 1.72440, twice that length, or half that length. Nor does the unit row remember that A had twenty supporting articles. Counts, magnitude, and article links must remain separate records.

5 A population of profiles makes a people basis

5.1 Center the rows around the population

Let nn now count the fitted people, so BB has shape n×Kn\times K. Its rows have unit length, but they need not average to zero. Denote their 1×K1\times K mean row by b‾\bar b. In the example it is

b‾≈(0.1276,0.0652,0.1407).\bar b\approx(0.1276,0.0652,0.1407).

Let 𝟏\mathbf1 denote a column of nn ones, so 𝟏b‾\mathbf1\bar b repeats the mean row for every person. Subtract it from BB and call the resulting n×Kn\times K centered matrix HH:

H=B−𝟏b‾.H=B-\mathbf1\bar b.

Let hp=bp−b‾h_p=b_p-\bar b denote person pp’s row of HH. For A,

hA=bA−b‾≈(0.3411,−0.6500,0.5213).h_A=b_A-\bar b\approx(0.3411,-0.6500,0.5213).

This second centering asks how A’s normalized coverage differs from the fitted population of normalized person profiles. The earlier subtraction compared articles inside a cell. The reference objects are different.

5.2 Covariance asks which coordinates move together

Let CPC_P denote the K×KK\times K people covariance. Using the fitted-person count nn, compute it as

CP=1nH𝖳H,C_P=\frac1nH^{\mathsf T}H,

Here jj and kk each index news coordinates from 1 to KK. Entry (j,k)(j,k) is the average product of those two centered coordinates across people. A large positive product tends to occur when both coordinates deviate in the same direction; a negative product occurs when they deviate oppositely. The diagonal entries are coordinate variances.

Each person contributes one row. A person supported by 10,000 articles does not enter the covariance a thousand times more heavily than a person supported by ten. More articles can affect the quality of a row, but they are not its weight in this fit. This is equal-person PCA: principal component analysis of normalized, centered coverage profiles.

Let ℓ\ell index a people direction, wℓw_\ell its K×1K\times1 unit eigenvector column, and eℓe_\ell its covariance eigenvalue. Each pair satisfies

CPwℓ=eℓwℓ.C_Pw_\ell=e_\ell w_\ell.

Multiplying the direction by the covariance stretches it by eℓe_\ell without turning it. As explained in the Eigen Times eigenvector chapter, the largest eigenvalue identifies the direction of greatest variation, and the next identifies the greatest variation perpendicular to it.

5.2.1 Read one matrix entry as an average

For the original four-person fixture, the first centered coordinate is approximately 0.3411 for A, 0.8207 for B, -0.3459 for C, and -0.8159 for D. The N1 variance therefore adds their four squares and divides by four, giving approximately 0.393847. The N1–N3 covariance instead multiplies each person’s first and third centered entries and averages the four products, giving approximately 0.138797. Both calculations use one product per person. The notebook prints the complete centered table and checks the covariance against independent reference entries.

5.2.2 Verify an eigenvector with two coordinates

Set the four-person matrix aside for a small covariance illustration. Let C*=(2112)C_*=\begin{pmatrix}2&1\\1&2\end{pmatrix}. The star marks this separate synthetic table. Let a+=(1,1)𝖳/2a_+=(1,1)^{\mathsf T}/\sqrt2 and a−=(1,−1)𝖳/2a_-=(1,-1)^{\mathsf T}/\sqrt2 be unit columns. Multiplying the first gives (3,3)𝖳/2=3a+(3,3)^{\mathsf T}/\sqrt2=3a_+. Multiplying the second gives (1,−1)𝖳/2=a−(1,-1)^{\mathsf T}/\sqrt2=a_-. They are eigenvectors with eigenvalues 3 and 1; their dot product is (1−1)/2=0(1-1)/2=0.

For a unit column aa, the scalar a𝖳C*aa^{\mathsf T}C_*a is variance along that direction. Along the first coordinate axis it is 2. Along a+a_+ it is 3; along a−a_- it is 1. To see why 3 is the maximum, express any unit direction as a=γa++δa−a=\gamma a_++\delta a_-, where γ,δ\gamma,\delta are scalar coefficients satisfying γ2+δ2=1\gamma^2+\delta^2=1. Its variance is 3γ2+δ2=1+2γ23\gamma^2+\delta^2=1+2\gamma^2, no larger than 3. The larger-eigenvalue direction captures the most spread because of this numerical property, not because its words are necessarily more important.

Multiplying an eigenvector by -1 leaves its line and eigenvalue unchanged. Its coordinates and resulting scores both reverse signs. A sign convention is useful for reproducible displays; it does not change the represented variation.

5.3 Keep two patterns in the toy world

The toy covariance has eigenvalues approximately 0.49700.4970, 0.28170.2817, and 0.18100.1810. Keeping the first two retains

0.4970+0.28170.4970+0.2817+0.1810≈0.8114\frac{0.4970+0.2817}{0.4970+0.2817+0.1810}\approx0.8114

of the between-person variance, or 81.14%. This is a reconstruction statement about this four-person cloud. It is not an accuracy estimate, an identity confidence, or the fraction of articles explained.

Let rr count the retained positive people directions; it is 2 in this toy example. The bare rr is a count, distinct from the contrast row rpr_p. The retained direction index ℓ\ell runs from 1 to rr. Put those eigenvectors, ordered by decreasing eigenvalue, into the columns of the K×rK\times r loading matrix WW:

W≈(0.7964−0.35160.10900.88380.59490.3087).W\approx\begin{pmatrix} 0.7964&-0.3516\\ 0.1090&0.8838\\ 0.5949&0.3087 \end{pmatrix}.

Its rows correspond to N1, N2, N3; its columns correspond to people patterns P1 and P2. The columns have length one and are perpendicular. Their signs are oriented so the largest-magnitude news loading is positive, matching the implementation’s sign convention. Flipping a column and all its scores would describe the same geometry. When two eigenvalues are close, a small change in data can rotate their directions substantially even while their combined subspace changes little. Fixing signs does not remove that ambiguity; it is another reason to keep model versions explicit.

Unit person profiles are centered, their covariance is decomposed, and two retained directions give each person two scores. The discarded third eigenvalue is visible rather than silently erased.

5.3.1 Account for the omitted variance

The three fitted eigenvalues sum to approximately 0.9596456. This is also the sum of the three diagonal entries of CPC_P, called its trace: total centered squared length per person, expressed in either coordinate system. The first two sum to approximately 0.7786802. Their ratio is 0.8114248; multiplying by 100 expresses it as a percentage. The omitted fraction is about 0.1885752.

The loading columns are unit directions, but a row of WW need not have length one. Similarly, the score rows need not be unit rows after population centering and projection. Keeping these distinctions prevents a loading, a coordinate, and a cosine from being treated as interchangeable numbers.

5.4 Locate a person on the patterns

Let upu_p denote person pp’s 1×r1\times r people-coordinate row. Project its centered profile hph_p onto the retained columns:

up=hpW=(bp−b‾)W.u_p=h_pW=(b_p-\bar b)W.

For A, the first coordinate is approximately

uA1≈0.3411(0.7964)−0.6500(0.1090)+0.5213(0.5949)≈0.5109.\begin{aligned} u_{A1}&\approx0.3411(0.7964)-0.6500(0.1090)\\ &\quad+0.5213(0.5949)\\ &\approx0.5109. \end{aligned}

The second uses the second column of WW. The resulting rows are:

Person P1 score P2 score
A 0.5109 -0.5334
B 0.5828 -0.1178
C 0.0813 0.8807
D -1.1750 -0.2294

P1 is a pattern, not person D, although D has the largest absolute P1 score in this small example. A negative score identifies the other pole of a direction. It does not say that D is opposed to the other people.

In the published models, there are 60 news coordinates and at most 24 directions with positive eigenvalues. The 217-person Hacker News model retains 84.54% of between-person variance; the 61-person general-news model retains 90.86%. These separately fitted percentages should not be read as a contest between corpora. There is no second varimax naming step and no whitening of people scores. The people directions are orthonormal in standardized news coordinates, not generally when mapped back into the original embedding geometry.

5.4.1 Show both dot products

A’s first score adds approximately 0.2717−0.0709+0.3101=0.51090.2717-0.0709+0.3101=0.5109. Its second coordinate is

uA2≈0.3411(−0.3516)−0.6500(0.8838)+0.5213(0.3087)≈−0.1199−0.5745+0.1609≈−0.5334.\begin{aligned} u_{A2}&\approx0.3411(-0.3516)-0.6500(0.8838)\\ &\quad+0.5213(0.3087)\\ &\approx-0.1199-0.5745+0.1609\\ &\approx-0.5334. \end{aligned}

These are projections of the same centered row onto different direction columns. Small differences in final digits arise if we multiply displayed four-decimal entries rather than the full-precision fixture.

The first score column has variance 0.4970 and the second 0.2817, the eigenvalues already reported. They are not divided by the square roots of those variances in this model. Dividing them would whiten the retained scores and would change distances and cosines; that is a different geometry.

6 One loading table has two entrances

6.1 Start with a news axis

The first row of the toy WW is (0.7964,−0.3516)(0.7964,-0.3516). It tells us how N1 loads on P1 and P2. Starting from N1, the interface can rank people patterns by the absolute sizes of those two entries. P1 comes first, with a positive loading; P2 comes second, with a negative loading.

6.2 Start with a people pattern

The first column of WW is (0.7964,0.1090,0.5949)𝖳(0.7964,0.1090,0.5949)^{\mathsf T}. It tells us how P1 combines N1, N2, and N3. Starting from P1, the interface ranks news axes by the absolute sizes of these entries: N1, then N3, then N2.

The coefficient linking N1 and P1 is exactly the same number in both views. No second model is needed to construct the reverse index. One index exposes rows of WW; the other exposes columns.

For the diagram, let jj index news axes, ℓ\ell retained people patterns, pp people, and aa articles. The scalar WjℓW_{j\ell} is row jj, column ℓ\ell of WW; zajz_{aj} and bpjb_{pj} are coordinate jj of the article and person-profile rows; upℓu_{p\ell} is coordinate ℓ\ell of the people-score row. The diagram uses parentheses for these same indices: for example, W(j,ℓ)W(j,\ell) means WjℓW_{j\ell}. Magnitude means absolute value, so either sign can produce a large magnitude.

A row of W answers news-to-pattern questions. A column answers pattern-to-news questions. The signed coefficient is shared, while each ranking compares a different set.

This does not make ranks reciprocal. N3’s strongest loading is P1, but N3 is only the second strongest news loading of P1. Nor are the entries probabilities. Some are negative, and the absolute values in a row or column need not sum to one.

6.3 People can be ranked in two distinct ways

Starting from news axis N1, one can rank people by their direct profile coordinate |bp1||b_{p1}|. B comes first in the toy example, with 0.9484. This asks whose normalized relative coverage extends farthest along N1, considering either pole.

Starting from people pattern P1, one can instead rank people by |up1||u_{p1}|. D comes first, with 1.1750. This asks whose centered profile extends farthest along that learned combination of news coordinates.

Those rankings differ because the questions differ. Neither is the same as finding the nearest neighbor of a selected person, which uses the entire row of people scores. A useful interface names the selected axis, mode, pole, model, and time window so a reader knows which question produced the list.

The audited Hacker News model supplies a real example. The loading joining news axis N56 to people pattern P19 is +0.4144960629+0.4144960629 in both stored indexes. N56’s terms include develop, interview, databas, ceo, manag, and github. The shared coefficient is a precise navigation fact. Calling it “the database CEO axis” would discard the mixed meaning of the direction and the evidence needed to identify a current CEO.

Check 5. In the toy matrix, N3 ranks P1 first. Where does P1 rank N3? Explain why the two answers are consistent.

7 Returning through the map loses detail

7.1 The return operation

When discussing one person without naming them, write bb, hh, and uu for their unit profile bpb_p, centered profile hph_p, and retained score row upu_p. The rows b,hb,h have KK coordinates and uu has rr. Let ĥ\hat h be the 1×K1\times K reconstruction of hh; the hat denotes an approximation returned from retained coordinates. After calculating u=hWu=hW, multiply by the transpose:

ĥ=uW𝖳=hWW𝖳.\hat h=uW^{\mathsf T}=hWW^{\mathsf T}.

Adding b‾\bar b gives the corresponding reconstruction of the unit news profile. It does not give back the original articles.

An orthogonal projector keeps a vector’s component in a chosen subspace and discards its perpendicular component. Let PP be the K×KK\times K forward-and-return matrix, defined by P=WW𝖳P=WW^{\mathsf T}. A symmetric matrix equals its transpose; an idempotent matrix has the same effect when applied twice as once. Because the columns of WW are orthonormal, PP has both properties:

P𝖳=P,P2=W(W𝖳W)W𝖳=P.P^{\mathsf T}=P,\qquad P^2=W(W^{\mathsf T}W)W^{\mathsf T}=P.

The first equality expresses symmetry and the second idempotence. Together these identify an orthogonal projector. Once a row lies in the retained plane, projecting it onto that plane again changes nothing.

The residual is the difference between the original centered row and its reconstruction, h−ĥh-\hat h. It is perpendicular to the retained directions. Consequently,

‖h‖2=‖ĥ‖2+‖h−ĥ‖2.\lVert h\rVert^2=\lVert \hat h\rVert^2+\lVert h-\hat h\rVert^2.

This is the familiar right-triangle rule, now applied to components of a vector. In the fitted toy example, A’s discarded squared length is approximately 0.2650. The example script checks the decomposition and the orthogonality of the residual using unrounded values.

Projecting a centered profile onto two people directions and returning reconstructs only its component in their plane. The residual is a measurable part of the original profile.

7.2 An even smaller exact example

To see the loss without decimals, set aside the fitted toy matrix for a moment. Let W*W_* denote this deliberately chosen 3×23\times2 matrix; the star distinguishes it from the fitted WW:

W*=(1/201/2001).W_* = \begin{pmatrix}1/\sqrt2&0\\1/\sqrt2&0\\0&1\end{pmatrix}.

This is an illustration, not another fit. It retains the shared direction of the first two coordinates and retains the third coordinate separately. A query h=(1,0,0)h=(1,0,0) goes forward to (1/2,0)(1/\sqrt2,0) and returns as (1/2,1/2,0)(1/2,1/2,0). The first-versus-second distinction was discarded.

A linear map preserves addition and scaling; multiplication by a fixed matrix is an example. An adjoint transfers a linear map to the other side of a dot product; for real matrices with Euclidean dot products, it is the transpose. Least squares means minimizing the sum of squared coordinate discrepancies. A Moore-Penrose pseudoinverse is a generalized inverse that gives least-squares solutions, choosing the shortest solution when several are possible. Because WW has orthonormal columns, W𝖳W^{\mathsf T} is both its adjoint and its pseudoinverse. It returns the least-squares reconstruction of a centered profile in the retained subspace. A two-sided inverse would undo the map in both directions for every input; this transpose cannot do that. With 24 retained directions out of 60, there is no way to recover every possible original 60-coordinate row.

7.2.1 Compute the residual and the shortest reconstruction

The first retained coordinate of h=(1,0,0)h=(1,0,0) is 1(1/2)+0(1/2)+0(0)=1/21(1/\sqrt2)+0(1/\sqrt2)+0(0)=1/\sqrt2; the second is zero. Returning multiplies the first column by 1/21/\sqrt2, giving (1/2,1/2,0)(1/2,1/2,0). Subtraction leaves (1/2,−1/2,0)(1/2,-1/2,0). Its squared length is 1/4+1/4=1/21/4+1/4=1/2, and its dot products with both retained columns are zero. The returned part also has squared length 1/21/2, so 1=1/2+1/21=1/2+1/2.

Any alternative reconstruction inside the retained plane can be written (t,t,s)(t,t,s), where tt and ss are real scalar choices. Its squared discrepancy from (1,0,0)(1,0,0) is

(1−t)2+t2+s2=2(t−1/2)2+s2+1/2.\begin{aligned} (1-t)^2+t^2+s^2 &=2(t-1/2)^2+s^2+1/2. \end{aligned}

Squares cannot be negative, so the minimum is 1/21/2, achieved at t=1/2t=1/2 and s=0s=0. This proves the least-squares claim for this example without asking the reader to trust an inverse formula.

7.3 Loss began before this projection

Even retaining every people direction would not recover an article list from a person profile. The earlier pipeline also reduced a 384-coordinate embedding to 60 news coordinates, averaged articles, subtracted backgrounds, and normalized lengths. Many different inputs can give the same output after those operations.

Navigation is valuable without being invertible. It provides linked views of a representation and routes back to stored evidence. The article links and factual sources are retained separately precisely because the vector cannot reconstruct them.

7.4 Why the singular value decomposition gives two views

Singular value decomposition (SVD) factors a matrix into two sets of orthonormal directions and nonnegative scale factors called singular values. For a longer prerequisite, revisit the Eigen Times SVD chapter. The operator min\min selects the smaller of two numbers. For the centered n×Kn\times K matrix HH, let k*=min⁡(n,K)k_* = \min(n,K). In the thin decomposition, UU is an n×k*n\times k_* matrix of left directions, indexed by people, and QQ is a K×k*K\times k_* matrix of right directions, indexed by news coordinates. Both have orthonormal columns. Let Σ\Sigma be the k*×k*k_*\times k_* diagonal matrix of singular values in descending order. Then

H=UΣQ𝖳.H=U\Sigma Q^{\mathsf T}.

Recall that rr counts retained positive people directions. Let UrU_r and QrQ_r contain the first rr columns of their respective matrices, and let Σr\Sigma_r contain the leading r×rr\times r diagonal block of Σ\Sigma. Thus UrU_r has shape n×rn\times r and QrQ_r shape K×rK\times r. Using the same ordering and orientation as the people eigendecomposition gives W=QrW=Q_r and

HW=UrΣr.HW=U_r\Sigma_r.

The scores include the singular values; they are not just the columns of UrU_r. Each squared singular value divided by nn equals the corresponding eigenvalue of the earlier people covariance. The right directions come from H𝖳HH^{\mathsf T}H, while the left directions come from HH𝖳HH^{\mathsf T}. These are two sides of the same centered profile matrix. An adjacency matrix instead records which graph nodes have edges between them. Neither calculation uses that matrix from Anthropology’s relationship graph. A graph community is a group of nodes relatively densely linked to one another. Communities and people patterns can be interesting to compare, but they are not the same construction.

Check 6. For W*W_* above, project h=(0,1,0)h=(0,1,0) forward and back. Compare its returned row with the return from (1,0,0)(1,0,0). What information can the map no longer distinguish?

7.4.1 A complete SVD that fits on a page

Use a separate centered table H*H_* with four rows (1,0)(1,0), (−1,0)(-1,0), (0,2)(0,2), and (0,−2)(0,-2). Each column sums to zero. Multiplication gives H*𝖳H*=diag⁡(2,8)H_*^{\mathsf T}H_*=\operatorname{diag}(2,8): the first column’s squares sum to two, the second’s to eight, and their paired products are zero.

The largest right direction is therefore (0,1)𝖳(0,1)^{\mathsf T}, followed by (1,0)𝖳(1,0)^{\mathsf T}. The singular values are the square roots of eight and two: 8\sqrt8 and 2\sqrt2. Projecting the rows onto the first right direction gives (0,0,2,−2)𝖳(0,0,2,-2)^{\mathsf T}. Divide that column by 8\sqrt8 to obtain its unit left direction (0,0,1/2,−1/2)𝖳(0,0,1/\sqrt2,-1/\sqrt2)^{\mathsf T}. The second left direction is (1/2,−1/2,0,0)𝖳(1/\sqrt2,-1/\sqrt2,0,0)^{\mathsf T}.

Now multiply back: the first left column times 8\sqrt8 times the first right row restores only the second coordinate; the second term restores only the first coordinate. Adding the two restores every entry of H*H_*. With four observations, the covariance eigenvalues are 8/4=28/4=2 and 2/4=0.52/4=0.5. Keeping only the first singular direction retains 8/(8+2)=0.88/(8+2)=0.8 of the total squared length. The discarded first-coordinate column has squared length two. This shows explicitly why scores contain singular values: the left direction has unit length, while its score column here has length 8\sqrt8.

The native OCaml teaching routine obtains a small SVD through a symmetric Gram matrix, the product H*𝖳H*H_*^{\mathsf T}H_*. That is adequate for these tiny well-scaled demonstrations; it is not presented as a numerically preferred algorithm for a large production fit. Both notebooks verify reconstruction, orthonormality, and the independent singular values.

8 Finding neighbors without inventing relationships

8.1 Compare complete rows in one coordinate system

Here pp and qq index two people, and up,uqu_p,u_q are their nonzero 1×r1\times r score rows from the same model and coverage window. The person index qq is distinct from a background row such as qgq_g. Cosine similarity, written cos⁡(up,uq)\operatorname{cos}(u_p,u_q), is their dot product divided by the product of their lengths:

cos⁡(up,uq)=up⋅uq‖up‖‖uq‖.\operatorname{cos}(u_p,u_q)=\frac{u_p\cdot u_q}{\lVert u_p\rVert\,\lVert u_q\rVert}.

It compares direction rather than length and is undefined if either row is zero. A value near 1 means similarly directed rows, 0 means perpendicular rows, and -1 means opposite directions in the specified centered representation. Negative similarity is not evidence of competition, disagreement, or hostility.

Anthropology’s neighbor view uses complete retained people-score rows within the same model and window. Do not mix a recent profile from one model with an all-history profile from another and call their raw dot product a meaningful similarity. Even matching dimensions are insufficient: the coordinates must share their meanings.

8.1.1 Coordinate closeness and directional closeness

Consider three synthetic two-coordinate score rows: a query a=(1,0)a=(1,0) and candidates b=(10,0)b=(10,0) and c=(1,1)c=(1,1). Coordinate distance means the length of the difference row. From aa to bb the difference is (−9,0)(-9,0) and the distance is 9. From aa to cc it is (0,−1)(0,-1) and the distance is 1. Thus coordinate distance prefers cc.

Cosine asks another question. The dot product of aa and bb is 10, their lengths are 1 and 10, and the cosine is 10/(1⋅10)=110/(1\cdot10)=1. The dot product of aa and cc is 1, their lengths are 1 and 2\sqrt2, and the cosine is 1/21/\sqrt2, approximately 0.7071. Directional similarity prefers bb. No contradiction exists: bb lies farther out along exactly the same direction.

If both rows are normalized, squared coordinate distance becomes twice one minus cosine. To check this, expand the squared difference into the two squared lengths minus twice the dot product. Each squared length is then one. For aa and normalized cc, this gives 2−2/22-2/\sqrt2, approximately 0.5858. This equivalence requires unit rows in the same geometry; the original people-score rows are not generally unit rows.

A single selected coordinate can be identical while all the other coordinates differ. Ranking by one axis therefore cannot replace a full-row neighbor calculation. Centering also changes the origin from which direction is measured, so a cosine of uncentered profiles answers a different question from a cosine of their centered people scores.

8.2 Compression can change the answer substantially

The toy example makes the danger visible. A and B have a cosine of only 0.0392 between their full centered three-coordinate profiles. Their cosine becomes 0.8211 after projection into the two retained people directions.

No arithmetic error occurred. The discarded direction contained much of their difference. The retained plane makes them look more alike. Keeping 81.14% of the population’s total variance did not preserve every pairwise angle.

Cosines before and after retaining two people directions. A and B demonstrate why a strong reduced-space neighbor score needs a full-space baseline.

The same check is useful in the real model. In the frozen Hacker News model, LeCun and Bengio have a retained-space cosine of 0.9587 and a full centered news-space cosine of 0.9118. Ellison and Siebel have 0.7198 and 0.5800 respectively, with 416 and 11 candidate articles. These values describe this corpus and representation; the much smaller support for one member deserves attention.

An appealing analogy can also fail. Zuckerberg and Gates have retained-space cosine 0.0782 in Hacker News and -0.1137 in the separate general-news model. Being prominent technology founders does not require the archives to cover them in the same way. A research system should let the evidence correct the analogy.

8.3 What a neighbor score leaves unresolved

Two profiles can be close because the same articles mention both people, because different articles discuss similar subjects, or because the archive repeatedly frames them in a similar way. A cosine alone cannot distinguish these explanations. One useful sensitivity analysis would remove their shared articles and recompute the comparison. That experiment was proposed in the paper; it was not part of the published evaluation.

For a nonzero centered row hh, retained energy is the fraction of its squared length preserved by projection, ‖hW‖2/‖h‖2\lVert hW\rVert^2/\lVert h\rVert^2. Support counts, full-space baselines, retained energy, dates, and example articles help a reader interpret a score. Calibration means agreement between predicted probabilities and observed frequencies. These aids do not turn a neighbor score into a calibrated probability of collaboration. A documented relationship still requires its own source. The named examples here demonstrate navigation, not a statistically representative assessment of retrieval quality.

Check 7. Does retaining 81.14% of total variance guarantee that every neighbor cosine changes by less than 18.86 percentage points? Use A and B to test the claim.

9 People can move while the map stays fixed

9.1 Reuse the ruler

A person can receive different coverage this year than ten years ago. We want to observe that movement without changing the ruler at the same time. Let ww index a coverage window, such as a decade or a recent period. Write bpwb_{pw} for person pp’s 1×K1\times K unit profile built from that window’s eligible articles, and upwu_{pw} for its 1×r1\times r people-coordinate row. A comma in up,wu_{p,w} separates the same indices without changing their meaning. Fit the people mean b‾\bar b and loading matrix WW once on the all-history profiles, then use them for every supported window:

upw=(bpw−b‾)W.u_{pw}=(b_{pw}-\bar b)W.

The article associations, supported cells, cell means, and cell backgrounds are recomputed for that window. Write mpgwm_{pgw} for person pp’s cell-gg mean in window ww, and qgwq_{gw} for that cell’s background mean, both 1×K1\times K rows. They use the same averaging rule as before, restricted to articles in the window. The people mean and directions remain the fitted all-history reference.

This makes score differences interpretable within a model version. Let w1w_1 and w2w_2 be two supported windows for the same person, and let Δp\Delta_p denote the Euclidean distance between that person’s two score rows. A simple movement statistic is

Δp=‖up,w2−up,w1‖.\Delta_p=\lVert u_{p,w_2}-u_{p,w_1}\rVert.

It is a distance between two supported coverage profiles. It is not, by itself, a measure of a person’s career change. A source mix can change, a name attribution can be corrected, or the comparison background can shift.

9.2 Separate movement from a change of mixture

Return to A’s two cell contrasts. Hold both contrasts fixed, but consider two synthetic count mixtures: 90:10 and 10:90. These are a sensitivity experiment, not two additional observed periods in the 45-article corpus. In each mixture both cells satisfy the support threshold.

Weighting the contrasts by those counts, normalizing, and applying the same fitted mean and WW gives people scores of approximately (0.696,−0.459)(0.696,-0.459) in the first mixture and (0.389,−0.547)(0.389,-0.547) in the second. Their apparent movement is about 0.319. Yet neither cell-specific contrast changed.

Equal supported-cell weighting keeps A at its original people coordinates (0.5109,−0.5334)(0.5109,-0.5334) in both mixtures. This illustrates one composition effect the balancing rule can reduce; it does not establish invariance to changes in the actual within-cell article distributions.

A change in article mixture can move an article-weighted profile even when its within-cell contrasts are fixed. Equal-cell weighting removes this particular movement, provided both cells remain supported.

It does not remove every composition effect. If a cell falls below the support threshold, the retained set changes. If qgwq_{gw} changes, a person’s contrast changes even if its own articles do not. Two points can be drawn in one coordinate frame while still differing in their comparison backgrounds.

9.3 Three clocks and a model version

An article’s publication date is one clock. The newest matched article used in a historical profile is another. The valid date of a documented role is a third. A fourth piece of context, the model version, identifies the numerical ruler rather than a date of an event.

The projection watermark records the latest article date represented in the local projected corpus. In the frozen people snapshot, matched profiles end on 4 September 2026 and the local projection watermark is 5 September. The current-news overlay is dated 1 October. Recent-36-month and latest-90-day profile windows are anchored to the builder’s latest matched article date, not automatically to the reader’s present clock. Publishing a larger identity catalog on 2 October did not refit those profiles or update every watermark.

An absent recent profile means insufficient supported coverage. It must not be plotted at the origin as if it were evidence that the person became average.

9.4 How a future incremental refresh could work

For each person-cell-window, let NpgwN_{pgw} count its distinct eligible associated articles, and let SpgwS_{pgw} be the 1×K1\times K sum of their standardized rows zaz_a. For Npgw>0N_{pgw}>0, the previously defined person-cell mean is mpgw=Spgw/Npgwm_{pgw}=S_{pgw}/N_{pgw}. These SS sums are unrelated to the singular-value matrix Σ\Sigma. Keep a separate deduplicated count and sum for the cell background. New qualified articles add to these statistics; expired rolling-window articles subtract from them; corrected attributions remove an old contribution and add the corrected one.

For these mean updates, the count and coordinate sum are sufficient summaries: they retain everything needed to recompute the mean without rereading unchanged articles. The Eigen Times streaming chapter explains this pattern more generally. The added complication here is the shared background: changing one background can require refreshing every person who uses that cell. Updating only the newly mentioned person would leave inconsistent comparisons.

After updating the means, reapply support gates, subtraction, equal-cell averaging, normalization, and the fixed projection. For one resulting supported window, abbreviate its centered profile as hp=bpw−b‾h_p=b_{pw}-\bar b, suppressing the window index for this calculation. Let QpQ_p denote its discarded squared length; this scalar is distinct from the SVD’s right-direction matrix QQ. Define the residual statistic as

Qp=‖hp−hpWW𝖳‖2.Q_p=\lVert h_p-h_pWW^{\mathsf T}\rVert^2.

It measures how much of that centered profile lies outside the retained people subspace. A sustained rise could motivate evaluating a new basis. It is not an automatic alarm that a person changed, and this QpQ_p uses people-profile geometry rather than silently inheriting every threshold from article-level Eigen Times statistics.

This refresh algorithm is proposed. The deployed latest-news overlay does not perform it. New versions should preserve their predecessors and be evaluated on later observations kept out of fitting, with adequate support and source-sensitive uncertainty checks before making claims about change.

Check 8. Can a person’s profile change when no new article has been associated with that person? Identify two ways the aggregation can produce that result.

9.4.1 Update a count and a sum separately

Cell E’s unique background has count 25 and coordinate sum (20,15,20)(20,15,20). Suppose one newly eligible article has row (0,0,0)(0,0,0) and is associated with an additional person outside the four fitted toy profiles. The background still consists of candidate-associated articles; none of A–D’s personal article sets changes. The count becomes 26 while the sum stays (20,15,20)(20,15,20), so the background becomes (20/26,15/26,20/26)(20/26,15/26,20/26), approximately (0.76923,0.57692,0.76923)(0.76923,0.57692,0.76923).

Every existing profile using E now subtracts that new background. It can move despite having no newly associated article. Removing the same article restores the count to 25 and the original background exactly. A’s personal E sum is separately (25,5,20)(25,5,20). Correcting one attribution by replacing (2,0,1)(2,0,1) with (0,2,2)(0,2,2) changes it to (23,7,21)(23,7,21) while leaving its count at 15. Addition, expiration, and attribution correction are different changes to the stored summaries.

10 Relevance and significance answer different questions

10.1 Direction is not attention

A topic coordinate tells us how strongly an article lies along a direction. Eigen significance supplies another scalar: a story’s attention score, based on the original story-ranking method. Anthropology keeps that signal alongside the vector rather than silently using it as a weight in the equal-person PCA fit.

The source’s story score is allocated to its member articles and normalized by the scored source’s total in the relevant period. Let αa\alpha_a be article aa’s resulting nonnegative attention share. Recall that zajz_{aj} is coordinate jj of the standardized article row zaz_a. Let rank⁡j(a)\operatorname{rank}_j(a) denote an attention-weighted topic score, not an ordinal position in a list. One selectable latest-article ranking uses

rank⁡j(a)=|zaj|αa.\operatorname{rank}_j(a)=|z_{aj}|\alpha_a.

The absolute coordinate measures topic strength on either pole; αa\alpha_a measures allocated attention. The product deliberately answers a combined question. A reader can also choose topic-only ordering or inspect one pole separately.

For a tiny illustration, suppose article X has topic strength 2 and attention share 0.01, while article Y has topic strength 0.5 and attention share 0.10. Topic-only ordering puts X first. The combined scores are 0.02 and 0.05, so attention-weighted ordering puts Y first. Nothing about the topic coordinates changed.

10.1.1 Allocate an attention score before multiplying

Use a separate synthetic story with attention score 12 and three member articles. An equal allocation rule would give each article 12/3=412/3=4 units. If that source’s scored total in the period is 40, each article’s attention share is 4/40=0.14/40=0.1. Their shares sum to 0.3, the story’s share 12/4012/40. This illustrates one declared allocation rule; it does not assert that every source allocates every story equally.

The topic product then has two separate inputs. An article with standardized topic coordinate -2 and attention share 0.1 receives an either-pole score |−2|(0.1)=0.2|-2|(0.1)=0.2. A positive-pole query would need an explicit sign restriction instead. Allocation, source normalization, pole selection, and multiplication each change the question, so each belongs in the ranking’s definition.

10.2 Person attention shares can overlap

If a 0.10-share article mentions A and B, that coverage can contribute to both people’s attention histories. Their combined credited share can exceed the article’s 0.10. These histories describe overlapping coverage, not a partition of total human importance.

Source normalization matters too. Averaging shares over active sources, including sources with no mention of a person, differs from pooling raw scores from a large and a small outlet. State the denominator before interpreting the percentage. The Eigen Times discussion of story energy and ranking provides the deeper article-level background; the person layer adds attribution and aggregation rather than redefining that original significance.

10.2.1 Compare two denominators numerically

Suppose two active sources have total attention scores 100 and 1,000. The selected person’s associated coverage receives 10 units in the first and zero in the second. Source-specific shares are 10/100=0.1010/100=0.10 and 0/1000=00/1000=0. Giving each source one vote yields (0.10+0)/2=0.05(0.10+0)/2=0.05. Pooling the raw scores instead gives (10+0)/(100+1000)=1/110(10+0)/(100+1000)=1/110, approximately 0.00909. Omitting the zero-mention source would produce yet another quantity, 0.10.

None of these denominators can be recovered merely by reading a displayed percentage. State the active source set, period, attribution rule, and aggregation rule. The two-person credit on one 0.10-share article likewise totals 0.10+0.10=0.200.10+0.10=0.20 because the histories overlap; it does not create another article or more source attention.

10.3 Today is a query with evidence requirements

The natural question “show today’s database news against today’s database CEOs” joins several tasks. News coordinates find topic-relevant articles. Dated role assertions identify people recorded as CEOs at the requested time. Supported person profiles locate those people in a historical coverage geometry. Qualified article-person associations establish which current articles actually concern them.

The current overlay can show relevant articles alongside historical profiles, but the audited overlay has zero newly qualified article-person joins. Proximity between an article and a profile is therefore a lead to inspect, not proof that the article concerns that person. Similarly, an old article calling someone a CEO does not establish that the role is current.

The Hacker News people model remains in its hn1 basis. New article embeddings can be measured in that fixed basis while carrying a scalar attention share from hn2. This is coherent only when the separation is explicit: the scalar is an overlay; it does not substitute a different set of coordinate axes. Model, corpus, dates, support, and role evidence remain visible parts of the query.

Check 9. Article X is more topic-aligned than article Y. Must X rank higher after attention weighting? Which separate evidence would be needed before calling either one news about a selected person?

11 Shared concepts across three sites

An eigenvector is a direction learned from variation. An ontology concept is a named idea that people agree to reuse. “Databases” can remain a useful concept while a database-related news direction changes between archives. Linking the two would let a reader follow one subject through Eigen Times, Eigen Hacks, and Anthropology without pretending that the sites share one coordinate system.

This chapter separates three stages: the personal tags already implemented in the 2 October 2026 paper edition; a proposed, reviewed table connecting concepts to axes; and a proposed learned model whose predictions would require evaluation.

11.1 A tag, an axis, and a hierarchy

The ontology picker lets a reader choose a concept and attach its identifier to a canonical node. Searching a label and clicking it records an annotation. It does not rotate the news basis, retrain a person vector, or establish a public factual relationship.

The pinned ontology has 157 concepts in 11 areas; 61 concepts have multiple parents. This is a polyhierarchy: one concept can belong under several broader ideas. For an invented illustration, “distributed databases” might belong under both “databases” and “distributed systems.” These parent links describe meaning. They need not form perpendicular directions, and a person can have several tags at once.

An axis label serves a different purpose. Its strongest stemmed terms help a reader interpret one fitted direction. A list containing databas, ceo, and github does not define a clean databases category. Moreover, the negative pole needs inspection too: it is not automatically “not databases.” The sign of an eigenvector is a convention; a reviewed meaning must be attached to the appropriate pole and model version.

For the mechanics behind named directions, return to From eigenvectors to named axes.

11.2 A small, reviewed bridge comes first

A crosswalk is a table of correspondences. Imagine a proposed row saying: in this particular Eigen Hacks basis, the positive pole of this axis has useful evidence for the databases concept. The row should retain the corpus, axis and sign, basis fingerprint, ontology version, representative articles, method, and review status.

The bridge should be many-to-many. Several axes may help retrieve database articles, and one mixed axis may connect to several concepts. A reviewer can leave a proposed correspondence unassigned. Shared concept identifiers then become destinations across the three sites, while each route retains its original numerical coordinates.

Concept-to-concept mappings also have different strengths: “exact,” “close,” “broader,” “narrower,” and “related” do different jobs in the W3C SKOS mapping vocabulary. An axis would first need a reviewed semantic interpretation before those conceptual distinctions could be used responsibly. A familiar-looking word is insufficient.

11.3 Learning a concept score, one multiplication at a time

A later model could learn from reviewed articles. Let mm count the concepts reviewers label, and let zz be one article’s 1×K1\times K standardized news row. Let AA be a K×mK\times m slope matrix: each slope measures how changing an input coordinate changes one concept’s linear score. Let β\beta be an m×1m\times1 intercept column supplying baseline scores at zero input. The notation ⊤\top is another form of the transpose symbol 𝖳\mathsf T. Denote the resulting 1×m1\times m prediction row by ss. A simple linear predictor multiplies inputs by slopes and adds the baseline:

s=zA+β⊤.s=zA+\beta^\top.

Each column of AA asks how the news coordinates contribute to one concept. Its predictions are scores; no probability interpretation has yet been established.

Consider an entirely illustrative two-coordinate, two-concept model:

z=(2,−2),A=(0.4000.4),β⊤=(0.5,0.5). \begin{aligned} z&=(2,-2),\\ A&=\begin{pmatrix}0.4&0\\0&0.4\end{pmatrix},\\ \beta^\top&=(0.5,0.5). \end{aligned}

The first prediction is 2(0.4)+(−2)(0)+0.5=1.32(0.4)+(-2)(0)+0.5=1.3. The second is 2(0)+(−2)(0.4)+0.5=−0.32(0)+(-2)(0.4)+0.5=-0.3. Nothing went wrong when these escaped the interval zero to one: they are scores from an unconstrained linear rule. We have not made a probability model.

To learn the coefficients, let NN count reviewed training articles, distinct from the fitted-person count nn used earlier. Stack their standardized rows into the N×KN\times K matrix ZZ and their matching concept labels into the N×mN\times m matrix YY. Row order agrees across the two matrices. For a binary label, an entry is one when reviewers judged the concept present and zero when they judged it absent. Missing judgments must remain missing, rather than becoming zeros.

A ridge fit balances squared prediction error against a penalty on large slopes. Let ρ≥0\rho\ge0 be the penalty strength. For a matrix, the Frobenius norm, denoted by double bars with subscript FF, is the square root of the sum of squared entries. Thus its square adds those squared entries. Here 𝟏\mathbf1 is a column of NN ones, so 𝟏β⊤\mathbf1\beta^\top repeats the baseline row for every article. The following objective assumes every target entry is observed; partially reviewed targets require a separately specified objective summing only observed errors. Choose AA and β\beta to minimize

‖ZA+𝟏β⊤−Y‖F2+ρ‖A‖F2. \lVert ZA+\mathbf1\beta^\top-Y\rVert_F^2 +\rho\lVert A\rVert_F^2.

The first term adds squared prediction errors; the second penalizes large slopes. The intercept is deliberately unpenalized. A held-out validation procedure, using articles reserved from parameter fitting, would choose ρ\rho rather than a convenient-looking result on the training articles.

The toy coefficients above can actually be fitted. Take four fictional, fully reviewed article rows:

Z=(−1−1−111−111),Y=(00011011). Z=\begin{pmatrix}-1&-1\\-1&1\\1&-1\\1&1\end{pmatrix},\qquad Y=\begin{pmatrix}0&0\\0&1\\1&0\\1&1\end{pmatrix}.

Each concept’s average label is 0.5, giving the intercept. Each coordinate column has sum of squares four and centered label-product sum two. With ρ=1\rho=1, its fitted slope is 2/(4+1)=0.42/(4+1)=0.4; the cross-slopes are zero. Training predictions are 0.1 or 0.9. The new row (2,−2)(2,-2) then extrapolates beyond the training coordinates, producing (1.3,−0.3)(1.3,-0.3). Training fit and behavior on new inputs are different tests.

An illustrative ridge fit: four invented article rows determine slopes and intercepts; an extrapolating row produces scores outside the probability interval. These are arithmetic examples, not Anthropology evaluation results.

Each corpus would need its own fitted coefficients. Shared target identifiers make the outputs comparable in meaning; they do not make the input bases interchangeable. Concepts overlap, so AA need not be orthogonal or even square. This semantic transformation does not inherit the reconstruction guarantees of WW.

11.3.1 Derive the ridge slope instead of memorizing it

A derivative describes how a scalar output changes per small change in a scalar input. Start with a function f(a)=a2f(a)=a^2, where aa is a real scalar, and let ε\varepsilon be a nonzero change in that input. The difference quotient is

f(a+ε)−f(a)ε=a2+2aε+ε2−a2ε=2a+ε.\frac{f(a+\varepsilon)-f(a)}{\varepsilon} =\frac{a^2+2a\varepsilon+\varepsilon^2-a^2}{\varepsilon} =2a+\varepsilon.

As ε\varepsilon approaches zero, this approaches 2a2a. That limiting slope is the derivative, written f′(a)f'(a). At a=2a=2, changes of 0.1 and 0.01 give difference quotients 4.1 and 4.01, approaching 4. A derivative is a local rate of change; it is not the function’s output, which also happens to equal 4 at this particular input.

Return to the first ridge target. Its centered labels are (−0.5,−0.5,0.5,0.5)(-0.5,-0.5,0.5,0.5), and the first input coordinate is (−1,−1,1,1)(-1,-1,1,1). Temporarily call its scalar slope aa and set the other slope to zero. Each of the four squared prediction errors is (a−0.5)2(a-0.5)^2. With penalty ρ=1\rho=1, the scalar objective, named J(a)J(a), is

J(a)=4(a−0.5)2+a2=5a2−4a+1.J(a)=4(a-0.5)^2+a^2=5a^2-4a+1.

Differentiate its expanded terms: J′(a)=10a−4J'(a)=10a-4. At an interior minimum a small move in either direction cannot improve the value, so the local slope must be zero. Solving 10a−4=010a-4=0 gives a=0.4a=0.4. Complete the square to verify this is a minimum: J(a)=5(a−0.4)2+0.2J(a)=5(a-0.4)^2+0.2. Its minimum value is 0.2. At a=0a=0 the value is 1; at a=0.5a=0.5 it is 0.25. Ridge accepts a little prediction error to reduce the coefficient penalty.

For any nonnegative penalty ρ\rho, this example instead has derivative 2(4+ρ)a−42(4+\rho)a-4, giving a=2/(4+ρ)a=2/(4+\rho). The input columns have zero cross-product, so the two slopes separate cleanly. In a general table they interact. Let ZcZ_c subtract each input column mean from ZZ, and let YcY_c subtract each target column mean from YY. Let II here be the K×KK\times K identity. The stationary equations for these centered tables are (Zc𝖳Zc+ρI)A=Zc𝖳Yc(Z_c^{\mathsf T}Z_c+\rho I)A=Z_c^{\mathsf T}Y_c. Solving this equation finds the slopes; restoring the target means minus the mean input’s fitted contribution gives the unpenalized intercepts.

The notation ∂J/∂a\partial J/\partial a denotes a partial derivative: change coefficient aa while holding the other coefficients fixed. The gradient collects all those partial derivatives in the coefficient table’s shape. For the full squared-error objective its slope gradient is 2[Zc𝖳(ZcA−Yc)+ρA]2[Z_c^{\mathsf T}(Z_cA-Y_c)+\rho A]. Setting the whole gradient to zero yields the same stationary equations. The notebooks check the scalar derivative with small finite changes, the completed-square minimum, and the fitted matrix equations independently.

11.4 A score becomes a probability only after another question

First define what is being predicted. One possible event is: “Under this annotation protocol, reviewers judge that this article substantially discusses databases.” It is not “this person is a database person.”

A logistic predictor maps a linear score through a smooth S-shaped function whose output lies between zero and one. Calibration means agreement between predicted probabilities and observed frequencies; the logistic bound alone does not establish it. The distinction is developed in Guo and colleagues’ study of calibration.

A reliability diagram offers a simple check. In an invented test set, collect 100 articles assigned probabilities near 0.8. Suppose reviewers mark 55 as positive. Plot their mean prediction near 0.8 horizontally and the observed fraction 0.55 vertically. A well-calibrated bin would lie near the diagonal, where those numbers agree. This example is hypothetical; Anthropology has not reported such a calibration experiment.

Repeat across probability ranges, report the number of articles in each bin and uncertainty, and inspect relevant sources and periods separately. Small bins can fluctuate substantially. Prevalence is the fraction of articles with a positive target. A model that predicts this overall fraction for every article can be calibrated yet poor at ranking individual articles. Discrimination is its ability to distinguish positive from negative cases. Calibration and discrimination therefore need separate evaluation, with articles reserved for testing rather than reused for fitting or tuning.

Checkpoint. The toy linear model predicts 1.3. Can the interface print “130% confidence”? Answer: No. It can display a labeled score. A probability needs a defined target, a suitable model, and held-out evidence that its estimates behave like probabilities.

11.4.1 Work through the logistic function and one fitting step

Let tt denote a real-valued linear score and let pp denote a modeled probability for one defined binary article label. Write exp⁡(t)=et\exp(t)=e^t for the exponential function, with base e≈2.71828e\approx2.71828; its output is positive for every finite real input. The logistic function, written σ(t)\sigma(t) in this section only, is

p=σ(t)=11+exp⁡(−t).p=\sigma(t)=\frac{1}{1+\exp(-t)}.

At zero, exp⁡(0)=1\exp(0)=1, so σ(0)=1/2\sigma(0)=1/2. At two, exp⁡(−2)≈0.135335\exp(-2)\approx0.135335, so σ(2)≈0.880797\sigma(2)\approx0.880797. At minus two, σ(−2)≈0.119203\sigma(-2)\approx0.119203. The denominator always exceeds one, placing the result strictly between zero and one for finite scores. This function’s σ\sigma is not the earlier standard deviation σj\sigma_j.

To fit such outputs, the notebook uses the first column of the four-row target table. Let θ\theta be a two-entry slope column and b0b_0 a scalar intercept, giving score ti=ziθ+b0t_i=z_i\theta+b_0 for training article index ii. Let yiy_i be its reviewed zero-or-one target. Binary log loss for one row is −yilog⁡pi−(1−yi)log⁡(1−pi)-y_i\log p_i-(1-y_i)\log(1-p_i), where log\log is the natural logarithm, the inverse of the exponential. A correct high-probability prediction gets a small loss; a confidently wrong one gets a large loss. All four losses are averaged. The demonstration adds η‖θ‖2/2\eta\lVert\theta\rVert^2/2 with penalty η=0.1\eta=0.1, keeping the intercept unpenalized.

We need three derivative rules before differentiating this model. The derivative of exp⁡(t)\exp(t) is exp⁡(t)\exp(t); the derivative of log⁡p\log p is 1/p1/p for positive pp; and the derivative of the reciprocal 1/q1/q with respect to nonzero qq is −1/q2-1/q^2. The chain rule multiplies the local rates when one function is placed inside another. We use these rules here and verify the resulting derivatives by small finite changes in the notebooks.

To differentiate the logistic function, name its denominator q(t)=1+exp⁡(−t)q(t)=1+\exp(-t). The inner negative sign has derivative -1, so q′(t)=−exp⁡(−t)q'(t)=-\exp(-t). Applying the reciprocal rule and the chain rule gives

σ′(t)=−q′(t)q(t)2=exp⁡(−t)(1+exp⁡(−t))2.\sigma'(t)=-\frac{q'(t)}{q(t)^2} =\frac{\exp(-t)}{(1+\exp(-t))^2}.

One factor is p=1/(1+exp⁡(−t))p=1/(1+\exp(-t)); the other is 1−p=exp⁡(−t)/(1+exp⁡(−t))1-p=\exp(-t)/(1+\exp(-t)). Hence the derivative is p(1−p)p(1-p). At t=0t=0, its value is (1/2)(1/2)=1/4(1/2)(1/2)=1/4.

Now let ℒ(p)\mathcal L(p) name a single row’s log loss and let yy be its fixed binary target. The notation dℒ/dpd\mathcal L/dp writes its derivative with respect to pp. Differentiating with respect to its probability gives

dℒdp=−yp+1−y1−p=−y(1−p)+(1−y)pp(1−p)=p−yp(1−p).\begin{aligned} \frac{d\mathcal L}{dp}&=-\frac{y}{p}+\frac{1-y}{1-p}\\ &=\frac{-y(1-p)+(1-y)p}{p(1-p)}\\ &=\frac{p-y}{p(1-p)}. \end{aligned}

Composing this loss with the logistic score multiplies by dp/dt=p(1−p)dp/dt=p(1-p), cancelling the denominator and leaving dℒ/dt=p−yd\mathcal L/dt=p-y. A slope changes the score at a rate equal to its input coordinate; the intercept changes it at rate one. Thus the slope gradient is the average of zi𝖳(pi−yi)z_i^{\mathsf T}(p_i-y_i) plus ηθ\eta\theta; the intercept gradient is the average of pi−yip_i-y_i.

Start all slopes and the intercept at zero. Every prediction is 0.5, so the four errors are (0.5,0.5,−0.5,−0.5)(0.5,0.5,-0.5,-0.5). Multiplying those by the first input column and averaging gives (−0.5−0.5−0.5−0.5)/4=−0.5(-0.5-0.5-0.5-0.5)/4=-0.5. The second slope gradient and intercept gradient are both zero. A gradient-descent step subtracts the gradient times a positive step size; with step size 0.2 the new slopes are (0.1,0)(0.1,0) and the intercept stays zero. Positive first-coordinate articles now receive σ(0.1)≈0.524979\sigma(0.1)\approx0.524979.

The native notebooks repeat that update 2,000 times and check a small final gradient and a separate reference slope. This verifies the training calculation for declared synthetic inputs. It supplies no real-world held-out calibration result. In the separate hypothetical reliability bin, 55/100=0.5555/100=0.55 observed frequency differs from 0.8 predicted frequency by 0.25. Constraining outputs to a probability range and testing their agreement with observations remain distinct tasks.

11.5 Why an article model cannot simply label a person

A person profile bb has been background-adjusted and normalized. It describes the direction of distinctive coverage. An article row zz has undergone neither of those person-level operations. Their equal number of coordinates does not make them the same kind of observation.

For a person score row uu, the earlier chapters reconstructed an approximate profile as b¯+uW⊤\overline b+uW^\top. Use the people basis and article predictor from the same corpus and news-coordinate version. Let spersons_{\mathrm{person}} denote the 1×m1\times m score row obtained by applying the article predictor to this reconstructed profile. Substitution gives

sperson=b¯A+β⊤+u(W⊤A). s_{\mathrm{person}}=\overline b A+\beta^\top+u(W^\top A).

The product W⊤AW^\top A has shape r×mr\times m: it describes how moving along each of the rr people directions would change those linear scores. This is a valid algebraic identity. Its interpretation as a useful prediction about a person remains an untested use outside the article model’s training distribution. Neither restoring the mean nor adding the intercept solves that problem.

Another possible quantity is the fraction of a person’s associated articles predicted to discuss a concept. That answers a coverage question, with uncertainty from both article attribution and classification. For a nonlinear predictor, predicting an average row and averaging individual predictions can differ. A future system must specify which quantity it means and evaluate it on independently reviewed person-level examples.

11.5.1 Check the algebra and then the meaning

Substitute the reconstructed row into the article rule directly: (b‾+uW𝖳)A+β𝖳(\bar b+uW^{\mathsf T})A+\beta^{\mathsf T}. Distribute multiplication over addition to obtain b‾A+uW𝖳A+β𝖳\bar bA+uW^{\mathsf T}A+\beta^{\mathsf T}, then group the last matrix product as u(W𝖳A)u(W^{\mathsf T}A). Grouping changes where we calculate intermediate tables, without changing their compatible dimensions or result. The notebook uses a separately declared three-input, two-output synthetic slope table because the earlier two-input ridge table does not fit the three-coordinate people fixture.

For a nonlinear example, take scalar article scores zero and two. Their average is one, so predicting after averaging gives σ(1)≈0.731059\sigma(1)\approx0.731059. Predicting separately and then averaging gives (σ(0)+σ(2))/2=(0.5+0.880797)/2≈0.690399(\sigma(0)+\sigma(2))/2=(0.5+0.880797)/2\approx0.690399. The difference is approximately 0.040660. Thus even the arithmetic choice of when to aggregate changes the output. A defined person-level target and separately reviewed evaluation remain necessary whichever quantity is chosen.

11.6 Aligning maps needs anchors and room to test them

Suppose two corpora contain reviewed representations of the same people. An anchor is a pair of rows known to describe the same entity in the two corpora. An orthogonal Procrustes fit tries to rotate or reflect one set of such rows toward the other. For this alignment only, let MM count paired anchors and kk count coordinates in each representation. Let XX and YY be their centered M×kM\times k tables, with matching entities in the same row; YY here is an anchor table, not the preceding concept-target matrix. Let QalignQ_{\mathrm{align}} be a k×kk\times k alignment matrix, distinct from the SVD factor QQ. The identity II now has shape k×kk\times k. The fit minimizes ‖XQalign−Y‖F2\lVert XQ_{\mathrm{align}}-Y\rVert_F^2 subject to the orthogonality constraint Qalign⊤Qalign=IQ_{\mathrm{align}}^\top Q_{\mathrm{align}}=I.

The constraint preserves distances within the chosen coordinate geometry. It cannot correct every difference in populations, source coverage, or scaling. Shared names are not enough; the anchors must be correctly identified and substantively comparable.

The published models share 60 canonical people. Sixty centered rows in 60 dimensions have rank at most 59: subtracting their mean makes the rows sum to zero, creating a dependence. Consequently, they cannot determine a full 60-dimensional alignment from independent evidence in every direction. Reserving anchors for testing leaves still fewer for fitting. A lower-dimensional or regularized approach is possible, but its assumptions and performance need evaluation. See Matching axes for the prerequisite distinction between comparing directions and assigning correspondences.

11.6.1 Fit an exact toy alignment and reserve a test row

Use four centered synthetic anchor rows (−1,−1)(-1,-1), (−1,1)(-1,1), (1,−1)(1,-1), and (1,1)(1,1) for XX. Stipulate a known map (0−110)\begin{pmatrix}0&-1\\1&0\end{pmatrix} only to construct test data YY. It maps a row (a,b)(a,b) to (b,−a)(b,-a). The first pair is therefore (−1,−1)(-1,-1) and (−1,1)(-1,1). The two tables have zero column means.

For the fit, multiply paired columns to obtain X𝖳Y=(0−440)X^{\mathsf T}Y=\begin{pmatrix}0&-4\\4&0\end{pmatrix}. Denote the SVD of this product by LΣalignT𝖳L\Sigma_{\mathrm{align}}T^{\mathsf T}, where LL and TT are two-by-two orthogonal direction matrices and Σalign\Sigma_{\mathrm{align}} contains its nonnegative singular values. The Procrustes solution is Qalign=LT𝖳Q_{\mathrm{align}}=LT^{\mathsf T}. The notebook computes it from the paired tables, then checks that it equals the stipulated map and gives zero training discrepancy.

Why this choice? Expanding the squared discrepancy yields the squared lengths of XX and YY minus twice the trace of Qalign𝖳X𝖳YQ_{\mathrm{align}}^{\mathsf T}X^{\mathsf T}Y. The first two quantities are fixed under an orthogonal map. In the singular coordinate frame, the trace is largest when each nonnegative singular value is multiplied by one rather than a smaller diagonal coefficient of an orthogonal matrix. The choice LT𝖳LT^{\mathsf T} achieves that maximum, and therefore the minimum discrepancy. This permits reflections as well as rotations, exactly as the stated constraint allows.

Reserve a fifth row (0.3,−0.7)(0.3,-0.7) from fitting. Its stipulated target is (−0.7,−0.3)(-0.7,-0.3); the estimated map reproduces that target with zero error. This noiseless test proves the toy algebra, not a real alignment’s usefulness. Actual held-out anchors could disagree because coverage, scale, or meaning differs between corpora.

The rank shortage has an equally small example. Center three two-coordinate rows: their sum becomes (0,0)(0,0), so the third row is the negative sum of the first two. There are at most two independent centered rows. With sixty centered anchors, the same dependence leaves at most fifty-nine independent rows. The notebooks also compute this bound using a synthetic centered sixty-by-sixty identity table; it is not a measurement of the real anchor matrix.

12 Following the evidence all the way back

The mathematics makes research paths visible. Following them well means keeping track of what each step establishes. This chapter uses the paper’s 2 October 2026 edition and its unchanged 1 October vector snapshot. The links are live entry points; the numbers below describe that dated audit.

12.1 A Jeff Dean journey in four steps

  1. Open Jeff Dean’s Anthropology profile. This is the identity and source-record starting point. The profile’s signal links include several news directions; no single one is “Jeff Dean’s eigenvector.”

  2. Open the explicit Dean view on news axis N24 and people pattern P2. The model selector matters: this is the Eigen Hacks hn1 people model. Its all-history Dean profile uses 253 qualified candidate articles. The displayed news coordinate is approximately +0.433 and the people coordinate +0.371. N24’s leading terms include learn, machin, model, and deep. The two scores measure different projections; they are not probabilities or two estimates of the same number.

  3. Inspect neighbors. The audit records Sanjay Ghemawat at cosine 0.7514 with 31 candidate articles, Andrew Ng at 0.7413 with 394, and Noam Shazeer at 0.7403 with 38. These cosines use the retained 24-dimensional people space. The corresponding centered 60-dimensional comparisons are 0.7082, 0.6380, and 0.6687. Compression changes the numerical similarity. Follow the Andrew Ng selection, then Ng’s canonical profile, to inspect another person’s evidence.

  4. Read an article. The Dean evidence includes a WIRED interview published on 14 December 2019. Its HN-derived record carries 15 December. Keep those dates distinct: the publisher and the archive record describe different observations. Then return to the map and ask which details in the source support the numerical association.

The related Eigen Hacks people view and Eigen Times people view extend the exploration into their separate corpora. Their axis numbers are local addresses, not universal concepts.

The audited current-news overlay has zero newly qualified article-person joins. It can display recent articles in the fixed news basis, but proximity to Dean’s historical profile does not establish that a recent article mentions him. This distinction is precisely where a promising research lead still needs evidence.

12.2 “Today’s database news against today’s database CEOs”

This query combines three tables: articles relevant to databases on a chosen date, sourced executive roles valid on that date, and qualified links between articles and people. The word “today” must resolve to an explicit date in all three.

Let tt be the queried date, tstartt_{\mathrm{start}} a role’s known start date, and tendt_{\mathrm{end}} its known exclusive end date. Exclusive means the role is no longer valid on that endpoint. The symbol ≤\leq includes equality, whereas << does not. Validity at the queried date can then be written tstart≤t<tendt_{\mathrm{start}}\leq t<t_{\mathrm{end}}. This is a useful interval convention, not permission to invent an end date. An unknown end does not by itself prove that a role continues today; retain the source’s as-of date and uncertainty.

The paper’s dated role query reaches executive records that have no supported profile in either published people model. The correct result is a visible gap, not a fabricated point on the map. Larry Ellison’s historical database coverage also cannot be substituted for a claim that he is the current CEO of a database company. A founder, a chair, and a current chief executive are distinct roles.

The research query can therefore show topic-ranked articles beside sourced roles while leaving missing joins visible. A completed current-CEO comparison would need additional profile coverage and article attribution. A larger identity catalog alone supplies neither.

12.3 One name, several kinds of evidence

Canonical identity answers which entity a record refers to. A Wikipedia/Wikidata correspondence can help confirm the node; it does not validate every article placed in a same-name bucket. Reviewed Devreal mappings reuse the existing canonical node. Name similarity alone never settles that correspondence.

A typed graph assertion answers another question: what relationship does a particular source state? “Employed by,” “reports to,” “coauthored,” and “competes with” have different meanings and sometimes different directions. Two people working for the same organization need not have overlapped or known each other. A co-mention records shared appearance in an article, leaving the relationship unspecified.

Dates and review states travel with the assertion. A source assertion, a machine classification, and a curator’s judgment are different kinds of support. A citation provides a route to inspection; it is not a measured probability of truth. Unknowns remain part of the record.

12.3.1 Follow a claim through a typed chain

Suppose a synthetic article’s record says it names canonical person A. That establishes the recorded article-person association under its review status. A topic model gives the article a coordinate on a database-related axis. That establishes a measurement in a particular news basis. A cosine places A near B. That establishes similar coverage directions in a particular representation. None of these records states that A employs B.

To add an employment assertion, inspect a source that actually states that relationship, identify the subject and organization, and retain its dates and review status. To add a colleague assertion, evidence must support the relevant overlap or relationship rather than merely a shared organization name. One can preserve a candidate lead without promoting it to an asserted fact. This chain keeps identity, attribution, numerical similarity, source assertion, and curator judgment separate enough to correct independently.

12.4 A financing is an event, not a complete investment graph

Use a fictional announcement: “FinchDB raised $4 million; Harbor Ventures led and Cedar Capital participated.” There is one financing event, one recipient, and two explicitly named investor roles. The statement does not tell us how much either investor contributed. A partner quoted about the company is not automatically a personal investor or the deal’s lead partner.

One mathematical picture is an incidence table: rows are entities and columns are events. An entry records participation, with a separate role label such as recipient, lead, or participant. This explanatory picture lets several entities meet at one event without inventing every possible pairwise relationship. The deployed catalog stores event records and optional investor relationships; it need not materialize this matrix.

Now imagine a separate fictional regulatory notice reports $1 million sold, later amended to $3 million. These are successive cumulative observations of an offering, so adding them to claim $4 million would double-count. An offering target, proceeds sold to date, and an announced round total answer different questions. First-sale, filing, and announcement dates also remain distinct until evidence reconciles the events.

The audited catalog contains 2,893 financing events and 81 investor/partner relationship records. The totals need not agree: an event can have several known investors or none. Eight historical investor associations are excluded from the financing total. Of the financings, 2,875 come from selected SEC issuer-reported notices; publication is not SEC verification, and the selected notices are not a census of venture rounds.

Checkpoint. A financing has no named investor. Is its event record necessarily incomplete collection? Answer: No. The source may establish the financing while withholding investor identities. Preserve the event and the unknown participants.

Checkpoint. Two profiles have cosine 0.95 and share an employer. May we add “colleagues”? Answer: Those facts motivate a search. They do not establish overlapping employment or a sourced colleague relationship.

12.5 Reading onward

Use The Dual Geometry of People and News for the compact derivation, implementation audit, limitations, and source references behind this companion. Return to The Mathematics of Eigen Times for the foundations, especially text representation, covariance, SVD, named axes, and projection. Then explore Anthropology with three questions in view: which coordinates are being compared, what evidence supports the link, and which date does the claim describe?

Notation and shapes

All observation vectors in this companion are rows. A matrix direction, such as a column of WW, is a column. The shapes below summarize the definitions introduced in the chapters. The letter nn counts fitted people; NN counts reviewed training articles for the proposed calibration fit.

Symbol Shape Meaning
xa,μx_a,\mu 1×d1\times d Article embedding and its fitted reference mean
VV d×Kd\times K Retained raw news directions
R,DR,D K×KK\times K Naming rotation and named-axis scale matrix
ca,ya,zac_a,y_a,z_a 1×K1\times K Raw, named, and standardized article coordinates
qg,mpgq_g,m_{pg} 1×K1\times K Cell background and person-cell mean
rp,bpr_p,b_p 1×K1\times K Balanced contrast and its unit direction
B,HB,H n×Kn\times K Unit person rows and their centered rows
b‾\bar b 1×K1\times K Mean unit person profile in the fitted population
CPC_P K×KK\times K Equal-person covariance
WW K×rK\times r Retained people-pattern loadings
upu_p 1×r1\times r A person’s retained people coordinates
WW𝖳WW^{\mathsf T} K×KK\times K Rank-rr projector in news-coordinate geometry
Z,YZ,Y N×KN\times K, N×mN\times m Article inputs and reviewed concept targets
A,βA,\beta K×mK\times m, m×1m\times1 Proposed concept slopes and intercepts
W𝖳AW^{\mathsf T}A r×mr\times m Algebraic slopes from people directions to concept scores

Here d=384d=384, K=60K=60, and r=24r=24 in the published models described by the paper; the main fictional example uses K=3K=3 and r=2r=2. The bare rr is a retained dimension count; the indexed rpr_p is a person’s contrast row. The hypothetical number of ontology targets, mm, is a modeling choice rather than a statement that every ontology concept has a trained classifier.

In the thin SVD H=UΣQ𝖳H=U\Sigma Q^{\mathsf T}, k*=min⁡(n,K)k_*=\min(n,K). The matrix UU is n×k*n\times k_*, Σ\Sigma is k*×k*k_*\times k_*, and QQ is K×k*K\times k_*. Selecting the first rr columns of QQ with the fitted orientation produces WW; the corresponding columns of UU and diagonal block of Σ\Sigma produce the score matrix UrΣrU_r\Sigma_r. Coordinate sums SpgwS_{pgw} and concept-score rows ss are different objects.

Time-window subscripts extend the same shapes: bpw,qgw,mpgwb_{pw},q_{gw},m_{pgw} have KK coordinates, while upwu_{pw} has rr. The displacement Δp\Delta_p, discarded squared length QpQ_p, attention share αa\alpha_a, and ranking score are scalars. In the alignment example, X,YX,Y are M×kM\times k anchor tables and QalignQ_{\mathrm{align}} is the k×kk\times k orthogonal map; that use of YY is local to alignment.

Answers to the checks

Check 1. One unique background article and up to two person-article associations. The background describes a set of articles; the person profiles describe associations with each individual.

Check 2. The forward result is 1×241\times24, the transpose is 24×6024\times60, and WW𝖳WW^{\mathsf T} is 60×6060\times60. The returned row hWW𝖳hWW^{\mathsf T} is 1×601\times60. The combined matrix has rank 24, so it cannot be the 60-dimensional identity.

Check 3. No. Diagonal entries describe individual variances. In the example the off-diagonal entries remain -0.6 after marginal standardization, so the two coordinates remain correlated.

Check 4. Repeating E1 inside the background would weight that reporting more heavily merely because it has two associated people. E1 should nevertheless enter both personal aggregations because each asks about the articles associated with a different person. Deduplicate within a set, not across two different research questions.

Check 5. P1 ranks N3 second: its loading 0.5949 is smaller in magnitude than N1’s 0.7964. N3’s own row ranks P1 ahead of P2 because 0.5949 exceeds 0.3087. The shared entry agrees exactly; the competing entries differ.

Check 6. (0,1,0)(0,1,0) also projects to (1/2,0)(1/\sqrt2,0) and returns as (1/2,1/2,0)(1/2,1/2,0). The two input rows differ only along a discarded direction, so the retained map cannot distinguish them.

Check 7. No. A and B move from cosine 0.0392 to 0.8211, a change of about 0.782. Total variance is an aggregate reconstruction quantity, not a bound on every pairwise angle.

Check 8. Yes. Its comparison background can change because other associated articles enter a cell, or older articles can leave a rolling window. A support threshold can also change which cells remain in the average. These mechanisms do not necessarily describe a change in the person’s activities.

Check 9. No. A smaller topic coordinate can be outweighed by a larger attention share. Calling an article news about a person additionally needs a qualified article-person association and the underlying identity evidence. A ranking product cannot supply that link.

The three later checkpoints distinguish a linear score from a probability, a sourced financing event from its possibly unknown investors, and a coverage neighbor from a documented colleague. Their answers appear where each question is introduced so the required evidence remains next to the proposed inference.

Reproducing the examples

The companion’s public source bundle includes the Markdown manuscript, scripts/examples.py, examples.json, the figure sources and rendered images, Mermaid diagrams, packaging configuration, and a compact paper-evidence.json snapshot of reported real-data measurements. It also includes the supplied shelf artwork, which is kept separate from the reading editions. No restricted news archive or full catalog export is required to reproduce the fictional calculations.

From the Anthropology repository root, regenerate and check the numerical examples with:

cd papers/mathematics-of-anthropology
python3 scripts/examples.py
python3 scripts/examples.py --check

NumPy supplies the numerical linear algebra. Figure generation retains editable SVGs and uses rsvg-convert for PNGs; Mermaid source files are included with their rendered diagrams. The README records the rendering and four-format build commands. The script verifies unit norms, covariance and eigenvector identities, the forward-return projection, residual orthogonality, cosine comparisons, and the ridge solution. These are arithmetic and implementation checks, not scientific validation of the real people model.

The reported Anthropology measurements retain their source dates and model identifiers. The Hacker News people model is eigenhacks:eigen:hn1:people:v1:bcaccd95b794; the general-news model is eigentimes:eigen:v2:people:v1:47debba3db48. The evidence snapshot records the source paper’s checksum and audited implementation revisions. The public source inventory records the companion inputs and build revision. A refreshed news overlay must not silently replace the dated evidence used in this edition.

The mathematical pipeline can be checked from synthetic inputs, while the empirical claims require their stated source artifacts and attribution rules. Keeping those two kinds of reproducibility distinct is part of making the model useful for research.

Glossary

Term Meaning in this companion
Adjoint Map transferring a linear operation across a dot product; the transpose for the real Euclidean coordinates here.
Anchor Reviewed pair representing the same entity in two maps for an alignment.
Article association Recorded candidate link between an article and a canonical person, with attribution evidence and review state.
Attention share Allocated scalar story attention divided by a stated source-period total.
Basis Independent directions used to express a space or subspace.
Calibration Agreement between modeled probabilities and observed label frequencies for a defined task.
Cell A specified source-and-decade grouping used for comparison and support.
Centering Subtracting the selected reference mean from observations.
Contrast A difference from a declared reference; a person contrast averages supported person-cell minus background rows.
Coordinate Amount along a named direction in a specified coordinate system.
Correlation Covariance divided by both coordinates’ positive standard deviations.
Covariance Average product of paired centered coordinates, with an explicit denominator convention.
Crosswalk Reviewed table connecting distinct vocabularies or representations while retaining their provenance.
Derivative Limiting rate of scalar output change per input change; a partial derivative varies one input while holding others fixed.
Discrimination Ability to distinguish positive and negative labeled cases, separate from calibration.
Eigenvalue and eigenvector A stretch factor and direction satisfying covariance times direction equals factor times direction.
Embedding Numerical representation of text produced by a specified model.
Frobenius norm Square root of the sum of squared matrix entries.
Gradient Collection of partial derivatives giving local changes of an objective with its parameters.
Held-out data Observations excluded from a stated fitting or tuning stage and reserved to evaluate it.
Incidence table Table marking which entities participate in which articles or events.
Intercept Baseline linear score when all input coordinates are zero.
Logistic function Map from a real score to a value between zero and one using the exponential function.
Mean Coordinate sum divided by its observation count, or a stated weighted counterpart.
Norm A length; the Euclidean norm here squares coordinates, adds them, and takes the square root.
Normalization Dividing a nonzero row by its own length to retain direction.
Ontology Reusable vocabulary of concepts and relationships among their meanings.
Orthogonal and orthonormal Perpendicular; perpendicular with unit length. An orthogonal square matrix changes orientation while preserving lengths.
PCA Principal component analysis: finding perpendicular directions of greatest variation in centered observations.
People pattern Learned direction of variation across person profiles.
Person profile A normalized, background-adjusted direction of one person’s associated news coverage.
Projection Measuring and retaining components along chosen directions.
Procrustes alignment Orthogonal fit aligning paired anchor rows by minimizing squared discrepancies.
Pseudoinverse Generalized inverse giving a least-squares solution, selecting the shortest when multiple solutions exist.
Rank Number of independent directions represented in a matrix.
Reconstruction Approximation rebuilt from retained coordinates.
Residual Original row minus reconstructed row, before any summary of its size.
Ridge penalty Added squared-slope cost that trades prediction fit against coefficient magnitude.
Singular value Nonnegative scale connecting left and right directions in an SVD.
Slope Coefficient multiplying one input coordinate in a linear prediction.
Standard deviation Nonnegative square root of variance.
Standardization Dividing coordinates by their reference standard deviations; mean subtraction must also be specified.
Subspace All weighted combinations of selected directions.
Support Eligible observed articles satisfying the stated profile gates; not a guarantee of statistical reliability.
SVD Singular value decomposition, factoring a table into left directions, nonnegative scales, and transposed right directions.
Trace Sum of a square matrix’s diagonal entries.
Variance Average squared centered coordinate value under a declared denominator convention.
Whitening Transformation giving unit variances and zero cross-coordinate covariance in the chosen retained space.

Subject index

This index points to explanations and worked calculations rather than to isolated mentions. Notebook links use the same local anchors as the book.

Aggregation and counting: Unique backgrounds; equal supported cells; incremental counts and sums.

Alignment: Anchor fit and held-out row; constraints and rank.

Attention: Allocation; source denominators; topic ranking.

Calculus and fitting: Difference quotient, derivative, and ridge minimum; logistic gradient step.

Covariance and variance: Paired-deviation table; people covariance entry; retained variance.

Eigenvectors: Two-coordinate multiplication and maximum variance; people basis.

Evidence: Identity, association, and claim chain; dated Dean journey and event examples.

Geometry: Coordinate versus directional neighbors; compression effects.

Matrices: Product and transpose by hand; notation and shapes.

Normalization and standardization: Length and unit rows; A’s exact normalization; different denominators.

Ontology and probability: Tags, crosswalks, and proposed models; logistic and calibration; person-level substitution.

Projection and residual: First residual; exact round trip; SVD reconstruction.

Rotation and whitening: Rotated covariance arithmetic; article coordinates.

Scores and loadings: Both person dot products; one loading table, two entrances.

Time and version: Fixed ruler and separate clocks; background-only updates.