Eigen Times and Eigen Hacks turn an article’s text into an embedding, a list of numbers produced by a text model. Anthropology does not learn a new text embedding for every person. It inherits a versioned news coordinate system and measures each associated article in that system.
Let index an article and count embedding coordinates. Let be article ’s embedding row, and the mean embedding used when the news basis was fitted. Let count retained news directions and let contain them as orthonormal columns, so has shape . Denote the resulting raw news-coordinate row by :
Read this in two steps: subtract the reference mean, then take a dot product with each retained direction. In the audited implementation, and . This projection has already discarded information before people enter the calculation.
Why these particular directions? They are the directions of greatest fitted variation, found from a covariance matrix. The Eigen Times covariance chapter develops that argument. For the moment, regard as a fixed measuring instrument, accompanied by its mean and version.
A synthetic three-coordinate article can make every operation visible. Set its embedding row to and its fitted reference mean to . Subtraction produces . Choose all three coordinate directions as and introduce the square naming rotation for this small illustration. The raw and named rows are then both . This identity choice isolates the scaling arithmetic in the next section; it is not a claim that the production embedding or rotation is an identity.
The notebook’s batch fixture uses the same synthetic reference mean and scales for all 45 articles. It constructs the embedding rows from the declared standardized rows and verifies that the forward operations recover those rows. The fitted production mean, directions, and scales instead come from their stated news model; fitting them anew to one person’s articles would change the measuring instrument.
A statistically useful direction need not have a readable name. Eigen applies a stored orthogonal rotation . Denote the resulting named news-coordinate row by :
Lexical loadings measure associations between words and numerical directions. Varimax is an orientation rule that makes these loadings more concentrated on individual axes. Eigen uses it to choose the rotation, helping some axes acquire recognizable word lists. The Eigen Times varimax explanation shows why this can improve interpretability.
An ontology is a reusable vocabulary of concepts and their
relationships. The word lists are labels to investigate, not ontology
definitions. Stemming shortens words by removing endings, so a label
such as databas represents a processed word form. A list
containing databas, ceo, and
interview can describe mixed coverage. An axis has two
poles: its positive and negative directions. Positive-loading words may
describe one pole better than the other. In the implementation, the
lexical naming fit uses raw scores divided by their standard deviations,
called standardized raw scores, while the stored rotation acts on
unstandardized article scores. Operations commute when changing their
order leaves the result unchanged. Scaling and rotation need not
commute, so the lexical lists are a naming heuristic rather than exact
term contributions to every final coordinate.
Some named directions vary more than others. A coordinate of 2 on one direction may be ordinary while 2 on another is unusual. Anthropology divides each named coordinate by its fitted standard deviation.
Let index named news axes, from 1 to , and let be the positive standard deviation of axis . Let be the diagonal matrix containing those scales: its row-, column- entry is , and all off-diagonal entries are zero. The superscript denotes a matrix inverse, which undoes the corresponding multiplication. Here has reciprocal scales on its diagonal, so multiplying by it divides coordinate by . Denote the standardized article row by :
For example, with standard deviations , named coordinates become standardized coordinates . When discussing a generic article we omit its subscript and write for such a standardized row; is its coordinate . We will aggregate these rows into people profiles.
An eigenvector of a covariance matrix is a direction that multiplication by that matrix stretches without turning; its eigenvalue is the stretch factor. Let index the retained raw axes, from 1 to , and let be their positive fitted variances, which are eigenvalues of the raw covariance. The coefficient is the entry in row , column of the rotation. Summing over all raw axes gives each named-axis variance:
Each rotated variance is a weighted combination of raw variances. The squared rotation coefficients supply the weights. The mean, directions, rotation, variances, and corpus identity must travel together. Axis N24 in one model is not automatically N24 in another.
With named row and fitted standard deviations , divide coordinate by coordinate: , , and . The standardized row has length , so standardization has not made it a unit row. Normalizing it would be a further operation, producing and losing its overall magnitude.
Each standard deviation belongs to a column of the fitted reference population. Each row length belongs to an individual observation. The first operation changes units; the second discards size. A zero standard deviation cannot be used as a divisor and requires an explicit modeling policy, not silent division. The formulas in this chapter assume positive fitted scales.
Marginal standardization gives each coordinate unit variance separately. Whitening would also make the covariance between different coordinates zero; we will see that these are different operations. In this illustration only, use two raw directions with variances 4 and 1. Rotate them by 45 degrees using the following matrix:
Write for the feature covariance matrix of the standardized article rows in this illustration. Both rotated coordinates have variance 2.5, but they also have covariance -1.5. Divide both by and their variances become 1 while their covariance becomes -0.6:
They now have the same marginal scale. They still move together. Correlation is covariance divided by the two standard deviations; when both variances are one, the correlation equals the covariance. Full whitening would remove these off-diagonal correlations. This pipeline performs marginal standardization, not full whitening.
For the raw row , rotating and then marginally scaling gives approximately . Scaling the raw coordinates first by and then rotating gives . The two operations answer different questions. Their order must be part of the model definition.
This also changes the geometry in which people are compared. Equal lengths in standardized news units need not be equal lengths in the original embedding. A Mahalanobis statistic measures squared distance from a reference mean while accounting for covariance, not just each coordinate’s separate scale. The sum is not generally that original statistic; see the Eigen Times section on residuals and T squared. Outside this two-dimensional illustration, the symbols retain their fitted-model meanings.
Check 3. Does a covariance matrix with ones on its diagonal necessarily describe uncorrelated coordinates? Use the two-dimensional example to answer.
Write the raw coordinate covariance as , where constructs a diagonal matrix from its listed entries. In the displayed rotation, the first named coordinate is the sum of the two raw coordinates divided by ; the second is their difference, with the first raw coordinate negative, divided by .
Because the raw coordinates have zero covariance, the first named variance is . The second uses the squared negative coefficient and has the same variance. Their covariance uses paired, unsquared coefficients: . Dividing the two named coordinates by divides this covariance by , leaving .
The row first rotates to and then scales to . Scaling first instead gives and rotating gives . These routes standardize relative to different coordinate systems, so they use different fitted scale matrices and define different maps. The named-coordinate scale here is a scalar multiple of the identity and actually commutes with the rotation. This example therefore does not test noncommutation of one fixed pair of matrices; each standardizer belongs to the coordinates whose variances it measures.