A Step-by-Step Companion to The Dual Geometry of People and News
4 October 2026 | Edition 1.1.0-9ce09c2a
Anthropology lets a reader start with a person, find the kinds of news associated with that person, follow those news patterns to other people, and return to the articles. Its research paper, The Dual Geometry of People and News, describes the system and its evidence. This companion explains the mathematics one operation at a time.
The starting point is The Mathematics of Eigen Times. That book explains how articles become vectors, how covariance reveals recurring news directions, and how projections describe a new story. Here we take the next step. A collection of articles associated with a person becomes a profile in news space. A population of profiles produces another set of directions: patterns of how people are covered. One loading matrix connects the two spaces.
You need arithmetic and a willingness to read a small table. The next two chapters introduce the vector and matrix ideas needed to continue. When a prerequisite deserves a longer explanation, a link takes you to the relevant Eigen Times section. We do not repeat its derivations of text weighting, efficient matrix decomposition, or clustering thresholds. We spend that space on the new questions: how to combine a person’s articles, what to compare that coverage with, what a people pattern means, why navigation works in both directions, and how shared ontology concepts could connect different news corpora.
One fictional example runs through the book. It has four people, three news coordinates, two source-and-decade cells, and 45 distinct articles. The running example and plotted numerical matrices are generated by the accompanying script. Short additional arithmetic illustrations are written out in the text. Figures use blue for positive values and rust for negative values; color describes a sign, never a judgment about a person. Displayed values are rounded, while computations use their full precision.
The empirical examples come from the Anthropology paper’s 2 October 2026 edition. Its people models retain the 1 October snapshot. The catalog has subsequently grown to 59,399 people and organizations, but the fitted models still contain 217 Hacker News profiles and 61 general-news profiles. A catalog identity and a supported mathematical profile are different objects. None of the fictional numbers below is a measurement of a real person.
Version 1.1.0. This teaching expansion works through lengths, residuals, covariance, eigenvectors, SVD, profile aggregation, neighbor geometry, attention, and proposed ontology models in smaller steps. The complete Python and native OCaml notebooks compute the new synthetic examples as well as the original fixtures. Original numerical evidence, figure inputs, and the dated Jeff Dean journey are unchanged.
We will also keep implemented operations separate from research proposals. The published system computes the profiles, people basis, reciprocal indexes, and dated news overlay described here. Incremental profile refresh, learned ontology calibration, and a validated alignment between the two corpora are extensions to evaluate. Explaining their mathematics does not make them deployed features.
| If you want more on | Read in the Eigen Times companion |
|---|---|
| Turning words into numbers | Text as vectors |
| Dot products and angles | Cosine similarity |
| Rows, columns, multiplication, transpose | Matrices pictured |
| Means, covariance, and eigenvectors | The axes of a point cloud |
| Singular value decomposition (SVD) and low-rank approximation | The singular value decomposition |
| Naming axes by rotation | Varimax |
| Projection and missing information | Projection, reconstruction, residual |
| Updating sums without rereading everything | Sums over blocks |
A formula is a compact record of operations. First identify the inputs, then the operation, then what its output measures. Each chapter follows that order. The original four-person example remains one connected calculation; smaller examples are explicitly synthetic and set aside that fitted model temporarily. A successful arithmetic check does not validate the archive’s identities, attribution, or social claims.
In either notebook, run the setup and then the cells in reading order. A kernel is the running language process. Restarting it and running every cell checks that no hidden earlier calculation is needed. An assertion is a check that stops execution when a stated identity or reference answer fails. Edit the declared inputs to experiment; when an input changes, a fixed reference answer can fail deliberately while an identity should still hold. The Python and OCaml versions compute independently and need no private archive.
A scalar is one number, a vector is an ordered list, and a matrix is a rectangular table. Here an observation vector is a row. A direction used for projection is a column. A shape such as means four rows and three columns. The real numbers, written , include negative values, zero, and fractions; means a table of that shape with real entries.
An index selects an entry: is coordinate of row , and is row , column of matrix . The index is not a multiplier. The notation means . Adjacent scalar symbols mean multiplication; adjacent matrices mean the row-by-column product taught below. Parentheses group operations, so work inside them first. The symbol means membership, and counts the members of a set .
A square means multiplied by itself. For nonnegative , is the nonnegative number whose square is . Single bars give a scalar’s absolute value, removing its sign; double bars give a vector’s length. A bar above a symbol denotes an average, and a hat denotes a reconstruction or estimate. The superscripts and both transpose a matrix, exchanging rows and columns. The symbol denotes a square identity matrix, with diagonal ones and other entries zero. A superscript on a square matrix means an inverse when it exists. The symbol warns of rounded or approximate equality.
This table is a reading map. The chapters introduce each object again with its dimensions before using it in a calculation.
| Object | Symbol and meaning |
|---|---|
| Article row | : article measured in standardized news coordinates. |
| Comparison within a cell | : unique-article background; : person ’s mean in cell . |
| Person before and after normalization | : balanced contrast; : its unit direction. |
| Population reference | : average fitted person profile; : centered row. |
| Learned people directions | : loadings mapping news coordinates to retained people coordinates . |
| Dimension counts | : embedding width; : news width; : retained people width. The scalar differs from row . |
| Sample counts | : fitted people; : reviewed training articles in the proposed concept model. |
| Concept model | : slopes; : intercepts; : predicted score row. |
| Local reuse of a letter | is an SVD direction matrix; is a scalar residual statistic; is an alignment matrix. |
Normalization makes one row have length one. Standardization divides each coordinate by its reference standard deviation. Centering subtracts a reference mean. These answer different questions and cannot be substituted for each other. A positive coordinate names one pole of an axis; its sign does not express approval, goodness, truth, or confidence.
Across volumes, equivalent roles sometimes use different symbols. The History Math companion’s unified notation uses for the retained people dimension count, for the direction matrix, for the score table, for concept coefficients, and for the concept count. This established Anthropology edition writes those roles as , the same , rows stacked together, , and , respectively. History counts people with and writes their covariance as ; this book uses and for those roles. Do not substitute letters without also matching their shapes and centering conventions. Here remains the people covariance, not the concept count. Here names the SVD’s right-direction matrix, while is a person’s scalar discarded squared length. These distinct objects never share a formula merely because their letters resemble one another.
The final notation table collects shapes. The glossary defines terms in words, and the subject index links to worked explanations. The Eigen Times companion sometimes writes observations as columns; transpose its complete formula when translating to this book’s row convention.
Imagine collecting articles that mention a database founder. Some concern database systems, some concern enterprise sales, and some concern unrelated activities. Each article can be measured on the same fixed news directions. The collection then has a mathematical shape: it spends more of its coverage in some directions than in others.
That is the sense in which a person becomes a source of news. It is a distribution of associated reporting. It does not mean the person wrote, approved, or caused the articles. Nor is the resulting vector a psychological portrait. Change the corpus (the article collection), the dates, or the attribution rules and the profile can change.
Three objects must remain separate throughout the calculation. An identity says which person a record denotes. An article association says that a particular article is a candidate match to that person. A relationship assertion says something specific, such as that one person reported to another during a stated period, with a source for the claim. A reliable identity does not automatically make every name match reliable. Two people appearing in an article does not establish a relationship between them.
The vector model starts from qualified candidate article associations. Anthropology’s graph stores separately sourced assertions and explicitly marked co-mentions. The same interface can display both, but one is not evidence for the other. A short distance in the vector view means similar coverage under that model. A line labeled employment in the graph needs employment evidence.
There will also be two meanings of people vector. A person profile describes one individual. A people pattern is a direction learned from the variation among many profiles. Think of the difference between a city’s weather measurements and a recurring weather pattern. A city has coordinates on the pattern; it is not itself the pattern.
The diagram is a roadmap of operations explained below. Project means measure along selected directions; rotate means change their orientation; standardize means divide by reference standard deviations. Center means subtract a comparison mean, and normalize means divide a nonzero row by its length. Principal component analysis (PCA) finds directions of greatest variation among centered observations. These names label the stages, not additional assumptions about a person’s identity.
The book follows this path in the order shown before explaining how the middle can be explored in both directions. At each step we will ask what the numbers mean and what information the operation discards.
Check 1. An article names two people. How many unique articles should it contribute to a background average, and how many person-article associations might it contribute? Keep your answer for the solutions at the end.
A scalar is a single number. A vector is an ordered list of numbers called coordinates. Our observation vectors are rows; when a matrix direction or another object is a column, we will say so. Let denote this example row:
It has three coordinates. The order matters: its first coordinate is 2, its second is 1, and its third is -1. In the running example the coordinates belong to three fixed news axes, called N1, N2, and N3. These names identify toy directions, not actual Anthropology topics.
Let be a second row of the same length. The index selects a coordinate: and are the numbers in position of the two rows. The summation sign means to add over all those positions, from 1 to 3 here. The dot product, written with a centered dot, multiplies corresponding coordinates and adds:
For these rows, the result is . An omitted multiplication sign between scalar entries also means multiply.
The Euclidean length, or norm, of is written . It is the square root of the sum of squared coordinates:
Dividing a nonzero vector by its length gives a unit vector. It keeps the direction while removing the size. The zero vector has no direction and cannot be normalized this way.
Single bars around a number mean its absolute value, or size without its sign. Thus . Single bars around a set will instead count its members; we will introduce those sets before using them. The symbol means approximately equal, as when values have been rounded.
Why square before adding? Simply adding the entries of gives zero even though the row is not zero. Squaring counts both directions as positive contributions. For a separate synthetic row , the squares are 9 and 16, their sum is 25, and the length is . This extends the right-triangle rule: perpendicular coordinate contributions add as squared lengths.
For the earlier row , the squared contributions are 4, 1, and 1. Its squared length is 6 and its length is approximately 2.44949. Dividing each coordinate by that same number gives approximately . Square these normalized coordinates and add: . That is the unit-length check; it does not require every individual coordinate to equal one.
Multiplying by 2 gives and doubles its length from 5 to 10. Normalizing either gives . Multiplication by -2 instead reverses the direction, giving unit row . Positive rescaling preserves direction; a negative rescaling reverses it. Neither normalization establishes that one observation has more supporting evidence than another.
Try it. Change the row to , then to . The first has unit direction . The second has length zero, and division by its length is undefined. A missing or directionless profile must not silently become an average profile.
A matrix is a rectangular table. Let count its rows and its columns; its shape is written . Put one three-coordinate person profile in each of four rows and the result is . A matrix entry uses its row index first and column index second. Diagonal entries have equal row and column indices; they run from the top left toward the bottom right.
Multiplication describes a weighted combination. Its numerical weights are coefficients; coefficients expressing directions in input coordinates are also called loadings. If is a row and is a loading matrix, then is a row. Its first number is the dot product of with the first column of . Its second number uses the second column. The adjacent dimensions must agree:
The transpose, written , exchanges rows and columns. It therefore has shape . Remembering shapes will prevent more mistakes than memorizing a complicated formula.
Use a separate synthetic transformation , with three input coordinates and two output coordinates:
For , the first output is . The second is . Thus . We multiplied a row by a table, producing a row. Each output column says how much of each input enters that output.
Transposing gives . Multiplying the returned row gives , which differs from . This deliberately arbitrary table does not have perpendicular unit columns, and its transpose is not an inverse. Shapes allow a multiplication; they do not promise what that multiplication will recover.
Suppose three measurements are 2, 3, and 7. Their mean is 4. Centering subtracts that mean, giving -2, -1, and 3. A negative centered value means below the chosen mean; it does not mean a negative amount of the original quantity.
For vector rows we average each column separately and subtract the resulting mean row. We will center twice for different purposes: first against a news background within a source-and-decade cell, then against the mean profile of the fitted population. These are two different comparisons, not a redundant repetition.
Variance measures spread by averaging squared deviations from a mean. A standard deviation is the nonnegative square root of a variance; it is zero when all values agree. Covariance averages the product of the centered values of two coordinates: a positive result means they tend to move together, a negative result means they tend to move oppositely. A covariance matrix collects those comparisons for every pair of coordinates. Its diagonal contains their individual variances. The fitted news model supplies its covariance convention; for the people calculation below we explicitly divide by the number of people.
The mean of 2, 3, and 7 is . The centered numbers are -2, -1, and 3; their sum is zero. Their squared deviations are 4, 1, and 9. For this descriptive example we divide by the observation count, three: variance is , approximately 4.6667. The standard deviation is , approximately 2.1602. Variance has squared coordinate units; standard deviation returns to the original units.
Pair these measurements with a second coordinate, 6, 4, and 2, in matching row order. The second mean is 4. Its deviations are 2, 0, and -2. The table keeps the pairings visible:
| Row | First deviation | Second deviation | First squared | Second squared | Product |
|---|---|---|---|---|---|
| 1 | -2 | 2 | 4 | 4 | -4 |
| 2 | -1 | 0 | 1 | 0 | 0 |
| 3 | 3 | -2 | 9 | 4 | -6 |
| Sum | 0 | 0 | 14 | 8 | -10 |
Divide the last three column totals by three. The resulting covariance matrix is
The repeated off-diagonal entry is the average paired product. Reversing the order of the two factors leaves each product unchanged, so the covariance matrix is symmetric. The negative sign records that a positive deviation in one coordinate tends to accompany a negative deviation in the other. It is not a causal claim.
Dividing centered columns by their standard deviations yields unit variances. Their covariance becomes the correlation , approximately -0.9449; standardization has not removed the association. Some estimators divide by one less than the count to estimate a population variance from a sample. Here the count denominator defines the spread of these declared rows; the people model below uses that same convention. Changing a denominator must be an explicit modeling choice.
Check the mechanism. If the second coordinate were constant, every second deviation would be zero. Its variance and every covariance involving it would be zero, and division by its standard deviation would be unavailable.
Perpendicular directions have dot product zero. Perpendicular unit directions are orthonormal. Let denote an identity matrix: it has ones on its diagonal and zeros elsewhere, so multiplying by it changes nothing. Its size matches the matrix product beside it. Let be a square matrix whose columns form a complete orthonormal system. Such a matrix is called orthogonal and satisfies . Multiplying by rotates or reflects coordinates and preserves all lengths and angles.
A tall matrix can also have orthonormal columns: . But when it has fewer columns than rows, multiplying by keeps only some directions. The length can decrease. The statement that an orthogonal change of coordinates preserves lengths applies to a complete rotation; a truncated projection preserves only the retained component. This distinction will explain the navigation duality.
Directions are linearly independent when none is a weighted combination of the others. A basis is a set of independent directions used to express coordinates. The subspace spanned by a collection of directions contains all their weighted combinations. The rank of a matrix counts its independent directions; keeping fewer directions is a low-rank representation.
The Eigen Times matrix pictures provide a longer visual introduction. When comparing a formula written with column vectors to this book, transpose it consistently rather than changing only one factor.
Check 2. A row has 60 news coordinates and its loading matrix has shape . What are the shapes of the forward result, the transpose, the combined matrix , and the returned row?
Take the synthetic row and retain only the horizontal unit direction . The dot product is . Rebuilding three units of that direction gives . The residual is what remains when we subtract the reconstruction from the original: .
The residual is a row, not initially a scalar error score. It tells us where the missing part points. Its squared length is . The reconstruction’s squared length is 9; the original’s is 25. Thus . Their dot product is zero, so the retained and residual parts are perpendicular. A later chapter applies this same calculation to the fitted people plane.
For comparison, multiplying by the complete orthogonal matrix gives . Its squared length is still 25. Multiplication by the transpose recovers exactly. A complete change of orientation has no discarded part; keeping only the horizontal coordinate did.
Eigen Times and Eigen Hacks turn an article’s text into an embedding, a list of numbers produced by a text model. Anthropology does not learn a new text embedding for every person. It inherits a versioned news coordinate system and measures each associated article in that system.
Let index an article and count embedding coordinates. Let be article ’s embedding row, and the mean embedding used when the news basis was fitted. Let count retained news directions and let contain them as orthonormal columns, so has shape . Denote the resulting raw news-coordinate row by :
Read this in two steps: subtract the reference mean, then take a dot product with each retained direction. In the audited implementation, and . This projection has already discarded information before people enter the calculation.
Why these particular directions? They are the directions of greatest fitted variation, found from a covariance matrix. The Eigen Times covariance chapter develops that argument. For the moment, regard as a fixed measuring instrument, accompanied by its mean and version.
A synthetic three-coordinate article can make every operation visible. Set its embedding row to and its fitted reference mean to . Subtraction produces . Choose all three coordinate directions as and introduce the square naming rotation for this small illustration. The raw and named rows are then both . This identity choice isolates the scaling arithmetic in the next section; it is not a claim that the production embedding or rotation is an identity.
The notebook’s batch fixture uses the same synthetic reference mean and scales for all 45 articles. It constructs the embedding rows from the declared standardized rows and verifies that the forward operations recover those rows. The fitted production mean, directions, and scales instead come from their stated news model; fitting them anew to one person’s articles would change the measuring instrument.
A statistically useful direction need not have a readable name. Eigen applies a stored orthogonal rotation . Denote the resulting named news-coordinate row by :
Lexical loadings measure associations between words and numerical directions. Varimax is an orientation rule that makes these loadings more concentrated on individual axes. Eigen uses it to choose the rotation, helping some axes acquire recognizable word lists. The Eigen Times varimax explanation shows why this can improve interpretability.
An ontology is a reusable vocabulary of concepts and their
relationships. The word lists are labels to investigate, not ontology
definitions. Stemming shortens words by removing endings, so a label
such as databas represents a processed word form. A list
containing databas, ceo, and
interview can describe mixed coverage. An axis has two
poles: its positive and negative directions. Positive-loading words may
describe one pole better than the other. In the implementation, the
lexical naming fit uses raw scores divided by their standard deviations,
called standardized raw scores, while the stored rotation acts on
unstandardized article scores. Operations commute when changing their
order leaves the result unchanged. Scaling and rotation need not
commute, so the lexical lists are a naming heuristic rather than exact
term contributions to every final coordinate.
Some named directions vary more than others. A coordinate of 2 on one direction may be ordinary while 2 on another is unusual. Anthropology divides each named coordinate by its fitted standard deviation.
Let index named news axes, from 1 to , and let be the positive standard deviation of axis . Let be the diagonal matrix containing those scales: its row-, column- entry is , and all off-diagonal entries are zero. The superscript denotes a matrix inverse, which undoes the corresponding multiplication. Here has reciprocal scales on its diagonal, so multiplying by it divides coordinate by . Denote the standardized article row by :
For example, with standard deviations , named coordinates become standardized coordinates . When discussing a generic article we omit its subscript and write for such a standardized row; is its coordinate . We will aggregate these rows into people profiles.
An eigenvector of a covariance matrix is a direction that multiplication by that matrix stretches without turning; its eigenvalue is the stretch factor. Let index the retained raw axes, from 1 to , and let be their positive fitted variances, which are eigenvalues of the raw covariance. The coefficient is the entry in row , column of the rotation. Summing over all raw axes gives each named-axis variance:
Each rotated variance is a weighted combination of raw variances. The squared rotation coefficients supply the weights. The mean, directions, rotation, variances, and corpus identity must travel together. Axis N24 in one model is not automatically N24 in another.
With named row and fitted standard deviations , divide coordinate by coordinate: , , and . The standardized row has length , so standardization has not made it a unit row. Normalizing it would be a further operation, producing and losing its overall magnitude.
Each standard deviation belongs to a column of the fitted reference population. Each row length belongs to an individual observation. The first operation changes units; the second discards size. A zero standard deviation cannot be used as a divisor and requires an explicit modeling policy, not silent division. The formulas in this chapter assume positive fitted scales.
Marginal standardization gives each coordinate unit variance separately. Whitening would also make the covariance between different coordinates zero; we will see that these are different operations. In this illustration only, use two raw directions with variances 4 and 1. Rotate them by 45 degrees using the following matrix:
Write for the feature covariance matrix of the standardized article rows in this illustration. Both rotated coordinates have variance 2.5, but they also have covariance -1.5. Divide both by and their variances become 1 while their covariance becomes -0.6:
They now have the same marginal scale. They still move together. Correlation is covariance divided by the two standard deviations; when both variances are one, the correlation equals the covariance. Full whitening would remove these off-diagonal correlations. This pipeline performs marginal standardization, not full whitening.
For the raw row , rotating and then marginally scaling gives approximately . Scaling the raw coordinates first by and then rotating gives . The two operations answer different questions. Their order must be part of the model definition.
This also changes the geometry in which people are compared. Equal lengths in standardized news units need not be equal lengths in the original embedding. A Mahalanobis statistic measures squared distance from a reference mean while accounting for covariance, not just each coordinate’s separate scale. The sum is not generally that original statistic; see the Eigen Times section on residuals and T squared. Outside this two-dimensional illustration, the symbols retain their fitted-model meanings.
Check 3. Does a covariance matrix with ones on its diagonal necessarily describe uncorrelated coordinates? Use the two-dimensional example to answer.
Write the raw coordinate covariance as , where constructs a diagonal matrix from its listed entries. In the displayed rotation, the first named coordinate is the sum of the two raw coordinates divided by ; the second is their difference, with the first raw coordinate negative, divided by .
Because the raw coordinates have zero covariance, the first named variance is . The second uses the squared negative coefficient and has the same variance. Their covariance uses paired, unsquared coefficients: . Dividing the two named coordinates by divides this covariance by , leaving .
The row first rotates to and then scales to . Scaling first instead gives and rotating gives . These routes standardize relative to different coordinate systems, so they use different fitted scale matrices and define different maps. The named-coordinate scale here is a scalar multiple of the identity and actually commutes with the rotation. This example therefore does not test noncommutation of one fixed pair of matrices; each standardizer belongs to the coordinates whose variances it measures.
Our fictional example has people A, B, C, and D. The 45 distinct articles are grouped below only to save space: articles in a batch have identical toy coordinates and associations. They are distinct observations, not duplicate records to count repeatedly. E and L identify two source-and-decade cells. A cell specifies both the source and the decade; the letters do not imply that every early or late article belongs together.
The coordinates in the table are already standardized news coordinates . Their reference basis and scales are stipulated for this toy example, not estimated from the 45 rows.
| Batch | Cell | Distinct articles | N1 | N2 | N3 | Associated people |
|---|---|---|---|---|---|---|
| E1 | E | 10 | 2 | 0 | 1 | A, B |
| E2 | E | 5 | 1 | 1 | 2 | A |
| E3 | E | 5 | 0 | 2 | 2 | C |
| E4 | E | 5 | -1 | 0 | -2 | D |
| L1 | L | 5 | 1 | -1 | 2 | A |
| L2 | L | 5 | 2 | 1 | -2 | B |
| L3 | L | 5 | 0 | 3 | 2 | B, C |
| L4 | L | 5 | -2 | 0 | -1 | D |
There are 25 distinct articles in E and 20 in L. A has 20 associated articles, B has 20, and C and D have 10 each. Thus there are 60 person-article associations but only 45 unique articles. The ten E1 articles each name two people, as do the five L3 articles. This is why adding person counts is not a way to count the archive.
An incidence matrix records these associations. Its rows are people, its columns are articles, and an entry is 1 when the article is associated with the person and 0 otherwise. It is a bookkeeping device; it does not turn the association into a verified identity match or a social relationship. The figure compresses identical article columns into one column per batch, with its multiplicity shown below.
Let index a person and a source-and-decade cell. Let be the set of unique eligible background articles in cell , and its subset associated with person . For any set , denotes its number of members, called its cardinality. The notation means that article belongs to that set; a sum with that condition includes each member once. Recall that is article ’s standardized news-coordinate row. Let be the background mean and the person-cell mean, both rows. For nonempty article sets, divide the coordinate sums by their respective article counts:
The background is drawn from qualified candidate-associated articles, not from every article in the news corpus.
Write for the first coordinate of cell E’s background row. It is
Computing the other coordinates in the same way gives
Now consider A. Its 15 E articles have mean ; its five L articles have mean . Subtracting each cell’s background gives
The question has changed from “what is A’s coverage like?” to “how does A’s coverage differ from its available comparison coverage in each cell?” A negative coordinate means relatively less in that direction, not an absence of articles or disagreement with a topic.
For E, add each batch row multiplied by its distinct-article count. The coordinate sums are : N2 receives , while N3 receives . Divide all three sums by 25 to obtain . For L the sums are and the count is 20, giving .
A’s E numerator includes only E1 and E2: . Divide by 15 to obtain . Its L numerator is ; division by five returns . Subtract corresponding backgrounds one coordinate at a time. In E the first contrast is . The other two are and . These exact fractions explain the rounded contrast row above.
Each denominator counts the observations named by that numerator. The background counts unique eligible articles in a cell. The personal mean counts that person’s associated articles within that cell. Neither denominator counts the number of named people in the article.
A has three times as many articles in E as in L. If we pooled the articles, E would dominate. Let be the set of cells with enough articles for person ; counts retained cells. In production, a cell needs at least five articles associated with that person. For a nonempty retained-cell set, let be the balanced contrast row. The implemented aggregation computes it by averaging the supported cell contrasts equally:
In the toy example both cells qualify for every person. A’s average contrast is therefore
For comparison, weighting A’s two contrasts in the article-count ratio 15:5 would give . Neither averaging rule is a law of nature. They encode different questions. Equal supported-cell weighting prevents a densely covered cell from receiving greater weight merely because it contains more articles.
Equal cell weights are not necessarily equal source weights. A source spanning three supported decades contributes three cells, while another source spanning one contributes one. Nor does background subtraction remove every archive bias. It changes a specified comparison, within the selected coverage that is available.
The production gates require at least five articles in a person-cell and at least ten across the retained cells for a person-window. The contrast must also have nonzero numerical length. A cell with four articles is not quietly padded, and an unsupported profile is missing rather than a zero vector. These are operational support rules, not a proof of statistical reliability.
For A, the first balanced coordinate is . The other two are and . Thus the exact contrast is . A denominator of two gives one vote to each retained cell.
Article weighting gives E the weight and L the weight . Its first coordinate is . The difference comes from weights, not from any changed article coordinate.
The order of the support gates matters. Counts of six and four sum to ten, but the four-article cell is removed first. Only six retained articles remain, so that person-window fails. Counts of fifteen and four leave one retained cell with fifteen articles; its contrast alone supplies the average. Counts of fifteen and five retain both cells and twenty supporting articles. Missing means have no invented zeros. The notebook checks all three cases explicitly.
Let denote person ’s unit profile. For a nonzero contrast, normalize it as follows:
For A the result is approximately . Multiplying by a positive constant would produce the same . This makes directions comparable without turning article volume into profile length. It also discards magnitude. A small, unstable contrast can become a full-length unit vector, which is one reason to display support counts and eventually evaluate uncertainty rather than trusting the norm alone.
Let be the matrix formed by stacking those unit profiles. In the toy example its four rows, in person order A, B, C, D, are
Each row describes one person’s direction of relative coverage in the fixed news space. We have not yet learned a people pattern.
Check 4. Why would counting E1 once for A and once for B inside the background change the question? Why is it nevertheless correct for E1 to enter both individual profiles?
The exact balanced contrast has squared length
Its length is , approximately 1.72440. Dividing each entry by that length cancels the common denominator, so . Its squared coordinates add to .
This cancellation explains both the convenience and the loss. We can recover the direction from the normalized row, but we cannot determine whether the contrast originally had length 1.72440, twice that length, or half that length. Nor does the unit row remember that A had twenty supporting articles. Counts, magnitude, and article links must remain separate records.
Let now count the fitted people, so has shape . Its rows have unit length, but they need not average to zero. Denote their mean row by . In the example it is
Let denote a column of ones, so repeats the mean row for every person. Subtract it from and call the resulting centered matrix :
Let denote person ’s row of . For A,
This second centering asks how A’s normalized coverage differs from the fitted population of normalized person profiles. The earlier subtraction compared articles inside a cell. The reference objects are different.
Let denote the people covariance. Using the fitted-person count , compute it as
Here and each index news coordinates from 1 to . Entry is the average product of those two centered coordinates across people. A large positive product tends to occur when both coordinates deviate in the same direction; a negative product occurs when they deviate oppositely. The diagonal entries are coordinate variances.
Each person contributes one row. A person supported by 10,000 articles does not enter the covariance a thousand times more heavily than a person supported by ten. More articles can affect the quality of a row, but they are not its weight in this fit. This is equal-person PCA: principal component analysis of normalized, centered coverage profiles.
Let index a people direction, its unit eigenvector column, and its covariance eigenvalue. Each pair satisfies
Multiplying the direction by the covariance stretches it by without turning it. As explained in the Eigen Times eigenvector chapter, the largest eigenvalue identifies the direction of greatest variation, and the next identifies the greatest variation perpendicular to it.
For the original four-person fixture, the first centered coordinate is approximately 0.3411 for A, 0.8207 for B, -0.3459 for C, and -0.8159 for D. The N1 variance therefore adds their four squares and divides by four, giving approximately 0.393847. The N1–N3 covariance instead multiplies each person’s first and third centered entries and averages the four products, giving approximately 0.138797. Both calculations use one product per person. The notebook prints the complete centered table and checks the covariance against independent reference entries.
Set the four-person matrix aside for a small covariance illustration. Let . The star marks this separate synthetic table. Let and be unit columns. Multiplying the first gives . Multiplying the second gives . They are eigenvectors with eigenvalues 3 and 1; their dot product is .
For a unit column , the scalar is variance along that direction. Along the first coordinate axis it is 2. Along it is 3; along it is 1. To see why 3 is the maximum, express any unit direction as , where are scalar coefficients satisfying . Its variance is , no larger than 3. The larger-eigenvalue direction captures the most spread because of this numerical property, not because its words are necessarily more important.
Multiplying an eigenvector by -1 leaves its line and eigenvalue unchanged. Its coordinates and resulting scores both reverse signs. A sign convention is useful for reproducible displays; it does not change the represented variation.
The toy covariance has eigenvalues approximately , , and . Keeping the first two retains
of the between-person variance, or 81.14%. This is a reconstruction statement about this four-person cloud. It is not an accuracy estimate, an identity confidence, or the fraction of articles explained.
Let count the retained positive people directions; it is 2 in this toy example. The bare is a count, distinct from the contrast row . The retained direction index runs from 1 to . Put those eigenvectors, ordered by decreasing eigenvalue, into the columns of the loading matrix :
Its rows correspond to N1, N2, N3; its columns correspond to people patterns P1 and P2. The columns have length one and are perpendicular. Their signs are oriented so the largest-magnitude news loading is positive, matching the implementation’s sign convention. Flipping a column and all its scores would describe the same geometry. When two eigenvalues are close, a small change in data can rotate their directions substantially even while their combined subspace changes little. Fixing signs does not remove that ambiguity; it is another reason to keep model versions explicit.
The three fitted eigenvalues sum to approximately 0.9596456. This is also the sum of the three diagonal entries of , called its trace: total centered squared length per person, expressed in either coordinate system. The first two sum to approximately 0.7786802. Their ratio is 0.8114248; multiplying by 100 expresses it as a percentage. The omitted fraction is about 0.1885752.
The loading columns are unit directions, but a row of need not have length one. Similarly, the score rows need not be unit rows after population centering and projection. Keeping these distinctions prevents a loading, a coordinate, and a cosine from being treated as interchangeable numbers.
Let denote person ’s people-coordinate row. Project its centered profile onto the retained columns:
For A, the first coordinate is approximately
The second uses the second column of . The resulting rows are:
| Person | P1 score | P2 score |
|---|---|---|
| A | 0.5109 | -0.5334 |
| B | 0.5828 | -0.1178 |
| C | 0.0813 | 0.8807 |
| D | -1.1750 | -0.2294 |
P1 is a pattern, not person D, although D has the largest absolute P1 score in this small example. A negative score identifies the other pole of a direction. It does not say that D is opposed to the other people.
In the published models, there are 60 news coordinates and at most 24 directions with positive eigenvalues. The 217-person Hacker News model retains 84.54% of between-person variance; the 61-person general-news model retains 90.86%. These separately fitted percentages should not be read as a contest between corpora. There is no second varimax naming step and no whitening of people scores. The people directions are orthonormal in standardized news coordinates, not generally when mapped back into the original embedding geometry.
A’s first score adds approximately . Its second coordinate is
These are projections of the same centered row onto different direction columns. Small differences in final digits arise if we multiply displayed four-decimal entries rather than the full-precision fixture.
The first score column has variance 0.4970 and the second 0.2817, the eigenvalues already reported. They are not divided by the square roots of those variances in this model. Dividing them would whiten the retained scores and would change distances and cosines; that is a different geometry.
The first row of the toy is . It tells us how N1 loads on P1 and P2. Starting from N1, the interface can rank people patterns by the absolute sizes of those two entries. P1 comes first, with a positive loading; P2 comes second, with a negative loading.
The first column of is . It tells us how P1 combines N1, N2, and N3. Starting from P1, the interface ranks news axes by the absolute sizes of these entries: N1, then N3, then N2.
The coefficient linking N1 and P1 is exactly the same number in both views. No second model is needed to construct the reverse index. One index exposes rows of ; the other exposes columns.
For the diagram, let index news axes, retained people patterns, people, and articles. The scalar is row , column of ; and are coordinate of the article and person-profile rows; is coordinate of the people-score row. The diagram uses parentheses for these same indices: for example, means . Magnitude means absolute value, so either sign can produce a large magnitude.
This does not make ranks reciprocal. N3’s strongest loading is P1, but N3 is only the second strongest news loading of P1. Nor are the entries probabilities. Some are negative, and the absolute values in a row or column need not sum to one.
Starting from news axis N1, one can rank people by their direct profile coordinate . B comes first in the toy example, with 0.9484. This asks whose normalized relative coverage extends farthest along N1, considering either pole.
Starting from people pattern P1, one can instead rank people by . D comes first, with 1.1750. This asks whose centered profile extends farthest along that learned combination of news coordinates.
Those rankings differ because the questions differ. Neither is the same as finding the nearest neighbor of a selected person, which uses the entire row of people scores. A useful interface names the selected axis, mode, pole, model, and time window so a reader knows which question produced the list.
The audited Hacker News model supplies a real example. The loading
joining news axis N56 to people pattern P19 is
in both stored indexes. N56’s terms include develop,
interview, databas, ceo,
manag, and github. The shared coefficient is a
precise navigation fact. Calling it “the database CEO axis” would
discard the mixed meaning of the direction and the evidence needed to
identify a current CEO.
Check 5. In the toy matrix, N3 ranks P1 first. Where does P1 rank N3? Explain why the two answers are consistent.
When discussing one person without naming them, write , , and for their unit profile , centered profile , and retained score row . The rows have coordinates and has . Let be the reconstruction of ; the hat denotes an approximation returned from retained coordinates. After calculating , multiply by the transpose:
Adding gives the corresponding reconstruction of the unit news profile. It does not give back the original articles.
An orthogonal projector keeps a vector’s component in a chosen subspace and discards its perpendicular component. Let be the forward-and-return matrix, defined by . A symmetric matrix equals its transpose; an idempotent matrix has the same effect when applied twice as once. Because the columns of are orthonormal, has both properties:
The first equality expresses symmetry and the second idempotence. Together these identify an orthogonal projector. Once a row lies in the retained plane, projecting it onto that plane again changes nothing.
The residual is the difference between the original centered row and its reconstruction, . It is perpendicular to the retained directions. Consequently,
This is the familiar right-triangle rule, now applied to components of a vector. In the fitted toy example, A’s discarded squared length is approximately 0.2650. The example script checks the decomposition and the orthogonality of the residual using unrounded values.
To see the loss without decimals, set aside the fitted toy matrix for a moment. Let denote this deliberately chosen matrix; the star distinguishes it from the fitted :
This is an illustration, not another fit. It retains the shared direction of the first two coordinates and retains the third coordinate separately. A query goes forward to and returns as . The first-versus-second distinction was discarded.
A linear map preserves addition and scaling; multiplication by a fixed matrix is an example. An adjoint transfers a linear map to the other side of a dot product; for real matrices with Euclidean dot products, it is the transpose. Least squares means minimizing the sum of squared coordinate discrepancies. A Moore-Penrose pseudoinverse is a generalized inverse that gives least-squares solutions, choosing the shortest solution when several are possible. Because has orthonormal columns, is both its adjoint and its pseudoinverse. It returns the least-squares reconstruction of a centered profile in the retained subspace. A two-sided inverse would undo the map in both directions for every input; this transpose cannot do that. With 24 retained directions out of 60, there is no way to recover every possible original 60-coordinate row.
The first retained coordinate of is ; the second is zero. Returning multiplies the first column by , giving . Subtraction leaves . Its squared length is , and its dot products with both retained columns are zero. The returned part also has squared length , so .
Any alternative reconstruction inside the retained plane can be written , where and are real scalar choices. Its squared discrepancy from is
Squares cannot be negative, so the minimum is , achieved at and . This proves the least-squares claim for this example without asking the reader to trust an inverse formula.
Even retaining every people direction would not recover an article list from a person profile. The earlier pipeline also reduced a 384-coordinate embedding to 60 news coordinates, averaged articles, subtracted backgrounds, and normalized lengths. Many different inputs can give the same output after those operations.
Navigation is valuable without being invertible. It provides linked views of a representation and routes back to stored evidence. The article links and factual sources are retained separately precisely because the vector cannot reconstruct them.
Singular value decomposition (SVD) factors a matrix into two sets of orthonormal directions and nonnegative scale factors called singular values. For a longer prerequisite, revisit the Eigen Times SVD chapter. The operator selects the smaller of two numbers. For the centered matrix , let . In the thin decomposition, is an matrix of left directions, indexed by people, and is a matrix of right directions, indexed by news coordinates. Both have orthonormal columns. Let be the diagonal matrix of singular values in descending order. Then
Recall that counts retained positive people directions. Let and contain the first columns of their respective matrices, and let contain the leading diagonal block of . Thus has shape and shape . Using the same ordering and orientation as the people eigendecomposition gives and
The scores include the singular values; they are not just the columns of . Each squared singular value divided by equals the corresponding eigenvalue of the earlier people covariance. The right directions come from , while the left directions come from . These are two sides of the same centered profile matrix. An adjacency matrix instead records which graph nodes have edges between them. Neither calculation uses that matrix from Anthropology’s relationship graph. A graph community is a group of nodes relatively densely linked to one another. Communities and people patterns can be interesting to compare, but they are not the same construction.
Check 6. For above, project forward and back. Compare its returned row with the return from . What information can the map no longer distinguish?
Use a separate centered table with four rows , , , and . Each column sums to zero. Multiplication gives : the first column’s squares sum to two, the second’s to eight, and their paired products are zero.
The largest right direction is therefore , followed by . The singular values are the square roots of eight and two: and . Projecting the rows onto the first right direction gives . Divide that column by to obtain its unit left direction . The second left direction is .
Now multiply back: the first left column times times the first right row restores only the second coordinate; the second term restores only the first coordinate. Adding the two restores every entry of . With four observations, the covariance eigenvalues are and . Keeping only the first singular direction retains of the total squared length. The discarded first-coordinate column has squared length two. This shows explicitly why scores contain singular values: the left direction has unit length, while its score column here has length .
The native OCaml teaching routine obtains a small SVD through a symmetric Gram matrix, the product . That is adequate for these tiny well-scaled demonstrations; it is not presented as a numerically preferred algorithm for a large production fit. Both notebooks verify reconstruction, orthonormality, and the independent singular values.
Here and index two people, and are their nonzero score rows from the same model and coverage window. The person index is distinct from a background row such as . Cosine similarity, written , is their dot product divided by the product of their lengths:
It compares direction rather than length and is undefined if either row is zero. A value near 1 means similarly directed rows, 0 means perpendicular rows, and -1 means opposite directions in the specified centered representation. Negative similarity is not evidence of competition, disagreement, or hostility.
Anthropology’s neighbor view uses complete retained people-score rows within the same model and window. Do not mix a recent profile from one model with an all-history profile from another and call their raw dot product a meaningful similarity. Even matching dimensions are insufficient: the coordinates must share their meanings.
Consider three synthetic two-coordinate score rows: a query and candidates and . Coordinate distance means the length of the difference row. From to the difference is and the distance is 9. From to it is and the distance is 1. Thus coordinate distance prefers .
Cosine asks another question. The dot product of and is 10, their lengths are 1 and 10, and the cosine is . The dot product of and is 1, their lengths are 1 and , and the cosine is , approximately 0.7071. Directional similarity prefers . No contradiction exists: lies farther out along exactly the same direction.
If both rows are normalized, squared coordinate distance becomes twice one minus cosine. To check this, expand the squared difference into the two squared lengths minus twice the dot product. Each squared length is then one. For and normalized , this gives , approximately 0.5858. This equivalence requires unit rows in the same geometry; the original people-score rows are not generally unit rows.
A single selected coordinate can be identical while all the other coordinates differ. Ranking by one axis therefore cannot replace a full-row neighbor calculation. Centering also changes the origin from which direction is measured, so a cosine of uncentered profiles answers a different question from a cosine of their centered people scores.
The toy example makes the danger visible. A and B have a cosine of only 0.0392 between their full centered three-coordinate profiles. Their cosine becomes 0.8211 after projection into the two retained people directions.
No arithmetic error occurred. The discarded direction contained much of their difference. The retained plane makes them look more alike. Keeping 81.14% of the population’s total variance did not preserve every pairwise angle.
The same check is useful in the real model. In the frozen Hacker News model, LeCun and Bengio have a retained-space cosine of 0.9587 and a full centered news-space cosine of 0.9118. Ellison and Siebel have 0.7198 and 0.5800 respectively, with 416 and 11 candidate articles. These values describe this corpus and representation; the much smaller support for one member deserves attention.
An appealing analogy can also fail. Zuckerberg and Gates have retained-space cosine 0.0782 in Hacker News and -0.1137 in the separate general-news model. Being prominent technology founders does not require the archives to cover them in the same way. A research system should let the evidence correct the analogy.
Two profiles can be close because the same articles mention both people, because different articles discuss similar subjects, or because the archive repeatedly frames them in a similar way. A cosine alone cannot distinguish these explanations. One useful sensitivity analysis would remove their shared articles and recompute the comparison. That experiment was proposed in the paper; it was not part of the published evaluation.
For a nonzero centered row , retained energy is the fraction of its squared length preserved by projection, . Support counts, full-space baselines, retained energy, dates, and example articles help a reader interpret a score. Calibration means agreement between predicted probabilities and observed frequencies. These aids do not turn a neighbor score into a calibrated probability of collaboration. A documented relationship still requires its own source. The named examples here demonstrate navigation, not a statistically representative assessment of retrieval quality.
Check 7. Does retaining 81.14% of total variance guarantee that every neighbor cosine changes by less than 18.86 percentage points? Use A and B to test the claim.
A person can receive different coverage this year than ten years ago. We want to observe that movement without changing the ruler at the same time. Let index a coverage window, such as a decade or a recent period. Write for person ’s unit profile built from that window’s eligible articles, and for its people-coordinate row. A comma in separates the same indices without changing their meaning. Fit the people mean and loading matrix once on the all-history profiles, then use them for every supported window:
The article associations, supported cells, cell means, and cell backgrounds are recomputed for that window. Write for person ’s cell- mean in window , and for that cell’s background mean, both rows. They use the same averaging rule as before, restricted to articles in the window. The people mean and directions remain the fitted all-history reference.
This makes score differences interpretable within a model version. Let and be two supported windows for the same person, and let denote the Euclidean distance between that person’s two score rows. A simple movement statistic is
It is a distance between two supported coverage profiles. It is not, by itself, a measure of a person’s career change. A source mix can change, a name attribution can be corrected, or the comparison background can shift.
Return to A’s two cell contrasts. Hold both contrasts fixed, but consider two synthetic count mixtures: 90:10 and 10:90. These are a sensitivity experiment, not two additional observed periods in the 45-article corpus. In each mixture both cells satisfy the support threshold.
Weighting the contrasts by those counts, normalizing, and applying the same fitted mean and gives people scores of approximately in the first mixture and in the second. Their apparent movement is about 0.319. Yet neither cell-specific contrast changed.
Equal supported-cell weighting keeps A at its original people coordinates in both mixtures. This illustrates one composition effect the balancing rule can reduce; it does not establish invariance to changes in the actual within-cell article distributions.
It does not remove every composition effect. If a cell falls below the support threshold, the retained set changes. If changes, a person’s contrast changes even if its own articles do not. Two points can be drawn in one coordinate frame while still differing in their comparison backgrounds.
An article’s publication date is one clock. The newest matched article used in a historical profile is another. The valid date of a documented role is a third. A fourth piece of context, the model version, identifies the numerical ruler rather than a date of an event.
The projection watermark records the latest article date represented in the local projected corpus. In the frozen people snapshot, matched profiles end on 4 September 2026 and the local projection watermark is 5 September. The current-news overlay is dated 1 October. Recent-36-month and latest-90-day profile windows are anchored to the builder’s latest matched article date, not automatically to the reader’s present clock. Publishing a larger identity catalog on 2 October did not refit those profiles or update every watermark.
An absent recent profile means insufficient supported coverage. It must not be plotted at the origin as if it were evidence that the person became average.
For each person-cell-window, let count its distinct eligible associated articles, and let be the sum of their standardized rows . For , the previously defined person-cell mean is . These sums are unrelated to the singular-value matrix . Keep a separate deduplicated count and sum for the cell background. New qualified articles add to these statistics; expired rolling-window articles subtract from them; corrected attributions remove an old contribution and add the corrected one.
For these mean updates, the count and coordinate sum are sufficient summaries: they retain everything needed to recompute the mean without rereading unchanged articles. The Eigen Times streaming chapter explains this pattern more generally. The added complication here is the shared background: changing one background can require refreshing every person who uses that cell. Updating only the newly mentioned person would leave inconsistent comparisons.
After updating the means, reapply support gates, subtraction, equal-cell averaging, normalization, and the fixed projection. For one resulting supported window, abbreviate its centered profile as , suppressing the window index for this calculation. Let denote its discarded squared length; this scalar is distinct from the SVD’s right-direction matrix . Define the residual statistic as
It measures how much of that centered profile lies outside the retained people subspace. A sustained rise could motivate evaluating a new basis. It is not an automatic alarm that a person changed, and this uses people-profile geometry rather than silently inheriting every threshold from article-level Eigen Times statistics.
This refresh algorithm is proposed. The deployed latest-news overlay does not perform it. New versions should preserve their predecessors and be evaluated on later observations kept out of fitting, with adequate support and source-sensitive uncertainty checks before making claims about change.
Check 8. Can a person’s profile change when no new article has been associated with that person? Identify two ways the aggregation can produce that result.
Cell E’s unique background has count 25 and coordinate sum . Suppose one newly eligible article has row and is associated with an additional person outside the four fitted toy profiles. The background still consists of candidate-associated articles; none of A–D’s personal article sets changes. The count becomes 26 while the sum stays , so the background becomes , approximately .
Every existing profile using E now subtracts that new background. It can move despite having no newly associated article. Removing the same article restores the count to 25 and the original background exactly. A’s personal E sum is separately . Correcting one attribution by replacing with changes it to while leaving its count at 15. Addition, expiration, and attribution correction are different changes to the stored summaries.
A topic coordinate tells us how strongly an article lies along a direction. Eigen significance supplies another scalar: a story’s attention score, based on the original story-ranking method. Anthropology keeps that signal alongside the vector rather than silently using it as a weight in the equal-person PCA fit.
The source’s story score is allocated to its member articles and normalized by the scored source’s total in the relevant period. Let be article ’s resulting nonnegative attention share. Recall that is coordinate of the standardized article row . Let denote an attention-weighted topic score, not an ordinal position in a list. One selectable latest-article ranking uses
The absolute coordinate measures topic strength on either pole; measures allocated attention. The product deliberately answers a combined question. A reader can also choose topic-only ordering or inspect one pole separately.
For a tiny illustration, suppose article X has topic strength 2 and attention share 0.01, while article Y has topic strength 0.5 and attention share 0.10. Topic-only ordering puts X first. The combined scores are 0.02 and 0.05, so attention-weighted ordering puts Y first. Nothing about the topic coordinates changed.
Use a separate synthetic story with attention score 12 and three member articles. An equal allocation rule would give each article units. If that source’s scored total in the period is 40, each article’s attention share is . Their shares sum to 0.3, the story’s share . This illustrates one declared allocation rule; it does not assert that every source allocates every story equally.
The topic product then has two separate inputs. An article with standardized topic coordinate -2 and attention share 0.1 receives an either-pole score . A positive-pole query would need an explicit sign restriction instead. Allocation, source normalization, pole selection, and multiplication each change the question, so each belongs in the ranking’s definition.
If a 0.10-share article mentions A and B, that coverage can contribute to both people’s attention histories. Their combined credited share can exceed the article’s 0.10. These histories describe overlapping coverage, not a partition of total human importance.
Source normalization matters too. Averaging shares over active sources, including sources with no mention of a person, differs from pooling raw scores from a large and a small outlet. State the denominator before interpreting the percentage. The Eigen Times discussion of story energy and ranking provides the deeper article-level background; the person layer adds attribution and aggregation rather than redefining that original significance.
Suppose two active sources have total attention scores 100 and 1,000. The selected person’s associated coverage receives 10 units in the first and zero in the second. Source-specific shares are and . Giving each source one vote yields . Pooling the raw scores instead gives , approximately 0.00909. Omitting the zero-mention source would produce yet another quantity, 0.10.
None of these denominators can be recovered merely by reading a displayed percentage. State the active source set, period, attribution rule, and aggregation rule. The two-person credit on one 0.10-share article likewise totals because the histories overlap; it does not create another article or more source attention.
The natural question “show today’s database news against today’s database CEOs” joins several tasks. News coordinates find topic-relevant articles. Dated role assertions identify people recorded as CEOs at the requested time. Supported person profiles locate those people in a historical coverage geometry. Qualified article-person associations establish which current articles actually concern them.
The current overlay can show relevant articles alongside historical profiles, but the audited overlay has zero newly qualified article-person joins. Proximity between an article and a profile is therefore a lead to inspect, not proof that the article concerns that person. Similarly, an old article calling someone a CEO does not establish that the role is current.
The Hacker News people model remains in its hn1 basis.
New article embeddings can be measured in that fixed basis while
carrying a scalar attention share from hn2. This is
coherent only when the separation is explicit: the scalar is an overlay;
it does not substitute a different set of coordinate axes. Model,
corpus, dates, support, and role evidence remain visible parts of the
query.
Check 9. Article X is more topic-aligned than article Y. Must X rank higher after attention weighting? Which separate evidence would be needed before calling either one news about a selected person?
An eigenvector is a direction learned from variation. An ontology concept is a named idea that people agree to reuse. “Databases” can remain a useful concept while a database-related news direction changes between archives. Linking the two would let a reader follow one subject through Eigen Times, Eigen Hacks, and Anthropology without pretending that the sites share one coordinate system.
This chapter separates three stages: the personal tags already implemented in the 2 October 2026 paper edition; a proposed, reviewed table connecting concepts to axes; and a proposed learned model whose predictions would require evaluation.
The ontology picker lets a reader choose a concept and attach its identifier to a canonical node. Searching a label and clicking it records an annotation. It does not rotate the news basis, retrain a person vector, or establish a public factual relationship.
The pinned ontology has 157 concepts in 11 areas; 61 concepts have multiple parents. This is a polyhierarchy: one concept can belong under several broader ideas. For an invented illustration, “distributed databases” might belong under both “databases” and “distributed systems.” These parent links describe meaning. They need not form perpendicular directions, and a person can have several tags at once.
An axis label serves a different purpose. Its strongest stemmed terms
help a reader interpret one fitted direction. A list containing
databas, ceo, and github does not
define a clean databases category. Moreover, the negative pole needs
inspection too: it is not automatically “not databases.” The sign of an
eigenvector is a convention; a reviewed meaning must be attached to the
appropriate pole and model version.
For the mechanics behind named directions, return to From eigenvectors to named axes.
A crosswalk is a table of correspondences. Imagine a proposed row saying: in this particular Eigen Hacks basis, the positive pole of this axis has useful evidence for the databases concept. The row should retain the corpus, axis and sign, basis fingerprint, ontology version, representative articles, method, and review status.
The bridge should be many-to-many. Several axes may help retrieve database articles, and one mixed axis may connect to several concepts. A reviewer can leave a proposed correspondence unassigned. Shared concept identifiers then become destinations across the three sites, while each route retains its original numerical coordinates.
Concept-to-concept mappings also have different strengths: “exact,” “close,” “broader,” “narrower,” and “related” do different jobs in the W3C SKOS mapping vocabulary. An axis would first need a reviewed semantic interpretation before those conceptual distinctions could be used responsibly. A familiar-looking word is insufficient.
A later model could learn from reviewed articles. Let count the concepts reviewers label, and let be one article’s standardized news row. Let be a slope matrix: each slope measures how changing an input coordinate changes one concept’s linear score. Let be an intercept column supplying baseline scores at zero input. The notation is another form of the transpose symbol . Denote the resulting prediction row by . A simple linear predictor multiplies inputs by slopes and adds the baseline:
Each column of asks how the news coordinates contribute to one concept. Its predictions are scores; no probability interpretation has yet been established.
Consider an entirely illustrative two-coordinate, two-concept model:
The first prediction is . The second is . Nothing went wrong when these escaped the interval zero to one: they are scores from an unconstrained linear rule. We have not made a probability model.
To learn the coefficients, let count reviewed training articles, distinct from the fitted-person count used earlier. Stack their standardized rows into the matrix and their matching concept labels into the matrix . Row order agrees across the two matrices. For a binary label, an entry is one when reviewers judged the concept present and zero when they judged it absent. Missing judgments must remain missing, rather than becoming zeros.
A ridge fit balances squared prediction error against a penalty on large slopes. Let be the penalty strength. For a matrix, the Frobenius norm, denoted by double bars with subscript , is the square root of the sum of squared entries. Thus its square adds those squared entries. Here is a column of ones, so repeats the baseline row for every article. The following objective assumes every target entry is observed; partially reviewed targets require a separately specified objective summing only observed errors. Choose and to minimize
The first term adds squared prediction errors; the second penalizes large slopes. The intercept is deliberately unpenalized. A held-out validation procedure, using articles reserved from parameter fitting, would choose rather than a convenient-looking result on the training articles.
The toy coefficients above can actually be fitted. Take four fictional, fully reviewed article rows:
Each concept’s average label is 0.5, giving the intercept. Each coordinate column has sum of squares four and centered label-product sum two. With , its fitted slope is ; the cross-slopes are zero. Training predictions are 0.1 or 0.9. The new row then extrapolates beyond the training coordinates, producing . Training fit and behavior on new inputs are different tests.
Each corpus would need its own fitted coefficients. Shared target identifiers make the outputs comparable in meaning; they do not make the input bases interchangeable. Concepts overlap, so need not be orthogonal or even square. This semantic transformation does not inherit the reconstruction guarantees of .
A derivative describes how a scalar output changes per small change in a scalar input. Start with a function , where is a real scalar, and let be a nonzero change in that input. The difference quotient is
As approaches zero, this approaches . That limiting slope is the derivative, written . At , changes of 0.1 and 0.01 give difference quotients 4.1 and 4.01, approaching 4. A derivative is a local rate of change; it is not the function’s output, which also happens to equal 4 at this particular input.
Return to the first ridge target. Its centered labels are , and the first input coordinate is . Temporarily call its scalar slope and set the other slope to zero. Each of the four squared prediction errors is . With penalty , the scalar objective, named , is
Differentiate its expanded terms: . At an interior minimum a small move in either direction cannot improve the value, so the local slope must be zero. Solving gives . Complete the square to verify this is a minimum: . Its minimum value is 0.2. At the value is 1; at it is 0.25. Ridge accepts a little prediction error to reduce the coefficient penalty.
For any nonnegative penalty , this example instead has derivative , giving . The input columns have zero cross-product, so the two slopes separate cleanly. In a general table they interact. Let subtract each input column mean from , and let subtract each target column mean from . Let here be the identity. The stationary equations for these centered tables are . Solving this equation finds the slopes; restoring the target means minus the mean input’s fitted contribution gives the unpenalized intercepts.
The notation denotes a partial derivative: change coefficient while holding the other coefficients fixed. The gradient collects all those partial derivatives in the coefficient table’s shape. For the full squared-error objective its slope gradient is . Setting the whole gradient to zero yields the same stationary equations. The notebooks check the scalar derivative with small finite changes, the completed-square minimum, and the fitted matrix equations independently.
First define what is being predicted. One possible event is: “Under this annotation protocol, reviewers judge that this article substantially discusses databases.” It is not “this person is a database person.”
A logistic predictor maps a linear score through a smooth S-shaped function whose output lies between zero and one. Calibration means agreement between predicted probabilities and observed frequencies; the logistic bound alone does not establish it. The distinction is developed in Guo and colleagues’ study of calibration.
A reliability diagram offers a simple check. In an invented test set, collect 100 articles assigned probabilities near 0.8. Suppose reviewers mark 55 as positive. Plot their mean prediction near 0.8 horizontally and the observed fraction 0.55 vertically. A well-calibrated bin would lie near the diagonal, where those numbers agree. This example is hypothetical; Anthropology has not reported such a calibration experiment.
Repeat across probability ranges, report the number of articles in each bin and uncertainty, and inspect relevant sources and periods separately. Small bins can fluctuate substantially. Prevalence is the fraction of articles with a positive target. A model that predicts this overall fraction for every article can be calibrated yet poor at ranking individual articles. Discrimination is its ability to distinguish positive from negative cases. Calibration and discrimination therefore need separate evaluation, with articles reserved for testing rather than reused for fitting or tuning.
Checkpoint. The toy linear model predicts 1.3. Can the interface print “130% confidence”? Answer: No. It can display a labeled score. A probability needs a defined target, a suitable model, and held-out evidence that its estimates behave like probabilities.
Let denote a real-valued linear score and let denote a modeled probability for one defined binary article label. Write for the exponential function, with base ; its output is positive for every finite real input. The logistic function, written in this section only, is
At zero, , so . At two, , so . At minus two, . The denominator always exceeds one, placing the result strictly between zero and one for finite scores. This function’s is not the earlier standard deviation .
To fit such outputs, the notebook uses the first column of the four-row target table. Let be a two-entry slope column and a scalar intercept, giving score for training article index . Let be its reviewed zero-or-one target. Binary log loss for one row is , where is the natural logarithm, the inverse of the exponential. A correct high-probability prediction gets a small loss; a confidently wrong one gets a large loss. All four losses are averaged. The demonstration adds with penalty , keeping the intercept unpenalized.
We need three derivative rules before differentiating this model. The derivative of is ; the derivative of is for positive ; and the derivative of the reciprocal with respect to nonzero is . The chain rule multiplies the local rates when one function is placed inside another. We use these rules here and verify the resulting derivatives by small finite changes in the notebooks.
To differentiate the logistic function, name its denominator . The inner negative sign has derivative -1, so . Applying the reciprocal rule and the chain rule gives
One factor is ; the other is . Hence the derivative is . At , its value is .
Now let name a single row’s log loss and let be its fixed binary target. The notation writes its derivative with respect to . Differentiating with respect to its probability gives
Composing this loss with the logistic score multiplies by , cancelling the denominator and leaving . A slope changes the score at a rate equal to its input coordinate; the intercept changes it at rate one. Thus the slope gradient is the average of plus ; the intercept gradient is the average of .
Start all slopes and the intercept at zero. Every prediction is 0.5, so the four errors are . Multiplying those by the first input column and averaging gives . The second slope gradient and intercept gradient are both zero. A gradient-descent step subtracts the gradient times a positive step size; with step size 0.2 the new slopes are and the intercept stays zero. Positive first-coordinate articles now receive .
The native notebooks repeat that update 2,000 times and check a small final gradient and a separate reference slope. This verifies the training calculation for declared synthetic inputs. It supplies no real-world held-out calibration result. In the separate hypothetical reliability bin, observed frequency differs from 0.8 predicted frequency by 0.25. Constraining outputs to a probability range and testing their agreement with observations remain distinct tasks.
A person profile has been background-adjusted and normalized. It describes the direction of distinctive coverage. An article row has undergone neither of those person-level operations. Their equal number of coordinates does not make them the same kind of observation.
For a person score row , the earlier chapters reconstructed an approximate profile as . Use the people basis and article predictor from the same corpus and news-coordinate version. Let denote the score row obtained by applying the article predictor to this reconstructed profile. Substitution gives
The product has shape : it describes how moving along each of the people directions would change those linear scores. This is a valid algebraic identity. Its interpretation as a useful prediction about a person remains an untested use outside the article model’s training distribution. Neither restoring the mean nor adding the intercept solves that problem.
Another possible quantity is the fraction of a person’s associated articles predicted to discuss a concept. That answers a coverage question, with uncertainty from both article attribution and classification. For a nonlinear predictor, predicting an average row and averaging individual predictions can differ. A future system must specify which quantity it means and evaluate it on independently reviewed person-level examples.
Substitute the reconstructed row into the article rule directly: . Distribute multiplication over addition to obtain , then group the last matrix product as . Grouping changes where we calculate intermediate tables, without changing their compatible dimensions or result. The notebook uses a separately declared three-input, two-output synthetic slope table because the earlier two-input ridge table does not fit the three-coordinate people fixture.
For a nonlinear example, take scalar article scores zero and two. Their average is one, so predicting after averaging gives . Predicting separately and then averaging gives . The difference is approximately 0.040660. Thus even the arithmetic choice of when to aggregate changes the output. A defined person-level target and separately reviewed evaluation remain necessary whichever quantity is chosen.
Suppose two corpora contain reviewed representations of the same people. An anchor is a pair of rows known to describe the same entity in the two corpora. An orthogonal Procrustes fit tries to rotate or reflect one set of such rows toward the other. For this alignment only, let count paired anchors and count coordinates in each representation. Let and be their centered tables, with matching entities in the same row; here is an anchor table, not the preceding concept-target matrix. Let be a alignment matrix, distinct from the SVD factor . The identity now has shape . The fit minimizes subject to the orthogonality constraint .
The constraint preserves distances within the chosen coordinate geometry. It cannot correct every difference in populations, source coverage, or scaling. Shared names are not enough; the anchors must be correctly identified and substantively comparable.
The published models share 60 canonical people. Sixty centered rows in 60 dimensions have rank at most 59: subtracting their mean makes the rows sum to zero, creating a dependence. Consequently, they cannot determine a full 60-dimensional alignment from independent evidence in every direction. Reserving anchors for testing leaves still fewer for fitting. A lower-dimensional or regularized approach is possible, but its assumptions and performance need evaluation. See Matching axes for the prerequisite distinction between comparing directions and assigning correspondences.
Use four centered synthetic anchor rows , , , and for . Stipulate a known map only to construct test data . It maps a row to . The first pair is therefore and . The two tables have zero column means.
For the fit, multiply paired columns to obtain . Denote the SVD of this product by , where and are two-by-two orthogonal direction matrices and contains its nonnegative singular values. The Procrustes solution is . The notebook computes it from the paired tables, then checks that it equals the stipulated map and gives zero training discrepancy.
Why this choice? Expanding the squared discrepancy yields the squared lengths of and minus twice the trace of . The first two quantities are fixed under an orthogonal map. In the singular coordinate frame, the trace is largest when each nonnegative singular value is multiplied by one rather than a smaller diagonal coefficient of an orthogonal matrix. The choice achieves that maximum, and therefore the minimum discrepancy. This permits reflections as well as rotations, exactly as the stated constraint allows.
Reserve a fifth row from fitting. Its stipulated target is ; the estimated map reproduces that target with zero error. This noiseless test proves the toy algebra, not a real alignment’s usefulness. Actual held-out anchors could disagree because coverage, scale, or meaning differs between corpora.
The rank shortage has an equally small example. Center three two-coordinate rows: their sum becomes , so the third row is the negative sum of the first two. There are at most two independent centered rows. With sixty centered anchors, the same dependence leaves at most fifty-nine independent rows. The notebooks also compute this bound using a synthetic centered sixty-by-sixty identity table; it is not a measurement of the real anchor matrix.
The mathematics makes research paths visible. Following them well means keeping track of what each step establishes. This chapter uses the paper’s 2 October 2026 edition and its unchanged 1 October vector snapshot. The links are live entry points; the numbers below describe that dated audit.
Open Jeff Dean’s Anthropology profile. This is the identity and source-record starting point. The profile’s signal links include several news directions; no single one is “Jeff Dean’s eigenvector.”
Open the explicit Dean
view on news axis N24 and people pattern P2. The model selector
matters: this is the Eigen Hacks hn1 people model. Its
all-history Dean profile uses 253 qualified candidate
articles. The displayed news coordinate is approximately
+0.433 and the people coordinate
+0.371. N24’s leading terms include learn,
machin, model, and deep. The two
scores measure different projections; they are not probabilities or two
estimates of the same number.
Inspect neighbors. The audit records Sanjay Ghemawat at cosine 0.7514 with 31 candidate articles, Andrew Ng at 0.7413 with 394, and Noam Shazeer at 0.7403 with 38. These cosines use the retained 24-dimensional people space. The corresponding centered 60-dimensional comparisons are 0.7082, 0.6380, and 0.6687. Compression changes the numerical similarity. Follow the Andrew Ng selection, then Ng’s canonical profile, to inspect another person’s evidence.
Read an article. The Dean evidence includes a WIRED interview published on 14 December 2019. Its HN-derived record carries 15 December. Keep those dates distinct: the publisher and the archive record describe different observations. Then return to the map and ask which details in the source support the numerical association.
The related Eigen Hacks people view and Eigen Times people view extend the exploration into their separate corpora. Their axis numbers are local addresses, not universal concepts.
The audited current-news overlay has zero newly qualified article-person joins. It can display recent articles in the fixed news basis, but proximity to Dean’s historical profile does not establish that a recent article mentions him. This distinction is precisely where a promising research lead still needs evidence.
This query combines three tables: articles relevant to databases on a chosen date, sourced executive roles valid on that date, and qualified links between articles and people. The word “today” must resolve to an explicit date in all three.
Let be the queried date, a role’s known start date, and its known exclusive end date. Exclusive means the role is no longer valid on that endpoint. The symbol includes equality, whereas does not. Validity at the queried date can then be written . This is a useful interval convention, not permission to invent an end date. An unknown end does not by itself prove that a role continues today; retain the source’s as-of date and uncertainty.
The paper’s dated role query reaches executive records that have no supported profile in either published people model. The correct result is a visible gap, not a fabricated point on the map. Larry Ellison’s historical database coverage also cannot be substituted for a claim that he is the current CEO of a database company. A founder, a chair, and a current chief executive are distinct roles.
The research query can therefore show topic-ranked articles beside sourced roles while leaving missing joins visible. A completed current-CEO comparison would need additional profile coverage and article attribution. A larger identity catalog alone supplies neither.
Canonical identity answers which entity a record refers to. A Wikipedia/Wikidata correspondence can help confirm the node; it does not validate every article placed in a same-name bucket. Reviewed Devreal mappings reuse the existing canonical node. Name similarity alone never settles that correspondence.
A typed graph assertion answers another question: what relationship does a particular source state? “Employed by,” “reports to,” “coauthored,” and “competes with” have different meanings and sometimes different directions. Two people working for the same organization need not have overlapped or known each other. A co-mention records shared appearance in an article, leaving the relationship unspecified.
Dates and review states travel with the assertion. A source assertion, a machine classification, and a curator’s judgment are different kinds of support. A citation provides a route to inspection; it is not a measured probability of truth. Unknowns remain part of the record.
Suppose a synthetic article’s record says it names canonical person A. That establishes the recorded article-person association under its review status. A topic model gives the article a coordinate on a database-related axis. That establishes a measurement in a particular news basis. A cosine places A near B. That establishes similar coverage directions in a particular representation. None of these records states that A employs B.
To add an employment assertion, inspect a source that actually states that relationship, identify the subject and organization, and retain its dates and review status. To add a colleague assertion, evidence must support the relevant overlap or relationship rather than merely a shared organization name. One can preserve a candidate lead without promoting it to an asserted fact. This chain keeps identity, attribution, numerical similarity, source assertion, and curator judgment separate enough to correct independently.
Use a fictional announcement: “FinchDB raised $4 million; Harbor Ventures led and Cedar Capital participated.” There is one financing event, one recipient, and two explicitly named investor roles. The statement does not tell us how much either investor contributed. A partner quoted about the company is not automatically a personal investor or the deal’s lead partner.
One mathematical picture is an incidence table: rows are entities and columns are events. An entry records participation, with a separate role label such as recipient, lead, or participant. This explanatory picture lets several entities meet at one event without inventing every possible pairwise relationship. The deployed catalog stores event records and optional investor relationships; it need not materialize this matrix.
Now imagine a separate fictional regulatory notice reports $1 million sold, later amended to $3 million. These are successive cumulative observations of an offering, so adding them to claim $4 million would double-count. An offering target, proceeds sold to date, and an announced round total answer different questions. First-sale, filing, and announcement dates also remain distinct until evidence reconciles the events.
The audited catalog contains 2,893 financing events and 81 investor/partner relationship records. The totals need not agree: an event can have several known investors or none. Eight historical investor associations are excluded from the financing total. Of the financings, 2,875 come from selected SEC issuer-reported notices; publication is not SEC verification, and the selected notices are not a census of venture rounds.
Checkpoint. A financing has no named investor. Is its event record necessarily incomplete collection? Answer: No. The source may establish the financing while withholding investor identities. Preserve the event and the unknown participants.
Checkpoint. Two profiles have cosine 0.95 and share an employer. May we add “colleagues”? Answer: Those facts motivate a search. They do not establish overlapping employment or a sourced colleague relationship.
Use The Dual Geometry of People and News for the compact derivation, implementation audit, limitations, and source references behind this companion. Return to The Mathematics of Eigen Times for the foundations, especially text representation, covariance, SVD, named axes, and projection. Then explore Anthropology with three questions in view: which coordinates are being compared, what evidence supports the link, and which date does the claim describe?
All observation vectors in this companion are rows. A matrix direction, such as a column of , is a column. The shapes below summarize the definitions introduced in the chapters. The letter counts fitted people; counts reviewed training articles for the proposed calibration fit.
| Symbol | Shape | Meaning |
|---|---|---|
| Article embedding and its fitted reference mean | ||
| Retained raw news directions | ||
| Naming rotation and named-axis scale matrix | ||
| Raw, named, and standardized article coordinates | ||
| Cell background and person-cell mean | ||
| Balanced contrast and its unit direction | ||
| Unit person rows and their centered rows | ||
| Mean unit person profile in the fitted population | ||
| Equal-person covariance | ||
| Retained people-pattern loadings | ||
| A person’s retained people coordinates | ||
| Rank- projector in news-coordinate geometry | ||
| , | Article inputs and reviewed concept targets | |
| , | Proposed concept slopes and intercepts | |
| Algebraic slopes from people directions to concept scores |
Here , , and in the published models described by the paper; the main fictional example uses and . The bare is a retained dimension count; the indexed is a person’s contrast row. The hypothetical number of ontology targets, , is a modeling choice rather than a statement that every ontology concept has a trained classifier.
In the thin SVD , . The matrix is , is , and is . Selecting the first columns of with the fitted orientation produces ; the corresponding columns of and diagonal block of produce the score matrix . Coordinate sums and concept-score rows are different objects.
Time-window subscripts extend the same shapes: have coordinates, while has . The displacement , discarded squared length , attention share , and ranking score are scalars. In the alignment example, are anchor tables and is the orthogonal map; that use of is local to alignment.
Check 1. One unique background article and up to two person-article associations. The background describes a set of articles; the person profiles describe associations with each individual.
Check 2. The forward result is , the transpose is , and is . The returned row is . The combined matrix has rank 24, so it cannot be the 60-dimensional identity.
Check 3. No. Diagonal entries describe individual variances. In the example the off-diagonal entries remain -0.6 after marginal standardization, so the two coordinates remain correlated.
Check 4. Repeating E1 inside the background would weight that reporting more heavily merely because it has two associated people. E1 should nevertheless enter both personal aggregations because each asks about the articles associated with a different person. Deduplicate within a set, not across two different research questions.
Check 5. P1 ranks N3 second: its loading 0.5949 is smaller in magnitude than N1’s 0.7964. N3’s own row ranks P1 ahead of P2 because 0.5949 exceeds 0.3087. The shared entry agrees exactly; the competing entries differ.
Check 6. also projects to and returns as . The two input rows differ only along a discarded direction, so the retained map cannot distinguish them.
Check 7. No. A and B move from cosine 0.0392 to 0.8211, a change of about 0.782. Total variance is an aggregate reconstruction quantity, not a bound on every pairwise angle.
Check 8. Yes. Its comparison background can change because other associated articles enter a cell, or older articles can leave a rolling window. A support threshold can also change which cells remain in the average. These mechanisms do not necessarily describe a change in the person’s activities.
Check 9. No. A smaller topic coordinate can be outweighed by a larger attention share. Calling an article news about a person additionally needs a qualified article-person association and the underlying identity evidence. A ranking product cannot supply that link.
The three later checkpoints distinguish a linear score from a probability, a sourced financing event from its possibly unknown investors, and a coverage neighbor from a documented colleague. Their answers appear where each question is introduced so the required evidence remains next to the proposed inference.
The companion’s public source bundle includes the Markdown
manuscript, scripts/examples.py,
examples.json, the figure sources and rendered images,
Mermaid diagrams, packaging configuration, and a compact
paper-evidence.json snapshot of reported real-data
measurements. It also includes the supplied shelf artwork, which is kept
separate from the reading editions. No restricted news archive or full
catalog export is required to reproduce the fictional calculations.
From the Anthropology repository root, regenerate and check the numerical examples with:
cd papers/mathematics-of-anthropology
python3 scripts/examples.py
python3 scripts/examples.py --checkNumPy supplies the numerical linear algebra. Figure generation
retains editable SVGs and uses rsvg-convert for PNGs;
Mermaid source files are included with their rendered diagrams. The
README records the rendering and four-format build commands. The script
verifies unit norms, covariance and eigenvector identities, the
forward-return projection, residual orthogonality, cosine comparisons,
and the ridge solution. These are arithmetic and implementation checks,
not scientific validation of the real people model.
The reported Anthropology measurements retain their source dates and
model identifiers. The Hacker News people model is
eigenhacks:eigen:hn1:people:v1:bcaccd95b794; the
general-news model is
eigentimes:eigen:v2:people:v1:47debba3db48. The evidence
snapshot records the source paper’s checksum and audited implementation
revisions. The public source inventory records the companion inputs and
build revision. A refreshed news overlay must not silently replace the
dated evidence used in this edition.
The mathematical pipeline can be checked from synthetic inputs, while the empirical claims require their stated source artifacts and attribution rules. Keeping those two kinds of reproducibility distinct is part of making the model useful for research.
| Term | Meaning in this companion |
|---|---|
| Adjoint | Map transferring a linear operation across a dot product; the transpose for the real Euclidean coordinates here. |
| Anchor | Reviewed pair representing the same entity in two maps for an alignment. |
| Article association | Recorded candidate link between an article and a canonical person, with attribution evidence and review state. |
| Attention share | Allocated scalar story attention divided by a stated source-period total. |
| Basis | Independent directions used to express a space or subspace. |
| Calibration | Agreement between modeled probabilities and observed label frequencies for a defined task. |
| Cell | A specified source-and-decade grouping used for comparison and support. |
| Centering | Subtracting the selected reference mean from observations. |
| Contrast | A difference from a declared reference; a person contrast averages supported person-cell minus background rows. |
| Coordinate | Amount along a named direction in a specified coordinate system. |
| Correlation | Covariance divided by both coordinates’ positive standard deviations. |
| Covariance | Average product of paired centered coordinates, with an explicit denominator convention. |
| Crosswalk | Reviewed table connecting distinct vocabularies or representations while retaining their provenance. |
| Derivative | Limiting rate of scalar output change per input change; a partial derivative varies one input while holding others fixed. |
| Discrimination | Ability to distinguish positive and negative labeled cases, separate from calibration. |
| Eigenvalue and eigenvector | A stretch factor and direction satisfying covariance times direction equals factor times direction. |
| Embedding | Numerical representation of text produced by a specified model. |
| Frobenius norm | Square root of the sum of squared matrix entries. |
| Gradient | Collection of partial derivatives giving local changes of an objective with its parameters. |
| Held-out data | Observations excluded from a stated fitting or tuning stage and reserved to evaluate it. |
| Incidence table | Table marking which entities participate in which articles or events. |
| Intercept | Baseline linear score when all input coordinates are zero. |
| Logistic function | Map from a real score to a value between zero and one using the exponential function. |
| Mean | Coordinate sum divided by its observation count, or a stated weighted counterpart. |
| Norm | A length; the Euclidean norm here squares coordinates, adds them, and takes the square root. |
| Normalization | Dividing a nonzero row by its own length to retain direction. |
| Ontology | Reusable vocabulary of concepts and relationships among their meanings. |
| Orthogonal and orthonormal | Perpendicular; perpendicular with unit length. An orthogonal square matrix changes orientation while preserving lengths. |
| PCA | Principal component analysis: finding perpendicular directions of greatest variation in centered observations. |
| People pattern | Learned direction of variation across person profiles. |
| Person profile | A normalized, background-adjusted direction of one person’s associated news coverage. |
| Projection | Measuring and retaining components along chosen directions. |
| Procrustes alignment | Orthogonal fit aligning paired anchor rows by minimizing squared discrepancies. |
| Pseudoinverse | Generalized inverse giving a least-squares solution, selecting the shortest when multiple solutions exist. |
| Rank | Number of independent directions represented in a matrix. |
| Reconstruction | Approximation rebuilt from retained coordinates. |
| Residual | Original row minus reconstructed row, before any summary of its size. |
| Ridge penalty | Added squared-slope cost that trades prediction fit against coefficient magnitude. |
| Singular value | Nonnegative scale connecting left and right directions in an SVD. |
| Slope | Coefficient multiplying one input coordinate in a linear prediction. |
| Standard deviation | Nonnegative square root of variance. |
| Standardization | Dividing coordinates by their reference standard deviations; mean subtraction must also be specified. |
| Subspace | All weighted combinations of selected directions. |
| Support | Eligible observed articles satisfying the stated profile gates; not a guarantee of statistical reliability. |
| SVD | Singular value decomposition, factoring a table into left directions, nonnegative scales, and transposed right directions. |
| Trace | Sum of a square matrix’s diagonal entries. |
| Variance | Average squared centered coordinate value under a declared denominator convention. |
| Whitening | Transformation giving unit variances and zero cross-coordinate covariance in the chosen retained space. |
This index points to explanations and worked calculations rather than to isolated mentions. Notebook links use the same local anchors as the book.
Aggregation and counting: Unique backgrounds; equal supported cells; incremental counts and sums.
Alignment: Anchor fit and held-out row; constraints and rank.
Attention: Allocation; source denominators; topic ranking.
Calculus and fitting: Difference quotient, derivative, and ridge minimum; logistic gradient step.
Covariance and variance: Paired-deviation table; people covariance entry; retained variance.
Eigenvectors: Two-coordinate multiplication and maximum variance; people basis.
Evidence: Identity, association, and claim chain; dated Dean journey and event examples.
Geometry: Coordinate versus directional neighbors; compression effects.
Matrices: Product and transpose by hand; notation and shapes.
Normalization and standardization: Length and unit rows; A’s exact normalization; different denominators.
Ontology and probability: Tags, crosswalks, and proposed models; logistic and calibration; person-level substitution.
Projection and residual: First residual; exact round trip; SVD reconstruction.
Rotation and whitening: Rotated covariance arithmetic; article coordinates.
Scores and loadings: Both person dot products; one loading table, two entrances.
Time and version: Fixed ruler and separate clocks; background-only updates.