An eigenvector is a direction learned from variation. An ontology concept is a named idea that people agree to reuse. “Databases” can remain a useful concept while a database-related news direction changes between archives. Linking the two would let a reader follow one subject through Eigen Times, Eigen Hacks, and Anthropology without pretending that the sites share one coordinate system.
This chapter separates three stages: the personal tags already implemented in the 2 October 2026 paper edition; a proposed, reviewed table connecting concepts to axes; and a proposed learned model whose predictions would require evaluation.
The ontology picker lets a reader choose a concept and attach its identifier to a canonical node. Searching a label and clicking it records an annotation. It does not rotate the news basis, retrain a person vector, or establish a public factual relationship.
The pinned ontology has 157 concepts in 11 areas; 61 concepts have multiple parents. This is a polyhierarchy: one concept can belong under several broader ideas. For an invented illustration, “distributed databases” might belong under both “databases” and “distributed systems.” These parent links describe meaning. They need not form perpendicular directions, and a person can have several tags at once.
An axis label serves a different purpose. Its strongest stemmed terms
help a reader interpret one fitted direction. A list containing
databas, ceo, and github does not
define a clean databases category. Moreover, the negative pole needs
inspection too: it is not automatically “not databases.” The sign of an
eigenvector is a convention; a reviewed meaning must be attached to the
appropriate pole and model version.
For the mechanics behind named directions, return to From eigenvectors to named axes.
A crosswalk is a table of correspondences. Imagine a proposed row saying: in this particular Eigen Hacks basis, the positive pole of this axis has useful evidence for the databases concept. The row should retain the corpus, axis and sign, basis fingerprint, ontology version, representative articles, method, and review status.
The bridge should be many-to-many. Several axes may help retrieve database articles, and one mixed axis may connect to several concepts. A reviewer can leave a proposed correspondence unassigned. Shared concept identifiers then become destinations across the three sites, while each route retains its original numerical coordinates.
Concept-to-concept mappings also have different strengths: “exact,” “close,” “broader,” “narrower,” and “related” do different jobs in the W3C SKOS mapping vocabulary. An axis would first need a reviewed semantic interpretation before those conceptual distinctions could be used responsibly. A familiar-looking word is insufficient.
A later model could learn from reviewed articles. Let count the concepts reviewers label, and let be one article’s standardized news row. Let be a slope matrix: each slope measures how changing an input coordinate changes one concept’s linear score. Let be an intercept column supplying baseline scores at zero input. The notation is another form of the transpose symbol . Denote the resulting prediction row by . A simple linear predictor multiplies inputs by slopes and adds the baseline:
Each column of asks how the news coordinates contribute to one concept. Its predictions are scores; no probability interpretation has yet been established.
Consider an entirely illustrative two-coordinate, two-concept model:
The first prediction is . The second is . Nothing went wrong when these escaped the interval zero to one: they are scores from an unconstrained linear rule. We have not made a probability model.
To learn the coefficients, let count reviewed training articles, distinct from the fitted-person count used earlier. Stack their standardized rows into the matrix and their matching concept labels into the matrix . Row order agrees across the two matrices. For a binary label, an entry is one when reviewers judged the concept present and zero when they judged it absent. Missing judgments must remain missing, rather than becoming zeros.
A ridge fit balances squared prediction error against a penalty on large slopes. Let be the penalty strength. For a matrix, the Frobenius norm, denoted by double bars with subscript , is the square root of the sum of squared entries. Thus its square adds those squared entries. Here is a column of ones, so repeats the baseline row for every article. The following objective assumes every target entry is observed; partially reviewed targets require a separately specified objective summing only observed errors. Choose and to minimize
The first term adds squared prediction errors; the second penalizes large slopes. The intercept is deliberately unpenalized. A held-out validation procedure, using articles reserved from parameter fitting, would choose rather than a convenient-looking result on the training articles.
The toy coefficients above can actually be fitted. Take four fictional, fully reviewed article rows:
Each concept’s average label is 0.5, giving the intercept. Each coordinate column has sum of squares four and centered label-product sum two. With , its fitted slope is ; the cross-slopes are zero. Training predictions are 0.1 or 0.9. The new row then extrapolates beyond the training coordinates, producing . Training fit and behavior on new inputs are different tests.
Each corpus would need its own fitted coefficients. Shared target identifiers make the outputs comparable in meaning; they do not make the input bases interchangeable. Concepts overlap, so need not be orthogonal or even square. This semantic transformation does not inherit the reconstruction guarantees of .
A derivative describes how a scalar output changes per small change in a scalar input. Start with a function , where is a real scalar, and let be a nonzero change in that input. The difference quotient is
As approaches zero, this approaches . That limiting slope is the derivative, written . At , changes of 0.1 and 0.01 give difference quotients 4.1 and 4.01, approaching 4. A derivative is a local rate of change; it is not the function’s output, which also happens to equal 4 at this particular input.
Return to the first ridge target. Its centered labels are , and the first input coordinate is . Temporarily call its scalar slope and set the other slope to zero. Each of the four squared prediction errors is . With penalty , the scalar objective, named , is
Differentiate its expanded terms: . At an interior minimum a small move in either direction cannot improve the value, so the local slope must be zero. Solving gives . Complete the square to verify this is a minimum: . Its minimum value is 0.2. At the value is 1; at it is 0.25. Ridge accepts a little prediction error to reduce the coefficient penalty.
For any nonnegative penalty , this example instead has derivative , giving . The input columns have zero cross-product, so the two slopes separate cleanly. In a general table they interact. Let subtract each input column mean from , and let subtract each target column mean from . Let here be the identity. The stationary equations for these centered tables are . Solving this equation finds the slopes; restoring the target means minus the mean input’s fitted contribution gives the unpenalized intercepts.
The notation denotes a partial derivative: change coefficient while holding the other coefficients fixed. The gradient collects all those partial derivatives in the coefficient table’s shape. For the full squared-error objective its slope gradient is . Setting the whole gradient to zero yields the same stationary equations. The notebooks check the scalar derivative with small finite changes, the completed-square minimum, and the fitted matrix equations independently.
First define what is being predicted. One possible event is: “Under this annotation protocol, reviewers judge that this article substantially discusses databases.” It is not “this person is a database person.”
A logistic predictor maps a linear score through a smooth S-shaped function whose output lies between zero and one. Calibration means agreement between predicted probabilities and observed frequencies; the logistic bound alone does not establish it. The distinction is developed in Guo and colleagues’ study of calibration.
A reliability diagram offers a simple check. In an invented test set, collect 100 articles assigned probabilities near 0.8. Suppose reviewers mark 55 as positive. Plot their mean prediction near 0.8 horizontally and the observed fraction 0.55 vertically. A well-calibrated bin would lie near the diagonal, where those numbers agree. This example is hypothetical; Anthropology has not reported such a calibration experiment.
Repeat across probability ranges, report the number of articles in each bin and uncertainty, and inspect relevant sources and periods separately. Small bins can fluctuate substantially. Prevalence is the fraction of articles with a positive target. A model that predicts this overall fraction for every article can be calibrated yet poor at ranking individual articles. Discrimination is its ability to distinguish positive from negative cases. Calibration and discrimination therefore need separate evaluation, with articles reserved for testing rather than reused for fitting or tuning.
Checkpoint. The toy linear model predicts 1.3. Can the interface print “130% confidence”? Answer: No. It can display a labeled score. A probability needs a defined target, a suitable model, and held-out evidence that its estimates behave like probabilities.
Let denote a real-valued linear score and let denote a modeled probability for one defined binary article label. Write for the exponential function, with base ; its output is positive for every finite real input. The logistic function, written in this section only, is
At zero, , so . At two, , so . At minus two, . The denominator always exceeds one, placing the result strictly between zero and one for finite scores. This function’s is not the earlier standard deviation .
To fit such outputs, the notebook uses the first column of the four-row target table. Let be a two-entry slope column and a scalar intercept, giving score for training article index . Let be its reviewed zero-or-one target. Binary log loss for one row is , where is the natural logarithm, the inverse of the exponential. A correct high-probability prediction gets a small loss; a confidently wrong one gets a large loss. All four losses are averaged. The demonstration adds with penalty , keeping the intercept unpenalized.
We need three derivative rules before differentiating this model. The derivative of is ; the derivative of is for positive ; and the derivative of the reciprocal with respect to nonzero is . The chain rule multiplies the local rates when one function is placed inside another. We use these rules here and verify the resulting derivatives by small finite changes in the notebooks.
To differentiate the logistic function, name its denominator . The inner negative sign has derivative -1, so . Applying the reciprocal rule and the chain rule gives
One factor is ; the other is . Hence the derivative is . At , its value is .
Now let name a single row’s log loss and let be its fixed binary target. The notation writes its derivative with respect to . Differentiating with respect to its probability gives
Composing this loss with the logistic score multiplies by , cancelling the denominator and leaving . A slope changes the score at a rate equal to its input coordinate; the intercept changes it at rate one. Thus the slope gradient is the average of plus ; the intercept gradient is the average of .
Start all slopes and the intercept at zero. Every prediction is 0.5, so the four errors are . Multiplying those by the first input column and averaging gives . The second slope gradient and intercept gradient are both zero. A gradient-descent step subtracts the gradient times a positive step size; with step size 0.2 the new slopes are and the intercept stays zero. Positive first-coordinate articles now receive .
The native notebooks repeat that update 2,000 times and check a small final gradient and a separate reference slope. This verifies the training calculation for declared synthetic inputs. It supplies no real-world held-out calibration result. In the separate hypothetical reliability bin, observed frequency differs from 0.8 predicted frequency by 0.25. Constraining outputs to a probability range and testing their agreement with observations remain distinct tasks.
A person profile has been background-adjusted and normalized. It describes the direction of distinctive coverage. An article row has undergone neither of those person-level operations. Their equal number of coordinates does not make them the same kind of observation.
For a person score row , the earlier chapters reconstructed an approximate profile as . Use the people basis and article predictor from the same corpus and news-coordinate version. Let denote the score row obtained by applying the article predictor to this reconstructed profile. Substitution gives
The product has shape : it describes how moving along each of the people directions would change those linear scores. This is a valid algebraic identity. Its interpretation as a useful prediction about a person remains an untested use outside the article model’s training distribution. Neither restoring the mean nor adding the intercept solves that problem.
Another possible quantity is the fraction of a person’s associated articles predicted to discuss a concept. That answers a coverage question, with uncertainty from both article attribution and classification. For a nonlinear predictor, predicting an average row and averaging individual predictions can differ. A future system must specify which quantity it means and evaluate it on independently reviewed person-level examples.
Substitute the reconstructed row into the article rule directly: . Distribute multiplication over addition to obtain , then group the last matrix product as . Grouping changes where we calculate intermediate tables, without changing their compatible dimensions or result. The notebook uses a separately declared three-input, two-output synthetic slope table because the earlier two-input ridge table does not fit the three-coordinate people fixture.
For a nonlinear example, take scalar article scores zero and two. Their average is one, so predicting after averaging gives . Predicting separately and then averaging gives . The difference is approximately 0.040660. Thus even the arithmetic choice of when to aggregate changes the output. A defined person-level target and separately reviewed evaluation remain necessary whichever quantity is chosen.
Suppose two corpora contain reviewed representations of the same people. An anchor is a pair of rows known to describe the same entity in the two corpora. An orthogonal Procrustes fit tries to rotate or reflect one set of such rows toward the other. For this alignment only, let count paired anchors and count coordinates in each representation. Let and be their centered tables, with matching entities in the same row; here is an anchor table, not the preceding concept-target matrix. Let be a alignment matrix, distinct from the SVD factor . The identity now has shape . The fit minimizes subject to the orthogonality constraint .
The constraint preserves distances within the chosen coordinate geometry. It cannot correct every difference in populations, source coverage, or scaling. Shared names are not enough; the anchors must be correctly identified and substantively comparable.
The published models share 60 canonical people. Sixty centered rows in 60 dimensions have rank at most 59: subtracting their mean makes the rows sum to zero, creating a dependence. Consequently, they cannot determine a full 60-dimensional alignment from independent evidence in every direction. Reserving anchors for testing leaves still fewer for fitting. A lower-dimensional or regularized approach is possible, but its assumptions and performance need evaluation. See Matching axes for the prerequisite distinction between comparing directions and assigning correspondences.
Use four centered synthetic anchor rows , , , and for . Stipulate a known map only to construct test data . It maps a row to . The first pair is therefore and . The two tables have zero column means.
For the fit, multiply paired columns to obtain . Denote the SVD of this product by , where and are two-by-two orthogonal direction matrices and contains its nonnegative singular values. The Procrustes solution is . The notebook computes it from the paired tables, then checks that it equals the stipulated map and gives zero training discrepancy.
Why this choice? Expanding the squared discrepancy yields the squared lengths of and minus twice the trace of . The first two quantities are fixed under an orthogonal map. In the singular coordinate frame, the trace is largest when each nonnegative singular value is multiplied by one rather than a smaller diagonal coefficient of an orthogonal matrix. The choice achieves that maximum, and therefore the minimum discrepancy. This permits reflections as well as rotations, exactly as the stated constraint allows.
Reserve a fifth row from fitting. Its stipulated target is ; the estimated map reproduces that target with zero error. This noiseless test proves the toy algebra, not a real alignment’s usefulness. Actual held-out anchors could disagree because coverage, scale, or meaning differs between corpora.
The rank shortage has an equally small example. Center three two-coordinate rows: their sum becomes , so the third row is the negative sum of the first two. There are at most two independent centered rows. With sixty centered anchors, the same dependence leaves at most fifty-nine independent rows. The notebooks also compute this bound using a synthetic centered sixty-by-sixty identity table; it is not a measurement of the real anchor matrix.