Abstract
A person can be treated as a source of news in a specific computational sense: the dated articles associated with that person form a distribution over recurring news patterns. Anthropology connects these distributions to sourced relationships, while Eigen Times and Eigen Hacks provide distinct news corpora and fixed topic spaces. This revision documents a 59,399-node catalog with 102,261 relationship records and 2,893 source-reported financing events. Relationship records distinguish 87,684 source assertions, 27 machine classifications, and 14,550 unclassified co-mentions. The preceding identity pass attempted all 106,907 NYT candidate name keys; 65,686 remain unresolved. Catalog confirmation does not verify every article attribution or expand the separately fitted people models. Their qualified articles are projected into 60-dimensional news spaces, standardized, balanced by supported source-and-decade cells, and summarized as unit person profiles. A second principal-component decomposition retains 24 people directions. One loading matrix supports news-to-people and people-to-news lookup; its transpose reconstructs a projection, not the articles or identities lost during aggregation. The unchanged models contain 217 Hacker News profiles and 61 general-news profiles, retaining 84.54% and 90.86% of between-person variance. We report descriptive navigation examples, resumable reconciliation, selected news extraction, issuer-reported SEC offerings, and immutable publication through Rust/PostgreSQL, Vercel, and S3. Implemented varimax naming and personal ontology annotations are distinguished from a proposed cross-site calibration layer. The result is an inspectable research system, not a validated measure of influence, expertise, or social relationship.
Keywords: computational social science; news archives; principal components; entity provenance; temporal retrieval; ontology alignment.
1 Introduction
The history of technology is carried by overlapping institutions and careers. Researchers become founders; early employees become investors; companies are acquired while their alumni establish new organizations. News records these trajectories unevenly, mixing technical work, public controversy, financial reporting, retrospective history, and incidental mentions. A useful research interface must support both kinds of question: Who was connected to whom? and What patterns of reporting surround this person, and how do those patterns appear in current news?
Anthropology, at anthropolo.gy, addresses the first question with documented relationships. Eigen Times and Eigen Hacks address the second through separate decompositions of general news and Hacker News. Their integration makes it possible to start with a person, inspect associated news directions, follow those directions to other people, and return to dated articles. The reverse journey begins with an article or topic and reaches people whose historical coverage occupies related directions. This is a retrieval system for generating and checking research questions. The evidence supporting a graph edge remains distinct from the numerical evidence supporting a coverage similarity.
The present paper extends the earlier Eigen Times systems account [1] by describing the implemented people layer and its three interfaces. Its contributions are a precise two-stage spectral construction, a distinction between navigational reciprocity and mathematical inversion, reproducible examples with support counts and comparison baselines, and a specification for ontology calibration that preserves provenance and uncertainty. This revision adds a completed identity reconciliation, organization assertions, selected news relationships, and first-class financing events. These expansions share a recoverable publication architecture while retaining different evidence statuses. Implementation statements refer to the audited source and artifacts in the appendix. Proposed methods are identified explicitly.
The mathematical lineage includes latent semantic retrieval [2] and principal-component analysis [3]. Temporal topic models [4] and diachronic representation studies [5] motivate care about changing language and coordinate systems. Temporal knowledge bases [6] motivate treating dated roles separately from timeless labels. None of these precedents supplies an accuracy estimate for the present system; that requires an independent evaluation.
2 Three interfaces and three evidence layers
| Interface | Primary object | Research entry point |
|---|---|---|
| Anthropology | Sourced person and organization records | Relationships, tribes, career waves, and canonical identities |
| Eigen Times | General-news corpus and its own news basis | Historical or current reporting, then people profiles |
| Eigen Hacks | Hacker News corpus and its own news basis | Technology discourse, then people profiles |
The curated seed graph contains 176 people and 108 organizations, connected by 380 sourced relationships; these seed counts are distinct from the expanded catalog below. Its 23 tribes are editorial research groupings; its 14 waves describe dated transitions. The graph cites 114 research sources and 28 canonical Devreal records. A relationship records its type, endpoints, evidence, and dates when supported. Coauthorship, postdoctoral supervision, employment, founding, investment, and acquisition are separate assertions. A group label such as “Sun Microsystems alumni” does not assert that all members met, agreed, or worked together simultaneously. Reviewed Devreal identities are reused rather than duplicated. [7]
The archive layer contains 21,264,134 stored article rows across the locally held corpora. These are rows, not unique people or necessarily unique articles. By 2 October 2026, the bounded identity pass had attempted all 106,907 NYT name-shaped keyword values: 41,221 had confirmed Wikidata/Wikipedia identities, 65,686 remained unresolved, and none remained pending. Several source keys can identify one person, while a keyword bucket can contain several namesakes. The identity edition published at 07:03 UTC contained 39,121 records, including 39,091 distinct Wikidata identities and 30 curated records without that confirmation. Subsequent organization and financing enrichment expanded the catalog to 59,399 nodes at 18:56 UTC, without changing the number of people. Identity confirmation neither verifies every associated article nor creates a documented relationship or a supported people-vector profile.
The third layer comprises personal organization tools. Free users can save node lists and select ontology tags; launch Pro adds personal records and identity-review requests. An approved owner’s statement is separate from historical evidence. An account annotation does not become a public consensus label, an independently verified relationship, or training data merely because it exists.
The homepage deliberately displays at most 60 nodes. Search and tribe filters provide other bounded views; the node list reaches the larger catalog and the path finder follows documented graph edges. Catalog growth does not automatically create new edges. Clickable count labels distinguish catalog coverage, relationship evidence, reported financings, and tribes. Unclassified co-mentions are excluded from the sourced-connection headline; machine classifications retain their review status. Figure 1 shows a live filtered view. Its force-directed positions display documented edges, not coordinates from the people-vector model.
3 Catalog reconciliation and recoverable publication
3.1 Decisions, identities, and article attribution
A completed matching pass means every queued key received a decision, not that every key became a known person. The NYT keyword pass supplies 106,907 of the 107,542 candidate rows; 284 curated graph rows and 351 archive rows supply the rest. Aliases are reconciled into canonical records, preserving Wikidata identifiers and Wikipedia links when confirmed. At the 07:03 UTC identity edition, the wider catalog included 37 reviewed Devreal mappings, compared with 28 in the smaller curated graph. Four explicit Devreal namesake nonmatches prevented exact-name collisions from becoming identity merges. Later enrichment increased reviewed mappings to 74 while preserving these scoped exclusions.
| Measurement | 07:03 UTC identity edition |
|---|---|
| Catalog people / organizations | 39,013 / 108 |
| Distinct Wikidata identities / lookup rows | 39,091 / 39,595 |
| Curated nodes / sourced edges / tribes | 284 / 380 / 23 |
| NYT keys attempted / pending | 106,907 / 0 |
| NYT keys confirmed / unresolved | 41,221 / 65,686 |
| Selected context-retry keys | 500 |
| Additional confirmed keys / distinct identities | 67 / 43 |
| Scoped source-context reviews / exclusions | 29 / 10 |
| Reviewed Devreal mappings / namesake nonmatches | 37 / 4 |
The first pass confirmed 41,154 NYT keys. A selected 500-key context retry added 67 confirmations and 43 distinct identities; its other outcomes were 273 no-matches, 157 needing context, one ambiguous case, and two outside the target scope. The retry requires a specific canonical-label anchor in an individual dated headline, a compatible structured life-date qualifier, or an explicit scoped review, in addition to authoritative identity/type and reciprocal Wikipedia-QID checks. This source-witness requirement applies to the retry, not retrospectively to every first-pass confirmation. Its 13.4% yield measures recovery within a selected queue, not precision or recall.
Review excluded three wrong-person retry associations before publication. A headline about one namesake cannot lend its date or organization to a different headline in the same bucket. Broad place or institution aliases are insufficient identity witnesses. The identity-stage registry contains 29 scoped source-context reviews: 28 mixed-person or unresolved-mention buckets and one discrepancy in a source birth fact. Ten newly reviewed mixed buckets separate supporting references from conflicting or unresolved references. The ten exact source-key/QID exclusions across the registry have narrower scope than a ban on an entire person.
For example, the keyword “Carr, David” supports the journalist through obituary and book references while also containing quarterback coverage. Confirming the journalist does not confirm the whole bucket. Similarly, Peter Norvig’s Wikidata alias resolves to Devreal’s canonical node; the filmmaker Mike White remains separate from Devreal’s Hootsuite engineer of the same name. These distinctions preserve both positive evidence and uncertainty. The 217/61 people-vector profiles predate this catalog pass and were neither retrained nor comprehensively re-attributed by it.
3.2 Recovery and bounded collection
After a workstation interruption, all 2,350 committed journal batches passed integrity, base, and corpus checks. The recovered state had 98,560 attempted keys and 8,347 pending. A pending-only continuation completed those keys in 643 seconds with 584 API requests and no request errors. This is one observed recovery run, not a throughput benchmark. The final reviewed manifest records journal sequence 2,586. The last pre-interruption worker error was operating-system file-table exhaustion (ENFILE); the record does not establish the cause of the workstation failure.
Collection is separate from web requests. Workers persist batches of at most 40 candidate decisions and cache API metadata for 24 hours. Uncached metadata API requests are serialized with a default 1.1-second start-to-start interval. Source robot policies, identification, maxlag, and retry/backoff controls accompany the collector. Wikimedia’s data-access guidance and API etiquette inform this bounded metadata workflow [8], [9]. The robots protocol describes crawler preferences [10]; API quotas and access controls remain separate mechanisms. This identity pass reconciled held archive metadata; subsequent collectors added explicit source assertions and bounded SEC records, as described below. No recurring crawl or collector is enabled. A future schedule can refresh authoritative identities and dated roles more often than historical catalog discovery, publishing only validated revisions.
3.3 Organization and news assertions
The first organization pass reads explicit, non-deprecated Wikidata employment, education, membership, and affiliation statements. Organization classes are checked through bounded ancestry; newly confirmed organizations require a Wikipedia page that reciprocally identifies the same Wikidata entity. Founder, chief-executive, and chairperson statements link only to people already in the canonical catalog. The pass added 17,735 organizations and 87,248 statement records, reaching 30,279 people. It preserves statement identifiers, revisions, references, qualifiers, and retrieval dates. Wikidata assertions are not independent editorial verification; some lack an external reference. Unambiguous year qualifiers provide role dates, while unknown dates remain unknown. A common employer does not establish colleague or reporting relationships. [7]
The news pass reads 8,188 selected local articles: 5,325 from Eigen Hacks and 2,863 from Eigen Times. Selection starts from earlier qualified mention indexes containing roughly 370 names, retaining articles with at least two canonical candidates. Within these articles, a full-name dictionary searches all confirmed catalog people and requires nearby affiliation or professional context. This expands discovery inside the selected set; it is not a rescan of the 21-million-row archive. The pass retained 16,623 qualified person–article associations, including 4,335 additional canonical associations. These qualifications remain machine decisions that can be wrong.
Typed relations require explicit local wording joining both people in one sentence. A unique surname can refer back to an established full name, but first names and pronouns are not resolved. The taxonomy distinguishes colleagues, subordinate-to-manager reporting, personal competition, analyst assessment, journalistic reporting, and explicit comments. Company rivalry does not establish personal rivalry; reporting information to someone does not establish managerial reporting. Negated, contested, hypothetical, uncertain, and quoted factual relationships are withheld. Analyst assessments preserve their attribution rather than asserting their contents as fact. Announced appointments retain their future or announced status; article dates are observations, never manufactured role intervals.
The resulting 14,577 news relationship records comprise 27 typed, machine-classified records and 14,550 unclassified co-mention pairs. The classifier produced 153 explicit sentence findings, withheld 43, and retained 110 findings across 96 articles; findings and deduplicated records are different units. An edge retains at most five representative source references; source counts use unique article IDs rather than repeated mentions. Public excerpts are complete sentences of at most 25 words, once per source URL. Full text and detailed candidate audits remain local. Machine classification denotes review status, not calibrated probability, human verification, or a factual relationship established by co-occurrence.
3.4 Financing events without invented investors
The financing layer treats the recipient’s reported financing as a first-class event, with optional linked investor relationships. An event can exist when no investor is named. Explicit investor capacities distinguish lead, co-lead, participant, unspecified, and investment partner. A partner acting for a named firm is not assumed to invest personally. Neither a title, a press-release quotation, nor board membership establishes that a person led a deal. Round totals, individual contributions, cumulative investment, commitments, and valuations retain separate scopes; repeated copies of a round total are never summed. Unknown stages and instruments remain unspecified. These records do not describe a complete cap table.
A separate discovery scan found 2,832 funding-language sentence candidates in 1,165 of the same 8,188 selected article records. All remained outside the published factual graph. Names and amount strings generate review leads, not confirmed parties or financings; uncertainty, fundraising targets, and contested claims retain separate statuses. Reviewed announcements and historical associations are curated separately, with event-level reconciliation of repeated reporting.
The SEC collector uses the Commission’s quarterly Form D datasets [11]. These preserve issuer-reported exempt-offering notices; neither filing nor inclusion in this catalog constitutes SEC verification. The bounded run inspected 60,990 filings from 2025 Q3 through 2026 Q2. It retained 2,876 offerings across 2,527 issuers and published 2,875 after withholding one ambiguous issuer identity. Eligible initial notices report an actual first sale and positive proceeds for a technology corporation. The selection excludes investment funds, nonoperating vehicles, business combinations, and incomplete notices. It is a selected regulatory sample, not a census of startup fundraising.
Explicit amendment chains update the same offering. The retained collection includes 173 amendments within 154 events; they contribute neither duplicate events nor additional copies of proceeds. Ambiguous initial notices, unanchored amendments, and corrected-date collisions are withheld. Reported proceeds mean securities sold as of the selected filing; the offering target is separate. A first-sale date is neither an announcement date nor proof of completion. Investor identities, lead roles, and stages are not inferred from SEC filings. Announcement and regulatory records require explicit event reconciliation across their different dates and amount scopes.
Issuer identity uses exact CIK mappings or documented contextual review. A similar name creates a review item, not a merge. Registry-only issuers remain distinct from curated and Wikipedia-confirmed organizations; a Wikidata CIK match without a reciprocal Wikipedia page does not count as Wikipedia confirmation. A legal issuer’s identity also does not establish its relationship to a similarly named product. The final edition retains 2,514 registry organizations and 74 reviewed Devreal mappings. It withholds the unresolved issuer overlap and reports no pending Devreal overlaps.
| Measurement | 18:56 UTC enriched edition |
|---|---|
| Catalog nodes: people / organizations | 59,399: 39,013 / 20,386 |
| Wikipedia-confirmed people / organizations | 38,996 / 17,853 |
| Registry organizations | 2,514 |
| Relationship records, all statuses | 102,261 |
| Source-asserted / machine-classified / co-mentions | 87,684 / 27 / 14,550 |
| Curated / Wikidata / news relationship origins | 436 / 87,248 / 14,577 |
| Financing events / historical investor associations | 2,893 / 8 |
| SEC / curated financing events | 2,875 / 18 |
| Explicit investor relationship records | 81 |
The merger pins its input revision, preserves unrelated edges and prior evidence, validates canonical endpoints and event recipients, and archives the exact bytes before atomic database publication. The final snapshot is 270,937,620 bytes; its revision and file hash agree across merge, database, and live receipts. The funding directory defaults to 2,893 reported financing events, while the combined event/association table has 2,901 records. Catalog enrichment leaves the curated seed graph and fitted people-model populations distinct. No recurring collection schedule was activated.
3.5 Deployed storage and access boundaries
Next.js on Vercel queries a separate Rust/Axum catalog service, also on Vercel, backed by Neon PostgreSQL in us-west-2. Catalog pages return at most 100 records; the service enforces 60 requests per fixed minute bucket per hashed client identity, and archive search allows 30. Versioned catalog tables and an atomic head switch let readers pin one immutable revision during publication. The restricted runtime can read the public catalog and update bounded rate-limit state; operator imports hold separate publication privileges. [7]
Private, versioned S3 storage preserves immutable catalog snapshots and selected article-reference metadata. Frontend evidence reads use a restricted production Vercel OIDC role rather than persistent AWS access keys [12]. Stored references contain titles, dates, URLs, and article identifiers; these exports do not distribute article bodies or account data. Bundled evidence provides a fallback when storage is unavailable. Account claims, lists, tags, and Pro records remain behind Verdun permissions in separate tables. S3 stores artifacts and evidence; PostgreSQL serves relational queries. No comparative cost or latency result is claimed.
4 From articles to named news coordinates
We use real-valued row vectors: denotes the real numbers and a superscript denotes matrix transpose. For any vector , denotes its Euclidean norm. Matrix dimensions specify rows followed by columns. For one corpus and one fixed version, let index an article, be the embedding dimension, and the number of retained news axes. Let be article ’s embedding, the fitted mean embedding, and the matrix whose columns are retained orthonormal eigenvectors. The audited pipeline uses a quantized BGE small English v1.5 checkpoint with , representing a cleaned title and a limited text prefix; the retained news dimension is . [13], [14], [15]
Let index raw news axes and index named news axes, each from 1 to . Let be the positive retained raw eigenvalues, their diagonal matrix, and the stored orthogonal naming rotation; is its row-, column- entry. Let be the positive marginal standard deviation of named axis under the fitted covariance, and let be the diagonal matrix with these standard deviations on its diagonal. Its off-diagonal entries are zero; the superscript denotes an inverse.
For article , denote its raw news-coordinate row by , its rotated named-coordinate row by , and its standardized named-coordinate row by , all in . The last row measures each named coordinate in marginal standard-deviation units. These projections and scales are computed as follows:
(1)
(2)
The means, eigenvectors, rotation, eigenvalues, and scaling matrix belong to a fingerprinted corpus version. Equal axis numbers in two corpora do not denote equal directions.
4.1 Naming is an unsupervised orientation
For the naming fit, let be the number of aligned article rows and the number of lexical features. Let contain their term-frequency–inverse-document-frequency (TF-IDF) rows, their mean row, and their raw coordinate rows . Throughout, denotes an all-ones column whose length matches the rows of the matrix being centered; here it has length . Let be the standardized raw score matrix, computed as ; the diagonal entries of are the inverse square roots of the corresponding eigenvalues. Finally, let denote lexical loadings on the raw axes and their rotated counterparts:
(3)
A pairwise varimax optimization makes the columns of more concentrated [16]. Printed labels are high positive-loading stemmed terms. They are not a supervised taxonomy, and they may describe one pole better than the other. The implementation fits this rotation on standardized raw scores and applies the stored rotation to unstandardized raw news coordinates before the marginal scaling in Equation 1. Because scaling and rotation do not generally commute, is a naming heuristic, not a literal term attribution for every standardized coordinate .
When article subscripts are omitted, , , and denote the corresponding raw, named, and standardized coordinate rows over the fitted article population; denotes their feature covariance matrix. Orthogonality preserves retained Euclidean energy and reconstruction between and : . It does not make the rotated covariance diagonal. Under the fitted covariance, , and
(4)
The diagonal of Equation 4 is one, but its off-diagonal entries generally remain nonzero. The operation is marginal standardization, not full whitening. Consequently, the sum of squared standardized named coordinates is not generally the original Mahalanobis statistic. This distinction matters because the people model learns geometry in these standardized news coordinates.
5 People as distributions of associated reporting
A “person as a news source” is an analytical aggregation, not an assertion that the person authored, endorsed, or caused the articles. In the frozen vector-input pipeline, exact-name matches pass conservative temporal and contextual gates. These article-person associations remain plausible, unreviewed candidates, independently of the later catalog reconciliation. The temporal gate uses the earliest documented professional context, not a person’s birth year; it can therefore exclude genuine early reporting. Retrospectives remain eligible. Publication time is not event time.
Let index a person, denote a coverage window, and a source-and-decade cell. Let be the set of unique qualified candidate articles in that cell and window, and its subset associated with person . Bars around a set denote its cardinality (number of articles); bars around a scalar later denote absolute value. In the averages below, runs over the indicated article set, and is article ’s standardized named-news coordinate row from Equation 1. Thus it is an article vector, not a person score or a scalar normalization constant. Let denote the cell’s background mean and the person’s cell mean. For nonempty article sets, define
(5)
The background mean is computed from candidate-associated articles, not the entire news corpus. An article mentioning two people is counted once in the background and once in each person’s relevant aggregation. Let be the set of cells with at least five articles associated with person in window ; its cardinality is the number of retained cells. Let denote the average background-subtracted residual across those cells, and its unit-normalized profile. A person-window profile also needs at least ten articles across retained cells, a nonempty retained-cell set, and a nonzero residual:
(6)
This gives equal weight to supported cells. It does not give every source equal total weight when sources span different numbers of decades. Background subtraction reduces one form of compositional dominance but does not eliminate archive selection, source, or historical bias. Unit normalization makes the direction comparable without using coverage volume as the profile’s magnitude; support counts must therefore remain visible separately.
Let be the number of supported people in the all-history training window. Their unit profiles form the rows of . Let denote their mean row and their centered profile matrix, computed as ; here has length . Let denote the population-normalized people covariance in news-coordinate space. Let be the number of retained positive people eigen-directions, their column matrix, the diagonal matrix of corresponding eigenvalues, and the identity matrix. These objects satisfy
(7)
The implementation retains at most 24 positive eigen-directions ( in the audited models) and gives each training person equal weight. Let index these directions from 1 to , and let denote column of , used as a column vector in matrix products. The implementation orients each direction so its largest-magnitude news loading is positive. This resolves a sign convention, not the ambiguity of nearly degenerate eigenspaces. There is no second varimax step and no whitening of people scores. The columns of are orthogonal in standardized news-coordinate geometry; their corresponding embedding-space directions generally are not mutually orthogonal.
Let denote the people-coordinate row for person in window . Every supported window uses the same all-history mean and loading matrix:
(8)
All-history, decade, recent-36-month, and latest-90-day profiles therefore share a fixed people basis. Their candidate backgrounds are nevertheless window-specific, so changes in a profile can reflect both its reporting and its comparison background. An absent supported window is missing data, not the zero vector.
6 The navigation duality
The phrase “people vector” has two related meanings. A row locates an individual in the retained people space. A column of describes a population pattern as a weighted combination of news directions. The individuals with large positive or negative scores on that column characterize its two poles. A people pattern is not a single person or a graph community.
6.1 One coefficient, two indexes
Write for the scalar entry in row , column of . Each coefficient appears twice in the public artifact: as the news loading of people pattern , and as the people-pattern loading associated with news axis . Selecting a news axis ranks patterns by across . Selecting a people pattern ranks news axes by the same quantity across . The interfaces expose a row and a column of one matrix. Their coefficients are identical, but their rankings compare different sets; they are not conditional probabilities or reciprocal rank positions. [7], [15]
For coordinate subscripts following a comma, is news coordinate of , and is people coordinate of . The person list has two distinct modes. News-first exploration ranks the direct unit-profile coordinate . People-first exploration ranks . Nearest-neighbor similarity is a third statistic: the cosine between complete retained people-coordinate rows, within the same model and window. For two nonzero rows, cosine is their dot product divided by the product of their Euclidean norms; it is undefined if either row is zero. These quantities answer different questions and should not be presented as interchangeable rankings.
6.2 Projection, adjoint, and information loss
For one supported person-window pair, abbreviate its unit profile as and its people-coordinate row as . Let denote the centered profile, computed as , and let denote its reconstruction from the retained people coordinates. The forward map is . The return map multiplies by the transposed loading matrix:
(9)
Let denote the forward-and-return matrix .
Proposition. If has orthonormal columns, then is an orthogonal projector of rank . Thus the forward-and-return operation preserves only the component of in the retained subspace.
A small example makes the distinction concrete. In this toy example only, set and . Suppose the three news directions describe databases, enterprise applications, and research, and the two-pattern loading matrix is
(10)
A database query represented as the centered direction projects to and returns as . Navigation reaches a meaningful shared enterprise pattern, while reconstruction loses the database-versus-application distinction. The example is illustrative, not a fit to the deployed data. Outside this example, and retain their fitted-model meanings. Even if all 60 components were retained, one could not invert name selection, averaging, background subtraction, normalization, or the earlier 384-to-60 article projection.
The singular-value decomposition also explains why these are population patterns. For the full decomposition of , let and contain orthonormal left and right singular vectors, respectively, and let be the rectangular diagonal matrix of nonnegative singular values in descending order. Then . Let and contain the first columns of the corresponding matrices, and the leading singular-value block. Taking the same ordering and orientation as the people eigendecomposition gives and retained person scores . The columns of diagonalize in news-coordinate space, while the corresponding columns of diagonalize in person space. This is the spectral duality behind the interface. It is not an eigendecomposition of Anthropology’s factual relationship adjacency matrix. The symbol here denotes singular values; above continues to denote standardized raw news scores.
7 Time, attention, and the manifestation in current news
The system separates three clocks: article publication time, the historical profile watermark, and the as-of date of a documented role. In the frozen vector snapshot, person-profile matches end on 4 September 2026, while the local corpus projection watermark is 5 September. Its article overlay is dated 1 October and contains 713 Hacker News records and 53 general-news records. Neither the 2 October identity publication nor the later relationship and financing releases refreshed this overlay. Refreshing an overlay neither refits the people basis nor backfills its name dictionary.
Eigen significance is retained as an additional signal, distinct from topic affinity. Original story rank combines story energy with member novelty. The score is allocated among canonical member articles and normalized by the scored source’s total for the relevant period. Person attention histories average coverage shares over active sources, including sources with no mention. A multi-person article can contribute coverage to several people, so these person shares do not partition a fixed total of human importance.
For the latest-article display in Anthropology, let denote coordinate of the standardized article row , and let be article ’s share of original Eigen attention in the edition. Let denote the resulting attention-weighted topic score used to order articles, rather than an ordinal rank position. A selectable ranking uses
(11)
Topic-only ordering and pole filters are also available. The static Eigen people pages use topic strength. Attention is an overlay for exploration, not a weight secretly inserted into the equal-person PCA fit.
A version boundary is especially important for Hacker News. Its people model remains in hn1; current article embeddings are reprojected into that unchanged space. Current scalar attention is drawn from hn2. This combination is explicit and permissible because the scalar does not replace an axis identity. The latest overlay currently has no newly qualified person-article join. An article near a person’s historical direction is therefore a candidate for investigation, not evidence that the article concerns that person.
7.1 Updating people while holding patterns fixed
A natural extension is to refresh supported person profiles more often than the population basis. For each qualified person-cell pair, maintain an article count and a coordinate sum; retain a separate deduplicated ledger for each cell background. New articles, corrected identities, and articles leaving a rolling window update these sufficient statistics. Recompute affected residuals and unit profiles, then apply the unchanged mean and . A background change can affect every person using that cell, so an implementation must invalidate more than the newly mentioned person’s profile. This incremental refresh is proposed; the current latest-news overlay does not perform it.
Within one version, let and be two supported coverage windows for person , and let denote the Euclidean distance between that person’s people-coordinate rows.
The change statistic is .
For a centered profile in either window, a complementary statistic is the discarded energy , which identifies reporting variation poorly represented by existing people directions. Neither statistic establishes a career change: both need support counts, underlying articles, and sensitivity to window-specific backgrounds. Source-stratified resampling and time holdouts could distinguish stable changes from sparse-cell variation. Sustained reconstruction deterioration would motivate a separately evaluated basis revision. Old and new artifacts should remain addressable, with explicit alignment evidence instead of silently moving every person’s coordinates.
8 Ontology calibration across the three sites
An ontology specifies reusable concepts and relations [17]; a vector basis specifies coordinates. Orthogonality, taxonomic ancestry, lexical resemblance, and probability calibration are different properties. The current integration does not yet contain a learned transformation from news axes to a shared ontology.
8.1 Implemented layers
Eigen Times and Eigen Hacks share numerical code for projection, varimax naming, and explicit basis identity. They do not share a single fitted basis. Their stemmed term lists supply interpretable entry points, but terms such as databas are not automatically equivalent to an ontology concept with the canonical identifier databases.
Anthropology reuses QueryGraph’s pinned seed-v3 ontology: 157 concepts across 11 areas, including 61 concepts with multiple parents. This polyhierarchy differs from an orthogonal factor space. The picker searches concept names, aliases, and lexical representations, and stores an explicitly clicked concept ID as a personal tag on a canonical node. The broader library contains recommendation algorithms, but this picker uses its text search rather than a fitted tag-to-vector transformation. Private tags currently do not affect news coordinates, people PCA, or the factual graph. [18]
The pinned dependency provides code-level ontology provenance. The current tag rows, however, do not themselves carry an explicit ontology snapshot, calibration version, reviewer judgment, or article-level evidence. A long-lived calibration system would need those fields. We therefore use “ontology calibration” below for a proposed alignment and validation layer, not as a claim that personal annotations are already a calibrated scientific measurement.
8.2 A versioned, many-to-many crosswalk
The first extension should be a reviewed crosswalk keyed by corpus, basis kind, version, fingerprint, signed axis or pole, ontology snapshot, concept ID, evidence, method, and review state. Reviewers should inspect representative articles and both poles. Mixed axes may map to several concepts; unsupported mappings should remain unassigned. Shared concept IDs would provide common destinations across the sites while preserving each corpus’s coordinate system. SKOS distinguishes exact, close, broader, narrower, and related concept mappings [19]; such distinctions can inform reviewed concept-to-concept links, without treating a numerical axis as an already defined concept.
This crosswalk can improve retrieval without altering any eigenvector. A query for “databases,” for example, could enumerate reviewed axis associations in both corpora, show their different loadings and support, and link to relevant Anthropology records. A taxonomy relationship such as “narrower than” would remain a conceptual relation, distinct from a documented employment edge or a large cosine.
8.3 Learned semantic scores and their calibration
A later supervised layer could use reviewed article-level multilabel targets. Here indexes a corpus, is its number of annotated training articles, and is the number of shared ontology concepts. Let contain those articles’ standardized coordinate rows, and let contain the corresponding reviewed concept targets. Each row refers to the same article in both matrices; the value in a target column records that concept’s annotation under the chosen target protocol. The fitted slope matrix is denoted by and the fitted intercept column by .
For the optimization, let and be candidate slopes and intercepts, and let be the chosen ridge penalty strength. The all-ones column here has length , so repeats the intercept row for every article. The subscript denotes the Frobenius norm (the square root of the sum of squared matrix entries); denotes the parameter values minimizing the stated objective. A simple baseline leaves the intercept unpenalized and fits
(12)
The displayed objective assumes every target entry is observed; unreviewed labels must not silently become negative targets. Partially annotated data would require a separately specified masked objective. The intercept represents concept prevalence even when coordinates are centered; omitting it would require explicit target centering and later restoration.
Regularized logistic heads are an alternative when binary concept membership is the intended target. Each corpus receives its own , with the same ontology identifiers but separately evaluated parameters. Since concepts overlap, need not be square or orthogonal; it does not preserve the news reconstruction identities. Its outputs are scores until out-of-sample probability calibration is established [20]. A public implementation must specify the target event, annotation protocol, prevalence, and abstention policy before displaying a probability.
For this comparison, take , , and from the people model of the same corpus used to train the head. In the linear baseline, relates directional changes in people coordinates to concept-score slopes. A person’s reconstructed profile is ; applying the article-trained head to it is an out-of-distribution use until independently validated. A unit-normalized, background-adjusted person profile is not an average article, and the intercept does not repair this mismatch. For nonlinear heads, averaging article predictions and evaluating an aggregated profile are different estimands. Neither yields a probability that the person “is” a concept without a separately defined and evaluated target.
8.4 Cross-corpus and cross-version alignment
Shared entity IDs provide potential anchors, not coordinate equivalence. An orthogonal Procrustes fit [21] can align reviewed, comparable anchor representations when its assumptions are justified. It is neither an ontology reasoner nor an identity resolver. The two published models share 60 canonical people. A centered 60-dimensional fit with only 60 anchors has rank at most 59 before any holdout; a full-space alignment would therefore be underdetermined. A lower-dimensional or regularized method still requires disjoint test anchors and an analysis of population mismatch.
The current code offers component matching, while persistent cross-version lineage, near-degenerate subspace alignment, and warm-started rotations remain proposed. Daily production measures new articles in fixed bases rather than silently refitting them. This paper treats the earlier design discussion of nightly refinement as a research direction, not an observed property of this release.
9 Worked examples
9.1 A concrete news-to-people-to-news route
In interface labels, the prefix N followed by an index denotes a named news axis; the prefix P followed by an index denotes a people pattern. Thus N56 means news axis 56, and P19 means people pattern 19, within the stated corpus and model version. In the Hacker News hn1 model, N56 is labeled by the terms develop, interview, databas, ceo, manag, and github. Its strongest absolute people-pattern loading is
(13)
Starting with N56, the interface ranks P19 first by loading magnitude. Starting with P19, it returns the same coefficient for N56. The reverse index contains exactly the same stored number. This is the operational navigation duality in Figure 3.
The association should not be renamed “the database CEO eigenvector.” The term list is mixed, P19 has other positive and negative news loadings, and the people ranked along it span several technology sectors. Ellison and Siebel’s whole-profile similarity below is not explained by this one axis. In the general-news v2 basis, N56 has entirely different terms. The complete namespace, not the number 56, identifies the direction. A reviewed ontology crosswalk would help interpret this ambiguity without erasing it.
9.2 Jeff Dean: from a profile to axes, neighbors, and historical evidence
In the examples below, “24D cosine” compares retained people-coordinate rows , whereas “60D cosine” compares their full centered news-coordinate profiles . “D” denotes the number of dimensions. Both comparisons use nonzero rows from the same corpus and all-history window.
A browser receipt at 18:57 UTC on 2 October records a concrete route from Jeff Dean’s profile to news axes and neighboring coverage profiles. The profile offered Eigen Hacks links to N16 (googl, android, search, app), N44 (agent, model, robot, intellig), and N45 (test, review, studi, perform). These stemmed labels aid navigation without declaring an ontology concept or a personal relationship.
In the separately selected N24/P2 all-history view, Dean has 253 qualified candidate articles, a news coordinate of +0.433, and a people coordinate of +0.371. N24′s terms begin learn, machin, model, and deep; these coordinates do not describe the N16/N44/N45 links above. His closest displayed profiles include Sanjay Ghemawat (24D cosine 0.7514; 31 articles), Andrew Ng (0.7413; 394), and Noam Shazeer (0.7403; 38). Their corresponding full centered 60D cosines are 0.7082, 0.6380, and 0.6687. These values are reproducible coverage associations, not proof of collaboration or independently reviewed article identity. Following Ng leads to his own canonical profile; cross-site links open Eigen Hacks and Eigen Times while preserving their separate basis namespaces.
The view also links to a historical WIRED interview with Dean, published on 14 December 2019 [22]. The HN-derived evidence row carries 15 December 2019, illustrating why a corpus timestamp must not silently replace the publisher’s date. This is an inspectable historical source, not a successful match to today’s news. The audited current overlay has zero newly qualified article–person joins. It therefore supports navigation among profiles, axes, and dated evidence while leaving the current-news identity gap visible.
9.3 AI lineage: similarity and relationships corroborate different claims
The factual graph records LeCun’s postdoctoral work with Hinton [23], separately from Sutskever’s doctoral supervision [24]. It connects the 2012 AlexNet coauthors [25], DNNresearch’s founding and acquisition [26], and later laboratory formation. The new AI umbrella contains 71 people-and-organization nodes, four subgroups, and four waves through TypeSafe AI. Diogo Almeida’s InstructGPT coauthorship [27] and TypeSafe’s 2024 founding [28] provide a sourced later branch; no direct Anthropic-to-TypeSafe personnel edge is inferred.
In the HN people model, LeCun and Bengio have an all-history cosine of 0.9587, supported by 370 and 178 candidate articles. Their full centered 60-dimensional cosine is 0.9118. The 2015 paper Deep learning independently establishes coauthorship [29]. The cosine establishes a reporting association, not that coauthorship, mentorship, or agreement. The increase after truncation shows why a strong reduced-space similarity needs a full-space baseline. A further sensitivity analysis should remove articles mentioning both people and recompute their profiles; shared coverage can contribute directly to their similarity. That exclusion experiment has not yet been run.
The graph is richer than the current profile coverage. Of 44 people touched by the AI module, 22 were already in the scan dictionary, 20 have HN people profiles, and four have general-news profiles. The TypeSafe founders have no supported profiles in either published people model. A research journey can therefore follow their sourced graph records while the vector interface honestly reports missing coverage.
9.4 Founder similarity: a positive example and a counterexample
| Pair | Corpus | 24D cosine | 60D cosine | Articles, first / second |
|---|---|---|---|---|
| LeCun / Bengio | HN | 0.9587 | 0.9118 | 370 / 178 |
| Ellison / Siebel | HN | 0.7198 | 0.5800 | 416 / 11 |
| Zuckerberg / Gates | HN | 0.0782 | 0.0672 | 3,884 / 664 |
| Zuckerberg / Gates | General news | −0.1137 | −0.1104 | 3,560 / 2,469 |
Ellison and Siebel are mutual nearest neighbors in the HN model. This is a useful lead for exploring enterprise-software coverage, but Siebel’s eleven articles are close to the support threshold and no corresponding general-news profile is available. The result is neither cross-corpus replication nor an uncertainty estimate.
The proposed intuition that Zuckerberg should align with Gates is not supported by these models. Their HN cosine is close to zero, and their general-news cosine is mildly negative. Zuckerberg’s closest HN profiles include Chris Hughes, Brian Acton, and Sheryl Sandberg; Gates’s include Steve Ballmer, Paul Allen, and Satya Nadella. These associations are compatible with differentiated coverage histories. A negative cosine under a centered, signed basis does not indicate opposition, antagonism, or disapproval.
9.5 Database news and CEOs: a dated query
“Today’s database news against today’s database CEOs” requires a date-specific combination: concept-relevant articles on date , sourced CEO role intervals valid at , and qualified article-person links. Founder status is not a substitute for current office. The retained 1 October 2026 role snapshot lists Oracle’s Clay Magouyrk and Mike Sicilia [30]; MongoDB’s interim CEO Dev Ittycheria [31]; Snowflake’s Sridhar Ramaswamy [32]; and Databricks’s Ali Ghodsi [33]. Ellison is represented separately as Oracle’s founder and chairman/CTO at that date.
The audited interface places this sourced role layer beside the dated article projection. None of the five listed CEOs has a supported people profile in either published model. The overlay also lacks a newly qualified person join. Topic-ranked news and sourced roles are visible together, but a successful CEO-to-article match is not asserted. Mixed N56 labels can retrieve off-topic items, so the selected axis is not a validated database-news classifier. Expanded catalog identity coverage does not fill either missing statistical profiles or article-person links.
10 Reproducibility, evaluation, and limitations
Before reporting the audit, we define its numerical diagnostics. Retained between-person variance is the sum of the retained people-covariance eigenvalues divided by the sum of all people-covariance eigenvalues. For a nonzero centered person row , retained energy is ; the table reports the minimum and median over supported people. Maximum orthogonality discrepancy is the largest absolute entry of . Forward/reverse loading discrepancy is the largest absolute difference between the two stored copies of a loading coefficient. Each quantity is measured within one fitted corpus model.
| Measurement | Eigen Hacks / hn1 | Eigen Times / v2 |
|---|---|---|
| Qualified candidate articles | 52,467 | 24,322 |
| People before support/profile construction | 341 | 109 |
| Supported all-history people | 217 | 61 |
| News dimensions / people directions | 60 / 24 | 60 / 24 |
| Retained between-person variance | 84.54% | 90.86% |
| Minimum per-person retained energy | 58.43% | 79.27% |
| Median per-person retained energy | 85.21% | 91.41% |
| Maximum orthogonality discrepancy | ||
| Forward/reverse loading discrepancy | 0 | 0 |
The accompanying extraction script checks the stored loading matrix, recomputes all exported person scores, and reproduces the examples and summary statistics. Forward and reverse indexes agree exactly in the serialized artifact; score recomputation agrees at the arithmetic order used in the audit. The historical catalog audit separately reconciles frozen matcher outputs, registries, and publication receipts. A second enrichment audit records the later organization, news, and financing releases, their source limitations, and exact artifact hashes; the extractor verifies count partitions and publication identity without rerunning the collectors. A complete journal and matching counts establish operational completeness, not identity accuracy. These consistency results do not establish entity-linking precision, retrieval relevance, calibrated uncertainty, or research utility.
The first empirical study should separate four evaluations. Relationship extraction additionally needs a stratified reviewed holdout by relation kind, source, era, and abstention status; SEC event reconciliation needs reviewed amendment and issuer-mapping cases. Neither operational completeness nor successful publication supplies those evaluations. Entity resolution requires a stratified, manually reviewed sample of person-article pairs, including common names, early historical mentions, and quarantined cases. Retrieval requires judged article and person relevance, comparing lexical search, direct news profiles, truncated people profiles, and the proposed ontology layer. Temporal robustness requires held-out time periods, source exclusions, threshold sensitivity, and bootstrap intervals for neighbors. User evaluation should measure whether researchers find and verify relevant relationships faster, with fewer unsupported inferences, using the small-graph interface and the two navigation directions.
Calibration targets need independent annotations and reliability analysis; a similarity value is not a confidence probability. Any concept classifier should be evaluated on time-, source-, and entity-separated splits, with prevalence and abstentions reported. Private annotations require consent and review before any reuse as model labels. Existing ontology recommendation likelihoods and descriptive section-purity statistics cannot be substituted for such a calibration study.
Several limitations remain structural. The name dictionary and available archives select the population. The relationship pass inherits its roughly 370-name article-selection bottleneck even while discovering additional names within those articles. The SEC pass is restricted by filing quarter, industry, legal form, notice completeness, and amendment rules; issuer-reported sales are not necessarily closed venture rounds. Canonical identity checks cannot validate all source claims or establish that a filing concerns a similarly named product. English-language reporting and prominent institutions are overrepresented. Co-mentions introduce dependent observations, and repeated reporting can inflate apparent support. Equal cell weighting gives small supported historical cells substantial influence. A cell entering or leaving the five-article threshold can make temporal profiles discontinuous. Unit normalization removes volume and can amplify an almost-zero residual; future exports should retain its pre-normalization norm. Low-rank projection can increase selected similarities. Model signs are conventional, and near-degenerate axes can rotate under refits. Source and event time are imperfectly observed. Source- and time-blocked resampling, shared-article exclusions, and threshold sensitivity are needed before interpreting changes in people coordinates.
Provenance should describe the entire measurement process. A useful future export could map documents and model artifacts to PROV-O entities, fitting and review to activities, and publishers or reviewers to agents [34]. The current JSON fields provide practical provenance, but do not claim RDF or PROV-O conformance. A fitted model ID and a whole exported-file checksum serve different purposes: the latest edition can change without changing the historical people basis.
11 Conclusion
Anthropology makes the human side of technology history explorable through sourced relationships, while people vectors expose recurring patterns in reporting across two independent news corpora. The same loading coefficient supports both directions of navigation: news into people patterns and people patterns back into news. That reciprocity is exact as an index relationship and approximate as reconstruction. Its value depends on keeping numerical association, documented fact, semantic annotation, and current office distinct.
The completed identity pass expands searchable coverage while preserving unresolved decisions, namesake exclusions, and a small explorable graph. Later organization assertions, selected news classifications, and reported financing events expand research routes without converting co-mentions into relationships or unknown investors into named parties. Recoverable collection and immutable publication make those distinct expansions auditable. The next research step is a reviewed ontology crosswalk, followed by evaluated semantic calibration with explicit versions and uncertainty. Stable identities, fixed coordinate systems, dated evidence, and links back to articles provide the interfaces for that work; coverage counts alone do not supply its validation.
Reproducibility appendix
Edition 1.0.1 is a notation-only patch to the edition previously labeled 2026-10-02.1 and delivered as semantic version 1.0.0. It defines symbols before use, clarifies the article vector in Equation 5, and distinguishes the singular-value matrix from standardized news scores. It preserves the final catalog published on 2 October 2026 at 18:56:31 UTC, the unchanged 1 October vector snapshot, the 07:03:29 UTC identity-stage audit, and the enrichment audit for organization, news, and financing publication. No observations, fitted models, numerical examples, or empirical conclusions changed. The source bundle includes the manuscript, diagrams, screenshot provenance, extraction script, evidence.json, historical catalog-audit.json, and enrichment-audit.json.
From the pinned Anthropology checkout, run node papers/people-news-duality/extract-evidence.mjs to recompute vector measurements and incorporate the two audits. The script verifies the current count partitions and the unchanged people-vector input. The seed graph’s record counts are unchanged, but its funding metadata and file hash have changed; the two editions’ hashes are therefore retained separately.
With the locally retained inputs, add --verify-catalog --artifact-root=/path/to/original-checkout to check every artifact’s exact byte length and SHA-256. For historical tracked files, the verifier reads Git objects at the original audit commit, rather than expecting current registry files to retain old bytes. Ignored raw inputs are resolved under the specified artifact root. The 105,338,217-byte identity snapshot and 270,937,620-byte enriched snapshot are not redistributed in the paper bundle. Their absence does not prevent building the manuscript from its included evidence and figures. The script does not crawl, refit, or mutate its inputs; compact audit manifests are not substitute datasets for independently repeating identity and relationship collection.
Source revisions.
Anthropology current implementation: 89058a16f3a7c34f9a117a3b246ee18633ff5aa8
Historical identity audit: 3a05569d67d99365a766add25db88197474fad5c
Eigen producer: ecef236a949241e7adf8576a450e5f0dcabe1668
QueryGraph ontology: 70cb97aa257afe58b96bd1a340a073b0514a1be2
People model identifiers.eigenhacks:eigen:hn1:people:v1:bcaccd95b794eigentimes:eigen:v2:people:v1:47debba3db48
Full fitted-model fingerprints (SHA-1 identifiers, not signatures).
HN: bcaccd95b7941b6e64fadde195d854b53ebdb8bc
News: 47debba3db480a75d579f015066ea49b23dd72ff
Enriched catalog revision (immutable logical identity).a83e02b2d9958f479f3481d8cad90637b63fc544e094931072ea81ea08a1a73a
Catalog file SHA-256.92285426b98123cad0e9f24e22a37fa61608411734998951d147be932f5e06f5
Frozen matcher input SHA-256; journal sequence 2,586.cb06b7089bdecc12808d7ecd3fdeca1fdf1db461d812cd1a900a40fb7548731a
Primary implementation locations. Eigen’s scripts/build-people-vectors.mjs creates the qualified exchange; crates/et-people/src/lib.rs constructs and decomposes person profiles; scripts/refresh-people-edition.mjs refreshes the overlay. Anthropology’s people-vector-explorer.tsx, vector-edition.tsx, and ontology-picker.tsx implement reciprocal navigation, dated article/role views, and personal annotation. Its scripts/match-wikipedia.mjs and scripts/lib/catalog-export.mjs reconcile and freeze identities; scripts/classify-news-relationships.mjs and scripts/collect-organization-relationships.mjs collect relationship evidence; scripts/collect-sec-funding.py collects offerings; scripts/merge-funding-events-catalog.mjs reconciles their identities and events. backend/src/lib.rs and backend/src/import.rs serve and publish revisions. src/lib/archive-evidence.ts and src/lib/server/archive-storage.ts preserve scoped evidence and bounded storage access. Full news-basis fingerprints and source hashes appear in the accompanying evidence and audit manifests.
The fitted-model fingerprint excludes mutable display metadata and current-edition content. The exported-file checksum in the accompanying evidence manifest is therefore needed to reproduce the exact snapshot. For a future snapshot, new data should be published under a new artifact identity; the figures in this paper must not be silently relabeled as current.
Public entry points: Anthropology vectors · Eigen Hacks people vectors · Eigen Times people vectors. The live graph screenshot records its requested URL, resolved canonical node, timestamp, viewport, and SHA-256 checksum. It contains the graph workspace only and reproduces no NYT image or article body.
References
- [1] A. Khrabrov, “Eigen Times: A Newspaper in the Eigenbasis of the News.” [Online]. Available: https://firstpair.org/library/eigentimes/
- [2] S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by latent semantic analysis,” Journal of the American Society for Information Science, vol. 41, no. 6, pp. 391–407, 1990, doi: 10.1002/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9.
- [3] I. T. Jolliffe and J. Cadima, “Principal component analysis: a review and recent developments,” Philosophical Transactions of the Royal Society A, vol. 374, no. 2065, p. 20150202, 2016, doi: 10.1098/rsta.2015.0202.
- [4] D. M. Blei and J. D. Lafferty, “Dynamic topic models,” in Proceedings of the 23rd International Conference on Machine Learning, 2006, pp. 113–120. doi: 10.1145/1143844.1143859.
- [5] W. L. Hamilton, J. Leskovec, and D. Jurafsky, “Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 2016, pp. 1489–1501. doi: 10.18653/v1/P16-1141.
- [6] J. Hoffart, F. M. Suchanek, K. Berberich, and G. Weikum, “YAGO2: A spatially and temporally enhanced knowledge base from Wikipedia,” Artificial Intelligence, vol. 194, pp. 28–61, 2013, doi: 10.1016/j.artint.2012.06.001.
- [7] QueryGraph, “Anthropology: sourced relationships, financing events, and vector navigation.” [Online]. Available: https://anthropolo.gy/vectors
- [8] Wikidata contributors, “Wikidata: Data access.” Accessed: Oct. 02, 2026. [Online]. Available: https://www.wikidata.org/wiki/Wikidata:Data_access
- [9] MediaWiki contributors, “API: Etiquette.” Accessed: Oct. 02, 2026. [Online]. Available: https://www.mediawiki.org/wiki/API:Etiquette
- [10] M. Koster, G. Illyes, H. Zeller, and L. Sassman, “Robots Exclusion Protocol,” technical report RFC 9309, 2022. doi: 10.17487/RFC9309.
- [11] U.S. Securities and Exchange Commission, “Form D Data Sets.” [Online]. Available: https://www.sec.gov/data-research/sec-markets-data/form-d-data-sets
- [12] Vercel, “AWS with Vercel OIDC.” Accessed: Oct. 02, 2026. [Online]. Available: https://vercel.com/docs/oidc/aws
- [13] S. Xiao, Z. Liu, P. Zhang, N. Muennighoff, D. Lian, and J.-Y. Nie, “C-Pack: Packed Resources For General Chinese Embeddings.” [Online]. Available: https://arxiv.org/abs/2309.07597v5
- [14] Beijing Academy of Artificial Intelligence, “BGE small English v1.5: Model card.” Accessed: Oct. 02, 2026. [Online]. Available: https://huggingface.co/BAAI/bge-small-en-v1.5
- [15] A. Khrabrov, “Eigen Times source and people-vector implementation.” [Online]. Available: https://github.com/alexy/eigentimes/tree/ecef236a949241e7adf8576a450e5f0dcabe1668
- [16] H. F. Kaiser, “The Varimax Criterion for Analytic Rotation in Factor Analysis,” Psychometrika, vol. 23, no. 3, pp. 187–200, 1958, doi: 10.1007/BF02289233.
- [17] T. R. Gruber, “A translation approach to portable ontology specifications,” Knowledge Acquisition, vol. 5, no. 2, pp. 199–220, 1993, doi: 10.1006/knac.1993.1008.
- [18] QueryGraph, “Ontology engineering and application specialization contract.” [Online]. Available: https://github.com/querygraph/ontology/blob/70cb97aa257afe58b96bd1a340a073b0514a1be2/ONTOLOGY.md
- [19] A. Miles and S. Bechhofer, Eds., “SKOS Simple Knowledge Organization System Reference.” [Online]. Available: https://www.w3.org/TR/skos-reference/
- [20] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in Proceedings of the 34th International Conference on Machine Learning, in Proceedings of Machine Learning Research, vol. 70. 2017, pp. 1321–1330. [Online]. Available: https://proceedings.mlr.press/v70/guo17a.html
- [21] P. H. Schönemann, “A generalized solution of the orthogonal Procrustes problem,” Psychometrika, vol. 31, no. 1, pp. 1–10, 1966, doi: 10.1007/BF02289451.
- [22] T. Simonite, “Google's AI Chief Wants to Do More With Less (Data).” [Online]. Available: https://www.wired.com/story/googles-ai-chief-do-more-less-data/
- [23] G. Hinton, “Geoffrey Hinton's former postdocs.” Accessed: Oct. 01, 2026. [Online]. Available: https://www.cs.toronto.edu/~hinton/postdocs.html
- [24] I. Sutskever, “Training Recurrent Neural Networks,” Doctoral dissertation, 2013. [Online]. Available: https://www.cs.utoronto.ca/~ilya/pubs/ilya_sutskever_phd_thesis.pdf
- [25] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in Advances in Neural Information Processing Systems, 2012. [Online]. Available: https://papers.nips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html
- [26] University of Toronto, “Google acquires U of T neural networks company.” Accessed: Oct. 01, 2026. [Online]. Available: https://www.utoronto.ca/news/google-acquires-u-t-neural-networks-company
- [27] L. Ouyang et al., “Training language models to follow instructions with human feedback.” [Online]. Available: https://arxiv.org/abs/2203.02155
- [28] TypeSafe AI, “TypeSafe AI Emerges From Stealth With 40 Million Dollars in Funding With New Model for Composable AI.” [Online]. Available: https://www.businesswire.com/news/home/20260915525333/en/TypeSafe-AI-Emerges-From-Stealth-With-%2440M-in-Funding-With-New-Model-for-Composable-AI
- [29] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, pp. 436–444, 2015, doi: 10.1038/nature14539.
- [30] Oracle, “Board of Directors.” [Online]. Available: https://www.oracle.com/corporate/executives/board-of-directors/
- [31] MongoDB, “Form 8-K: September 2026 leadership transition.” [Online]. Available: https://investors.mongodb.com/static-files/dc7beb9f-6084-46eb-9c15-cf65c30b999c
- [32] Snowflake, “Form 8-K: CEO appointment.” [Online]. Available: https://www.sec.gov/Archives/edgar/data/1640147/000164014724000033/snow-20240227.htm
- [33] Databricks, “From Soda Hall to the Football Field: Databricks Returns to Berkeley Roots With Cal.” [Online]. Available: https://www.databricks.com/company/newsroom/press-releases/soda-hall-football-field-databricks-returns-berkeley-roots-cal
- [34] T. Lebo, S. Sahoo, and D. McGuinness, Eds., “PROV-O: The PROV Ontology.” [Online]. Available: https://www.w3.org/TR/2013/REC-prov-o-20130430/