← First Pair Library

5 A population of profiles makes a people basis

5.1 Center the rows around the population

Let nn now count the fitted people, so BB has shape n×Kn\times K. Its rows have unit length, but they need not average to zero. Denote their 1×K1\times K mean row by b‾\bar b. In the example it is

b‾≈(0.1276,0.0652,0.1407).\bar b\approx(0.1276,0.0652,0.1407).

Let 𝟏\mathbf1 denote a column of nn ones, so 𝟏b‾\mathbf1\bar b repeats the mean row for every person. Subtract it from BB and call the resulting n×Kn\times K centered matrix HH:

H=B−𝟏b‾.H=B-\mathbf1\bar b.

Let hp=bp−b‾h_p=b_p-\bar b denote person pp’s row of HH. For A,

hA=bA−b‾≈(0.3411,−0.6500,0.5213).h_A=b_A-\bar b\approx(0.3411,-0.6500,0.5213).

This second centering asks how A’s normalized coverage differs from the fitted population of normalized person profiles. The earlier subtraction compared articles inside a cell. The reference objects are different.

5.2 Covariance asks which coordinates move together

Let CPC_P denote the K×KK\times K people covariance. Using the fitted-person count nn, compute it as

CP=1nH𝖳H,C_P=\frac1nH^{\mathsf T}H,

Here jj and kk each index news coordinates from 1 to KK. Entry (j,k)(j,k) is the average product of those two centered coordinates across people. A large positive product tends to occur when both coordinates deviate in the same direction; a negative product occurs when they deviate oppositely. The diagonal entries are coordinate variances.

Each person contributes one row. A person supported by 10,000 articles does not enter the covariance a thousand times more heavily than a person supported by ten. More articles can affect the quality of a row, but they are not its weight in this fit. This is equal-person PCA: principal component analysis of normalized, centered coverage profiles.

Let ℓ\ell index a people direction, wℓw_\ell its K×1K\times1 unit eigenvector column, and eℓe_\ell its covariance eigenvalue. Each pair satisfies

CPwℓ=eℓwℓ.C_Pw_\ell=e_\ell w_\ell.

Multiplying the direction by the covariance stretches it by eℓe_\ell without turning it. As explained in the Eigen Times eigenvector chapter, the largest eigenvalue identifies the direction of greatest variation, and the next identifies the greatest variation perpendicular to it.

5.2.1 Read one matrix entry as an average

For the original four-person fixture, the first centered coordinate is approximately 0.3411 for A, 0.8207 for B, -0.3459 for C, and -0.8159 for D. The N1 variance therefore adds their four squares and divides by four, giving approximately 0.393847. The N1–N3 covariance instead multiplies each person’s first and third centered entries and averages the four products, giving approximately 0.138797. Both calculations use one product per person. The notebook prints the complete centered table and checks the covariance against independent reference entries.

5.2.2 Verify an eigenvector with two coordinates

Set the four-person matrix aside for a small covariance illustration. Let C*=(2112)C_*=\begin{pmatrix}2&1\\1&2\end{pmatrix}. The star marks this separate synthetic table. Let a+=(1,1)𝖳/2a_+=(1,1)^{\mathsf T}/\sqrt2 and a−=(1,−1)𝖳/2a_-=(1,-1)^{\mathsf T}/\sqrt2 be unit columns. Multiplying the first gives (3,3)𝖳/2=3a+(3,3)^{\mathsf T}/\sqrt2=3a_+. Multiplying the second gives (1,−1)𝖳/2=a−(1,-1)^{\mathsf T}/\sqrt2=a_-. They are eigenvectors with eigenvalues 3 and 1; their dot product is (1−1)/2=0(1-1)/2=0.

For a unit column aa, the scalar a𝖳C*aa^{\mathsf T}C_*a is variance along that direction. Along the first coordinate axis it is 2. Along a+a_+ it is 3; along a−a_- it is 1. To see why 3 is the maximum, express any unit direction as a=γa++δa−a=\gamma a_++\delta a_-, where γ,δ\gamma,\delta are scalar coefficients satisfying γ2+δ2=1\gamma^2+\delta^2=1. Its variance is 3γ2+δ2=1+2γ23\gamma^2+\delta^2=1+2\gamma^2, no larger than 3. The larger-eigenvalue direction captures the most spread because of this numerical property, not because its words are necessarily more important.

Multiplying an eigenvector by -1 leaves its line and eigenvalue unchanged. Its coordinates and resulting scores both reverse signs. A sign convention is useful for reproducible displays; it does not change the represented variation.

5.3 Keep two patterns in the toy world

The toy covariance has eigenvalues approximately 0.49700.4970, 0.28170.2817, and 0.18100.1810. Keeping the first two retains

0.4970+0.28170.4970+0.2817+0.1810≈0.8114\frac{0.4970+0.2817}{0.4970+0.2817+0.1810}\approx0.8114

of the between-person variance, or 81.14%. This is a reconstruction statement about this four-person cloud. It is not an accuracy estimate, an identity confidence, or the fraction of articles explained.

Let rr count the retained positive people directions; it is 2 in this toy example. The bare rr is a count, distinct from the contrast row rpr_p. The retained direction index ℓ\ell runs from 1 to rr. Put those eigenvectors, ordered by decreasing eigenvalue, into the columns of the K×rK\times r loading matrix WW:

W≈(0.7964−0.35160.10900.88380.59490.3087).W\approx\begin{pmatrix} 0.7964&-0.3516\\ 0.1090&0.8838\\ 0.5949&0.3087 \end{pmatrix}.

Its rows correspond to N1, N2, N3; its columns correspond to people patterns P1 and P2. The columns have length one and are perpendicular. Their signs are oriented so the largest-magnitude news loading is positive, matching the implementation’s sign convention. Flipping a column and all its scores would describe the same geometry. When two eigenvalues are close, a small change in data can rotate their directions substantially even while their combined subspace changes little. Fixing signs does not remove that ambiguity; it is another reason to keep model versions explicit.

Unit person profiles are centered, their covariance is decomposed, and two retained directions give each person two scores. The discarded third eigenvalue is visible rather than silently erased.

5.3.1 Account for the omitted variance

The three fitted eigenvalues sum to approximately 0.9596456. This is also the sum of the three diagonal entries of CPC_P, called its trace: total centered squared length per person, expressed in either coordinate system. The first two sum to approximately 0.7786802. Their ratio is 0.8114248; multiplying by 100 expresses it as a percentage. The omitted fraction is about 0.1885752.

The loading columns are unit directions, but a row of WW need not have length one. Similarly, the score rows need not be unit rows after population centering and projection. Keeping these distinctions prevents a loading, a coordinate, and a cosine from being treated as interchangeable numbers.

5.4 Locate a person on the patterns

Let upu_p denote person pp’s 1×r1\times r people-coordinate row. Project its centered profile hph_p onto the retained columns:

up=hpW=(bp−b‾)W.u_p=h_pW=(b_p-\bar b)W.

For A, the first coordinate is approximately

uA1≈0.3411(0.7964)−0.6500(0.1090)+0.5213(0.5949)≈0.5109.\begin{aligned} u_{A1}&\approx0.3411(0.7964)-0.6500(0.1090)\\ &\quad+0.5213(0.5949)\\ &\approx0.5109. \end{aligned}

The second uses the second column of WW. The resulting rows are:

Person P1 score P2 score
A 0.5109 -0.5334
B 0.5828 -0.1178
C 0.0813 0.8807
D -1.1750 -0.2294

P1 is a pattern, not person D, although D has the largest absolute P1 score in this small example. A negative score identifies the other pole of a direction. It does not say that D is opposed to the other people.

In the published models, there are 60 news coordinates and at most 24 directions with positive eigenvalues. The 217-person Hacker News model retains 84.54% of between-person variance; the 61-person general-news model retains 90.86%. These separately fitted percentages should not be read as a contest between corpora. There is no second varimax naming step and no whitening of people scores. The people directions are orthonormal in standardized news coordinates, not generally when mapped back into the original embedding geometry.

5.4.1 Show both dot products

A’s first score adds approximately 0.2717−0.0709+0.3101=0.51090.2717-0.0709+0.3101=0.5109. Its second coordinate is

uA2≈0.3411(−0.3516)−0.6500(0.8838)+0.5213(0.3087)≈−0.1199−0.5745+0.1609≈−0.5334.\begin{aligned} u_{A2}&\approx0.3411(-0.3516)-0.6500(0.8838)\\ &\quad+0.5213(0.3087)\\ &\approx-0.1199-0.5745+0.1609\\ &\approx-0.5334. \end{aligned}

These are projections of the same centered row onto different direction columns. Small differences in final digits arise if we multiply displayed four-decimal entries rather than the full-precision fixture.

The first score column has variance 0.4970 and the second 0.2817, the eigenvalues already reported. They are not divided by the square roots of those variances in this model. Dividing them would whiten the retained scores and would change distances and cosines; that is a different geometry.