Eigenvectors come with two kinds of freedom that the mathematics does not fix and a newspaper must.
If is an eigenvector, so is : the same line pointing the other way. A score is measured by a dot product with the direction, so flipping the direction also flips its scores. The contribution to reconstruction is unchanged because for a scalar score . Its positive and negative poles, the two ends of the line, exchange names. Choosing a sign gives the display a consistent convention without changing the represented observations.
Eigen Times fixes orientation using skewness, a measure of asymmetry in the score distribution. Here a distribution describes how frequently different score values occur, and a tail is the part extending toward unusually large or small scores. A central moment averages powers of deviations from the mean. We already know the second central moment: averaging squared deviations gives variance. The third central moment instead averages cubed deviations. Cubing retains signs, so a large negative deviation contributes negatively and a large positive deviation contributes positively.
Let and denote the second and third central moments on one axis. For the notebook’s scores −5, −1, 0, 1, 1, the mean is −0.8. The centered scores are −4.2, −0.2, 0.8, 1.8, 1.8. Squaring and averaging gives ; cubing and averaging gives . The negative third moment says that, by this signed cubic measure, the negative side is more pronounced.
When , standardized skewness is . The power means the standard deviation cubed: . This division removes the units of the scores, but it does not change the sign. The orientation rule therefore needs only the sign of : flip if , making the third moment positive. Reversing all centered scores preserves their squares and reverses their cubes. This is a convention based on the measured third moment, not a claim that every asymmetric distribution has one unambiguous longer tail. A zero third moment supplies no preferred sign. The streaming chapter computes these moments without retaining every score.
Equal eigenvalues are called degenerate: any perpendicular unit pair in their plane is an eigenbasis. For example, a covariance equal to twice the identity multiplies every direction by two. There is no uniquely preferred diagonal in its plane because every unit direction has variance two. The notebook rotates this covariance and verifies that its entries remain unchanged.
Close eigenvalues can make individual directions sensitive to small data changes, even when their shared subspace is stable (Figure 9). A perturbation is such a change to the covariance; the eigenvalue gap is the separation from neighbouring eigenvalues. The tracking chapter explains the bound relating them. Even distinct principal axes can be difficult to name: variance-maximising directions may blend subjects such as markets and party politics. Naming keeps the selected subspace and chooses different axes inside it; the rotated axes need no longer be covariance eigenvectors. The purpose of the new coordinates is readability, while the retained geometric information stays the same.
Recall the retained orthonormal basis . Let be a orthogonal rotation and let be the named basis; the prime here means rotated. If is a raw coordinate column, its named column is (also sometimes written ). Why does the coordinate transformation use the transpose? The coordinates must change oppositely to the basis so that the represented point stays fixed. Substituting and using gives
Both bases therefore reconstruct the same projection. Subtracting that same projection from an observation leaves the same residual, so its length is unchanged. The covariance-adjusted distance called in §7 is also unchanged when its covariance is transformed consistently; it is not determined by the subspace alone.
Kaiser’s varimax criterion selects a rotation making each direction load strongly on some features and weakly on others. A criterion, or objective, is a numerical rule for comparing candidate choices. A larger value of this criterion indicates greater contrast in the sizes of loadings within the columns. It favors concentration but does not force entries to become exactly zero.
Let count the features, let index them, and let index axes. Write for row , column of the rotated loading table. Hold one column fixed. Square each loading, average those squared loadings, then measure how much the squared loadings vary around their mean. Squaring treats equally strong positive and negative loadings equally; it is concentration of magnitude that matters here.
We already derived the identity “variance equals mean square minus square of mean.” Apply it to the values . Their squares are , so this column’s variance is the average fourth power minus the square of the average second power. Add those variances over columns. That sequence of operations gives the criterion
The fourth power appears because it is the square of a squared loading. It is not a new assumption that extreme feature values should be counted four times. For a two-feature column with loadings 1 and 0, squared loadings are 1 and 0, their mean is 0.5, and the variance of those squared loadings is . Spread the same total squared loading equally by using loadings and . The squared loadings are now 0.5 and 0.5, so their variance is zero. The first column has greater concentration even though both columns have squared length one. In a rotation we cannot choose each column independently: the whole loading table must turn together. The notebooks compute the variance of squared loadings by directly subtracting each column’s mean and also by the displayed formula; the two methods agree.
Figure 10 shows six variables in two groups mixed by 30°. Their small secondary loadings make the computed maximizing rotation approximately , rather than an exact reversal of the mixing angle. The rotation concentrates each variable on one axis without making its other loading exactly zero.
Eigen Times sweeps over every pair of axes, rotating just that pair by its best angle. This two-coordinate operation is a Givens rotation. It leaves the other columns unchanged, reducing one step of a many-axis search to choosing one angle.
Let (Greek phi) denote that angle. Angles may be measured in degrees, with 360° in a full turn, or in radians, where angle is arc length divided by circle radius. One full turn is radians, with the circle constant. On the unit circle, is the horizontal coordinate and the vertical coordinate after turning through from the positive horizontal direction. Their squared sum is 1 because the point lies on that circle. A two-dimensional rotation matrix is
For one pair only, let be its loading columns, distinct from article and score vectors elsewhere. Multiplying the loading table by changes row ’s pair to
These two entries’ squared sum remains . Rotation redistributes a row’s contribution between axes while preserving its total squared size.
We can evaluate the criterion for many trial angles, but the two-column case also has a direct formula. Elementwise arithmetic treats each row separately. Define as the difference of the squared loadings and as twice their product. Define four scalar sums
These local capital letters are scalars, not earlier matrices. The first two sums measure the column totals of the new quantities; the last two collect their squared and cross-product terms. Define centered combinations and . The subscript star marks these adjusted scalar values.
Here is the route from the criterion to the angle. After rotation, the difference of the two squared loadings is , obtained by expanding the two squared rotated entries and using the double-angle identities. In those identities, and . Substituting this difference into the variance criterion and expanding once more leaves an angle-independent part plus
The appearance of comes from squaring expressions that already involve . We can maximize the displayed part geometrically: it is a dot product between the fixed vector and the unit vector , divided by a positive constant. The dot product is largest when the unit vector points along the fixed vector. The function returns the angle of the point with horizontal coordinate and vertical coordinate , preserving the quadrant that a ratio alone would lose. It supplies ; divide by four to get
This is a formula for choosing a maximizing angle, not a requirement to memorize trigonometric manipulation. The notebooks check its objective against a dense grid of trial angles and verify the angle-dependent expression. If both arguments of are zero, that expression is constant and the pair supplies no preferred rotation.
A sweep visits every axis pair. Sweeps stop when no pair rotates by more than radians (one millionth of a radian), or after fifty sweeps. Exact pairwise maximization cannot decrease the criterion, so its bounded objective improves monotonically, meaning it never moves downward. This does not guarantee the global best orientation: the choices across many column pairs interact. The reported alternative, an SVD-based fixed-point iteration (repeatedly applying one update rule), oscillated between two mirror-image states on a symmetric configuration, motivating the pairwise method.
For the embedding basis, naming rotates associations with words rather than unexplained embedding coordinates. The articles supply the common rows connecting the two descriptions: each has term measurements and raw eigen-coordinates.
Let be the article TF‑IDF matrix, its article mean, and a column of ones. Let be an table of raw eigen-coordinates divided by their positive raw-axis standard deviations. Dividing a coordinate by its standard deviation makes one unit mean one reference standard deviation on every axis. Because the unrotated eigen-coordinates already have zero reference mean and zero reference cross-covariance, this scaling gives unit reference variance and zero reference cross-covariance. Together those properties are called whitening. Scaling arbitrary correlated coordinates to unit variance would not by itself remove their cross-covariances.
Let (Greek beta) be the term-loading matrix. Its entry for one term and one axis sums, over articles, the centered term value multiplied by that article’s standardized raw score. Positive products arise when the term and score are both above their respective references, or both below. The sum measures association between term use and the direction; it is not a probability and has not been divided by the number of articles. Matrix multiplication performs all these sums together:
Here and . Lexical whitening floors its raw variance denominators at before taking square roots: if a variance is smaller, the implementation uses that positive minimum to avoid division by zero or an excessively small scale. Exact unit-variance whitening applies to positive axes whose denominators have not been changed by this safeguard, under the same reference weights used to fit them.
The embedding naming fit rotates ; the LSA fit rotates its right singular directions, already indexed by terms. Applying the resulting to the corresponding raw coordinate map gives named directions. Because scaling and rotation need not give the same result in opposite orders, word associations fitted from whitened scores are labels for investigation, not exact term-by-term decompositions of every stored named score. The block-sum chapter expands the centered loading formula. The older shorthand assumes compatible centered, scaled score columns; the SVD factor must not generally be substituted for . The distinction is practical: a unit left singular direction, an article coordinate, and a variance-standardized coordinate carry different scales even when they point along related patterns.