The question is now about one observation: how much of it can the retained directions reconstruct, and what do they leave out? Here an observation is an article, or a representative vector chosen for a story. It is already a numerical vector; we are not yet assigning it a claim about historical novelty.
Let count its original coordinates and count retained directions. Write for the observation column and for the fitted article mean. The matrix contains the retained raw directions as columns. They are orthonormal: each has length one and different columns have dot product zero. The following steps have different jobs.
Subtract the reference point. The centred column is the displacement from the fitted mean. Projection concerns this displacement, so the mean must be removed before measuring coordinates.
Measure along each retained direction. Define the coordinate column . Entry , with , is the dot product with retained direction . Stacking those dot products is matrix multiplication:
Rebuild what those coordinates describe. Multiplying by takes the weighted sum of the retained directions. Adding the mean puts the result back in the original coordinate system. Call this reconstruction ; the hat marks an approximation:
Subtract the reconstruction from the observation. The difference is the residual: what remains after the attempted reconstruction. The word names a difference vector, not its length and not its probability:
For a complete synthetic example, retain the first two coordinate directions in three dimensions. Take and . Then
The first coordinate survives unchanged, the second survives unchanged, and the third is absent from the retained plane. Subtraction reveals that absence explicitly: in the last residual entry.
An orthogonal projection keeps the component in the chosen subspace and leaves a residual perpendicular to it. Here this can be checked by taking its coordinates again. Since , where is the identity,
Thus every retained direction has zero dot product with the residual. This is also why the reconstruction is a least-squares one. Any change to the reconstructed point within the retained subspace adds another perpendicular component to the error. Its squared length adds to the existing residual squared length, so it cannot improve the fit. No differentiation is needed for this geometric argument.
Recall the two distinct operations in a length calculation. Squaring entries and adding them gives squared length; taking its square root gives length. For our residual,
The centred observation has squared length . The retained part has squared length . The Pythagorean identity reads : retained variation and residual variation account for the whole squared length. We use squared lengths below because perpendicular contributions add in this way. The ordinary lengths do not add: is not .
We need two different comparisons. One asks how far the story lies along the retained directions, taking each direction’s usual spread into account. The other asks how much of the story lies outside those directions. A statistic is a numerical summary calculated from data; calling something a statistic does not by itself assign a probability to it.
Recall that variance is an average squared deviation from a mean, and standard deviation is its nonnegative square root. Raw coordinate has fitted mean zero and variance . For this calculation retain only positive eigenvalues, so dividing by is possible. The quotient expresses the coordinate in units of its own standard deviation. A displacement of two units is ordinary on an axis whose usual spread is two, but larger relative to an axis whose usual spread is one.
A z-score subtracts a reference mean and divides by a positive reference standard deviation. The mean subtraction has already happened for the raw coordinates. Their variances can differ, but their fitted cross-covariances are zero. Dividing each by its own standard deviation therefore gives whitened coordinates: unit variance in every direction and zero cross-covariances under that same fitted reference.
Define , Hotelling’s squared covariance-adjusted distance, by squaring these standardized coordinates and adding. Here is the statistic’s name, not the square of the TF‑IDF matrix . With ,
Our synthetic story has . Suppose the fitted raw variances are and , hence standard deviations and . Its standardized coordinates are . Squaring and adding gives
Both retained coordinates are one reference standard deviation from their means, although their original sizes differ.
For compact matrix notation let be the diagonal matrix of those positive raw eigenvalues. A superscript denotes an inverse, the operation undoing multiplication by an invertible matrix. For a diagonal matrix, inversion replaces each nonzero diagonal entry by its reciprocal. Consequently first divides coordinate by ; the dot product with then multiplies by again and adds:
Define to be residual squared length:
For the same story, . This is a different measurement from . A point can be far along a retained direction and have no residual at all; another can lie close to the mean within the retained directions but have a substantial perpendicular residual. A residual can contain discarded familiar variation as well as genuinely new material. Neither statistic alone proves that an event is unprecedented.
Figure 12 shows why reference spreads matter: equal Euclidean distances can have different covariance-adjusted distances. With 60 fitted positive directions, the fitted average of under the matching covariance weights is 60. To see why, average each term : the numerator’s fitted average is precisely , so every term contributes one. A value of 400 is much larger than that fitted average. This arithmetic does not yet provide the probability of reaching 400; that would require a justified probability model or a separately assessed reference distribution.
The calculation uses raw eigen-coordinates before the naming rotation. The rotation changes the coordinate descriptions without changing the retained point. Recall the orthogonal naming rotation . In the named coordinate column , define the covariance . Its off-diagonal entries need not be zero: named coordinates can vary together even though the raw eigen-coordinates did not.
To measure the same distance in this new coordinate system, use its full covariance:
The inverse now adjusts for both different spreads and shared movement between coordinates. Reusing the old diagonal eigenvalues would attach old scales to new directions. Dividing by the named coordinates’ own variances is also insufficient when their covariances are nonzero. That operation is marginal standardization: it gives each coordinate unit variance separately but leaves correlations. Full whitening removes cross-covariances as well. The notebooks rotate the synthetic story and verify that the full calculation stays unchanged while the diagonal shortcut generally changes it.
The implementation sums only for positive raw eigenvalues and contributes zero otherwise. This is an explicit handling rule for zero-variance directions, not permission to divide by zero. The inverse formulas above assume positive retained eigenvalues.
The two statistics originated in monitoring variation against a principal-component reference model. For browsing, it is also useful to ask what share of a centred observation remains outside the retained subspace. Let (Greek nu) denote this novelty ratio. When , meaning the observation differs from the fitted mean, define
The denominator counts all centred squared length; the numerator counts only the residual part. In our example, . About 64.3% of this observation’s squared length lies outside the retained plane. That is a statement about this representation, not a 64.3% probability of a historically new event.
Orthogonality makes the denominator the sum of retained and residual squared lengths, both nonnegative. The ratio therefore lies between zero and one. Zero means the centred vector is fully represented; one means it lies perpendicular to the retained subspace. At , both numerator and denominator are zero, so the ratio is undefined. The implementation returns zero in that case as a software convention. This geometric use of “energy” means squared length; the attention score below uses the same ordinary word for a different quantity.
These equations describe one vector. The site’s stored story statistics are means of the member articles’ , , and values. The order matters: measure every member, then average the measurements. Measuring the mean embedding answers a different question.
For example, two member coordinate columns and under variances each have , so their mean statistic is one. Their mean coordinate column is zero, whose is zero. Averaging first cancels their opposite departures. Squared distances and ratios generally do not commute with averaging; that is why an article statistic, an average article statistic, and a centroid statistic must be kept distinct.
These three objects answer questions at three different levels. The dominant axis selects a story’s label. The spectrum combines the day’s stories. A profile describes the stories usually carrying one label.
For story , let be its named-coordinate column, with entry on axis . Here identifies a story. The separate notation denotes the empirical standard deviation of named coordinate across stories in the profile window: the latest ten years of loaded data. This is a reference scale, estimated separately from the raw eigenvalue . The implementation replaces scales smaller than by that positive floor to avoid division by zero or an extremely small denominator.
Divide each named coordinate by its reference scale. The dominant archetype is the axis with the largest signed result. The notation means “return the index attaining the largest value”, whereas would return the value itself. Thus the rule is .
Take the notebook’s synthetic story and scales . The scores are and , so the second axis wins despite its smaller raw coordinate. For a second story , the scores are and . The second wins again because is larger. Taking absolute values would change that decision. This particular rule does not subtract an empirical axis mean, so it is scale division rather than a general mean-subtracted z-score.
One possible “mixed” heuristic compares the positive top score with the runner-up . Their relative gap is ; the denominator is the top score. A rule declaring a mixture when that gap is at most would accept scores and , whose gap is . This is an illustrative rule retained from the original explanation, not a rule implemented by the current site.
A simple average gives every story equal influence. A weighted mean first multiplies each value by its nonnegative weight, adds those products, and divides by the positive total weight. Division by the total makes the result an average rather than a total that grows automatically when more stories are present.
Let denote story ’s attention weight, called story energy in the site:
The natural logarithm grows slowly; adding one makes the argument positive even for a zero count. This weight uses source and article counts. It is distinct from article fit weights, eigenvalues, and vector squared length.
For one chosen day, let be its spectrum column. The bar indicates averaging, and the sum includes that day’s stories only:
The notebooks use three synthetic story rows , , and , with source counts and article counts . Their attention weights are , , and . For the first spectrum coordinate, multiply by those weights, add, and divide by their sum. Repeating for the second coordinate gives approximately . The more heavily weighted third story pulls the mixture towards its second coordinate; no new direction has been fitted in this averaging step.
The era statistics are each spectrum coordinate’s mean and standard deviation over all loaded day spectra. The algorithm named after Welford updates three running summaries: how many values have arrived, their current mean, and the sum of squared deviations from that current mean. It adjusts the deviation sum when the mean changes; simply squaring deviations from an outdated mean would give the wrong answer. The statistical-threshold chapter walks through its calculation.
Each spectrum z-score subtracts the coordinate’s era mean and divides by its positive era standard deviation. The front page highlights values beyond (plus or minus two) and lists values below under Silence. These are comparisons with all loaded data, including dates later than some selected historical dates. They are not historical forecasts computed using only the past available at the time.
An archetype’s profile uses a different reference set: the stories it dominates in the profile window. For each named coordinate, it records a mean and standard deviation among those stories. Same as always combines strong positive profile and story signals with agreement within one profile standard deviation; agreement alone is insufficient. Different this time lists departures exceeding 1.5 profile standard deviations. The symbol in these descriptions is shorthand for the relevant reference standard deviation, not an SVD singular value. A zero reference deviation needs an explicit fallback before any z-score can be computed. Always identify the reference population before interpreting a standardized score.