Definitions are repeated here for lookup. Each discussion link returns to the chapter that develops the idea and its scope.
Absolute value. The size of a scalar without its sign. Discussion.
Accumulator. A stored partial total to which later block contributions are added. Discussion.
Anisotropy. Unequal distribution across directions. A shared embedding component can make unrelated texts point partly the same way. Discussion.
Assignment. A one-to-one pairing chosen to maximize a total similarity score, instead of choosing each match independently. Discussion.
Attention energy. A nonnegative story weight used in the spectrum or ranking. It is distinct from squared vector length and from a covariance fitting weight. Discussion.
Basis. Independent directions whose weighted combinations describe a space. Coordinates specify the weights in those combinations. Discussion.
Block. A manageable group of rows processed together; compatible accumulated sums can be combined across blocks. Discussion.
Canon. Representative high-scoring stories associated with an axis; episode deduplication prevents repeated days from dominating the collection. Discussion.
Centring. Subtracting a specified mean from every observation. The choice of mean is part of the calculation. Discussion.
Centroid. The coordinate-wise average of a group’s vectors; a centroid can hide variation among its members. Discussion.
Connected component. A maximal set of graph vertices joined by paths. Direct similarity is not required between every pair in a component. Discussion.
Coordinate. An amount along a chosen direction; changing the basis changes the coordinates even when the underlying vector is unchanged. Discussion.
Corpus. The collection of documents being analysed. Discussion.
Correlation. Covariance divided by two standard deviations, when both are positive; it compares linear variation on a common scale. Discussion.
Cosine similarity. The dot product divided by the two vector lengths. It measures directional agreement and is undefined for a zero vector. Discussion.
Covariance. The average product of two coordinates’ centred deviations; it describes how those coordinates vary together. Discussion.
Covariance matrix. The square table of all coordinate-pair covariances, calculated using specified observations and weights. Diagonal entries are variances. Discussion.
Damping. Reducing how quickly a contribution grows; the logarithmic term-frequency rule gives diminishing increments for repeated words. Discussion.
Dense storage. Allocating a place for every entry of a matrix, including zeros. Discussion.
Derivative. The limiting rate of output change for a small input change. It helps locate a minimum but needs a separate argument that the point is a minimum. Discussion.
Diagonal matrix. A matrix whose entries are zero wherever row and column indices differ. Usually square here, but the full SVD also uses a rectangular diagonal factor. Discussion.
Dimension. The number of independent directions needed to describe a space; in a table’s shape, the word also refers to its row or column count. Discussion.
Document frequency. The number of documents containing a term, regardless of its number of occurrences in each document. Discussion.
Dominant axis. The named direction with the largest signed scale-adjusted coordinate under the stated rule. Taking absolute values would give a different rule. Discussion.
Dot product. Multiply corresponding vector entries and add the products. The vectors must have the same number of entries. Discussion.
Eigengap. Separation between an eigenvalue or eigenvalue group and competing eigenvalues outside it; small gaps make individual directions less stable. Discussion.
Eigenspace. The space of vectors associated with one eigenvalue, including zero; a repeated eigenvalue can allow several independent directions. Discussion.
Eigenvalue. The scalar by which a matrix stretches an associated eigenvector. Covariance eigenvalues measure variance along unit eigenvectors. Discussion.
Eigenvector. A nonzero vector that a square matrix sends to a scalar multiple of itself. Its sign is arbitrary; repeated eigenvalues permit changes of direction within the space sharing that eigenvalue. Discussion.
Embedding. A vector computed by a trained model to represent text numerically, often allowing related meanings to have similar directions. Discussion.
Energy. A context-dependent quantity. Squared Euclidean length is geometric energy; story attention energy is a separate weighting convention. Discussion.
Episode. A group linking related stories across days under a stated similarity and time rule. Discussion.
Euclidean norm. A vector’s length: the square root of the sum of its squared entries. Discussion.
Explained variance. Variation captured by retained covariance directions; its fraction is their eigenvalue sum divided by the total eigenvalue sum. Discussion.
Fit weight. A nonnegative amount of influence assigned to an observation when calculating a mean and covariance. Discussion.
Frobenius norm. The square root of the sum of all squared matrix entries. It measures the total size of a matrix error. Discussion.
Function. A rule assigning an output to an input; for example a squared-error function assigns a cost to each proposed fitted value. Discussion.
Gaussian distribution. The symmetric bell-shaped probability model used to interpret illustrative standard-deviation cutoffs; real news scores need not obey it. Discussion.
Gram matrix. A matrix of pairwise dot products, formed by multiplying a table’s transpose by the table when comparing its columns. Discussion.
Gram–Schmidt. Making independent input directions perpendicular by subtracting their projections onto earlier unit directions, then normalizing the remainders. Discussion.
Graph. Vertices connected by edges. Here articles or stories are vertices and a similarity rule determines edges. Discussion.
Greedy matching. Taking the best currently available pair at each step. It can miss the best total one-to-one assignment. Discussion.
Identity matrix. The square matrix with diagonal ones and other entries zero; multiplication by it leaves compatible vectors unchanged. Discussion.
Inverse. A matrix operation that undoes another square matrix operation when such an undoing exists; a zero diagonal scale cannot be inverted. Discussion.
Latent semantic analysis. A text representation obtained by retaining leading singular directions of a document–term table. Discussion.
Least squares. Choosing an allowed fit to minimize the sum of squared residual entries. Discussion.
Linear combination. A sum of vectors multiplied by scalar weights. Discussion.
Loading. A number associating an input coordinate or vocabulary term with a direction. Its scaling and centring must be specified. Discussion.
Logarithm. The inverse of exponentiation; the natural logarithm uses base e and is defined for positive real inputs. Discussion.
Marginal scaling. Dividing each coordinate by its own standard deviation. This fixes each variance but does not generally remove correlations. Discussion.
Matrix. A rectangular table of numbers; its shape records rows first, then columns. Discussion.
Mean. A coordinate-wise average. For a weighted mean, multiply by nonnegative weights, add, and divide by their positive total. Discussion.
Moment. An average of powers or products of values. Centred moments use deviations from a specified mean. Discussion.
Naming rotation. An orthogonal change of coordinates within the retained space, chosen to make terms concentrate on interpretable directions. Discussion.
Norm. A rule measuring size; here Euclidean vector length, Frobenius matrix length, or an explicitly identified matrix operator norm. Discussion.
Novelty ratio. Residual squared length divided by total centred squared length. It measures unexplained geometric fraction, not social importance or a probability of novelty. Discussion.
Operator norm. The largest matrix output length over unit-length inputs; it measures the largest possible amplification. Discussion.
Orthogonal. Perpendicular: two vectors have zero dot product. A square orthogonal matrix preserves lengths. Discussion.
Orthonormal. Mutually perpendicular and individually of unit length. Discussion.
Orthonormalization. Replacing spanning directions by mutually perpendicular unit directions spanning the same space, discarding dependent inputs. Discussion.
Outer product. A column times a row, making a matrix of all pairwise products of their entries. Discussion.
Overnight matching. Associating newly fitted directions with previous directions so that labels remain useful across updates. Discussion.
Oversampling. Using additional working directions beyond the desired final rank in a randomized decomposition. Discussion.
Percentile. A cutoff in an ordered collection or distribution. It describes a proportion of values, not the probability that an individual match is correct. Discussion.
Perturbation. A change to a matrix or other object; stability asks how much an output can change in response. Discussion.
Positive semidefinite. A symmetric matrix for which every quadratic form is nonnegative; a covariance matrix has this property. Discussion.
Power iteration. Repeated matrix products that amplify larger singular directions relative to smaller ones, with normalization for numerical stability. Discussion.
Power sum. A sum of values raised to a specified power; the first three sums can recover a third centred moment. Discussion.
Precedent. An earlier story selected for comparison by similarity in a specified representation. Discussion.
Precision. Correct accepted links divided by all accepted links; it needs a convention if none are accepted. Discussion.
Principal component. A covariance eigendirection, or an observation’s score along it when referring to component scores; leading directions capture the largest fitted variance. Discussion.
Procrustes alignment. An orthogonal rotation or reflection used to compare bases describing nearly the same subspace even if individual columns differ. Discussion.
Profile. An archetype’s per-axis means and standard deviations over the stories it dominates in its reference window. Discussion.
Projection. Keeping the component of a vector inside a chosen subspace. Orthogonal projection leaves a perpendicular residual. Discussion.
Quadratic form. A scalar made by multiplying a row vector, a square matrix, and the matching column; it generalizes a weighted sum of squared coordinates. Discussion.
Randomized SVD. An approximation that first finds a smaller space with random test directions, then decomposes the data inside that space. Discussion.
Range. The set of outputs a matrix can produce from all compatible input vectors. Discussion.
Rank. The number of independent directions in a matrix; retaining fewer singular directions produces a lower-rank approximation. Discussion.
Recall. Correct accepted links divided by all true links in a specified evaluated set; it needs a convention if there are no true links. Discussion.
Reconstruction. An estimate made by combining retained directions with their coordinates and restoring the mean if it was removed. Discussion.
Relative error. Residual norm divided by the positive original norm; it compares geometric sizes, not the fraction of correct articles. Discussion.
Residual. The difference between an observation and its reconstruction. It is a vector before its length or squared length is measured. Discussion.
Rotation invariance. A quantity remains unchanged when coordinates are rotated consistently, including its covariance or metric where needed. Discussion.
Score. A coordinate along a fitted direction, or a ranking value when explicitly identified as such. Discussion.
Shape. The row and column counts of a matrix, or length of a vector; compatible shapes are required for multiplication. Discussion.
Single linkage. Clustering by connected components of the graph of sufficiently similar pairs; chains can connect unlike endpoints. Discussion.
Singular value. A nonnegative scale in a singular value decomposition, usually ordered largest first. Discussion.
Singular value decomposition. Factoring a matrix into orthogonal directions, separate nonnegative scales, and orthogonal directions on the other side. Discussion.
Sketch. A smaller collection of matrix-output vectors used to approximate the important output space. Discussion.
Skewness. Asymmetry of a distribution, measured here through a standardized third centred moment when variance is positive. Discussion.
Sparse storage. Storing nonzero values and their positions rather than every cell of a mostly zero matrix. Discussion.
Spectrum. In the newspaper interface, a day’s attention-weighted vector of named coordinates; in eigenvalue analysis, a collection of eigenvalues. Context identifies the meaning. Discussion.
Standard deviation. The nonnegative square root of variance, returning squared deviations to the original coordinate scale. Discussion.
Standardization. Subtracting a reference mean and dividing by its positive standard deviation. Discussion.
Story. A cluster of related articles under the chosen within-day rule. Discussion.
Streaming. Processing manageable blocks while carrying forward sufficient intermediate quantities instead of retaining every input row. Discussion.
Subspace. A collection containing all linear combinations of selected directions, including zero. Discussion.
SVD. Abbreviation for singular value decomposition. Discussion.
Term frequency. A term’s count within a document, optionally damped by the logarithmic rule specified here. Discussion.
TF–IDF. Term frequency multiplied by inverse document frequency, followed here by row normalization. Discussion.
Threshold. A chosen decision boundary. Its meaning depends on the representation, population, and inequality used. Discussion.
Trace. The sum of a square matrix’s diagonal entries. The covariance trace is total variance. Discussion.
Transpose. Exchanging rows and columns, denoted by a superscript top symbol. Discussion.
Truncation. Discarding later singular or eigen directions while retaining the leading ones. Discussion.
T² statistic. Hotelling’s variance-scaled squared distance in retained coordinates. It uses positive retained variances or an appropriate full covariance inverse. Discussion.
Uncorrelated. Having covariance zero under the specified reference distribution; this does not generally imply statistical independence. Discussion.
Unit vector. A vector with Euclidean length one. Discussion.
Variance. An average squared deviation from a mean; the denominator and weights depend on whether describing a fitted population or estimating from a sample. Discussion.
Varimax. An orthogonal rotation criterion that encourages uneven squared loadings within each direction, supporting simpler term-based names. Discussion.
Vector. An ordered list of numbers; the order determines which coordinate each entry represents. Discussion.
Weighted average. A sum of values multiplied by nonnegative weights, divided by a positive weight total. Discussion.
Welford update. An incremental running mean and sum of squared deviations, avoiding subtraction of two large nearly equal sums. Discussion.
Whitening. A transformation giving fitted coordinates identity covariance; unlike marginal scaling, it also removes pairwise linear correlations. Discussion.
Z-score. A deviation from a reference mean divided by its positive standard deviation; it is a standardized coordinate, not itself a probability. Discussion.