A Worked Companion to Eigen Times Math History
4 October 2026 | Edition 1.1.0-99aa176a
Eigen Times History Math teaches techniques that the earlier Mathematics of Eigen Times and Mathematics of Anthropology do not yet develop as complete executable lessons. Its parent, Eigen Times Math History, explains the historical arc and original publications. Here the emphasis is on learning how to perform a calculation, understanding why it works, and discovering when its answer should not be trusted.
The five editions share one text: PDF, EPUB, HTML, a Python notebook, and a native OCaml notebook. The notebooks contain the explanations, definitions, examples, exercises, glossary, references, and index. Code cells follow the relevant explanations. Each numerical example starts from declared inputs and is computed in both languages. Saved outputs are a record of a checked run; changing an input and running the cells again is how to investigate a new question.
This is a complement to the other two companions. A covariance calculation is not repeated merely because it occurs in another historical chapter. Conversely, calling a library routine is not the same as explaining the algorithm inside it. For example, the earlier OCaml setup already diagonalizes matrices and orthogonalizes columns. The new lessons expose the rotations, projections, reflections, and conditioning questions that a learner could otherwise miss.
Begin with the Eigen Times Math reader if vectors, matrix multiplication, covariance, or singular values are unfamiliar. Principal component analysis (PCA) finds orthogonal directions of variation in centered observations; its construction is taught there. Use the Anthropology Math reader for person profiles, people directions, reciprocal navigation, and proposed ontology predictors. The prerequisite catalog names exact lessons and the Python and OCaml cells that compute them. The delivered study bundle includes unchanged copies and reading versions of those four notebooks from their frozen 1.0.0 baselines, preserving the exact lesson and cell references used by this prerequisite catalog. This companion is the expanded 1.1.0 edition; separately delivered expanded companions provide their current teaching. The baseline citations are not silently repinned to changed cells. The public study repository provides a study landing page; the local bundle does not require access to a private repository.
The new companion has two connected routes. The numerical route moves from fitting observations to reliable matrix calculations, graphs, and retrieval. The information and learning route moves from uncertainty and forecast quality to formal concepts, derivatives, predictive embeddings, attention, and calibration. Read each route in order on a first pass. Later, the subject index can take you directly to a method.
A useful teaching rhythm is to read the definitions, predict the output, run the cell, inspect the invariant being tested, and change one input. An invariant is a property that the stated operation should preserve, such as an orthogonal transformation preserving squared length. A reference assertion about a fixed toy answer may need changing after a deliberate edit; an invariant should still hold whenever its assumptions hold. Cross-language agreement is an additional check, not a replacement for such mathematical reasoning.
All new arrays and graphs in this book are synthetic. They are small enough to inspect and require no credentials, newspaper archive, model download, or live service. A two-layer derivative exercise does not train BERT. A scaled-attention calculation does not reproduce the BGE embedding model. A limited triple-rule closure is not an implementation of an OWL reasoner. Quantizing a handful of numbers illustrates rounding and error; it does not reproduce a production model’s complete quantization procedure.
These distinctions matter because the three sites combine mathematics with choices about data, identities, timestamps, and evidence. A cosine or learned representation can suggest related coverage. It cannot establish an employment relationship, an investment, or the reliability of a source. Likewise, a score between zero and one is not automatically a calibrated probability. The new lessons explain the computations while keeping those evidential boundaries.
Numerical examples are computed with ordinary finite-precision arithmetic. Rounding error is the discrepancy introduced when a mathematical number or operation cannot be represented exactly by the machine. Approximation error is the loss intentionally accepted by a model or truncation. A stable algorithm can reduce the first while leaving the second unchanged. Neither is the same as an incorrect identity match or an unsupported factual claim.
The notebooks execute without network access after their runtimes and packages are installed. Optional reading links and the browser’s mathematics renderer may need a connection. Use a fresh kernel and run all cells from the beginning; later lessons can reuse helpers defined earlier. The setup code is a toolbox, not a prerequisite vocabulary test: its operations are introduced in the corresponding lessons. Python uses NumPy where a library operation helps make a comparison; OCaml performs the calculations natively.
An observation is one measured item. A scalar is one number. A vector is an ordered list of numbers, and a matrix is a rectangular table of them. Observation vectors are rows in this companion; projection directions are columns. A matrix with rows and columns has shape , where and here are positive integer dimensions. Multiplying a matrix by a matrix gives a matrix, with another positive dimension. Each product entry is a sum of coordinatewise products.
The symbol denotes the real numbers. Superscript means transpose, which exchanges rows and columns. For a vector , its Euclidean norm is the square root of the sum of squared coordinates. For a matrix , the Frobenius norm is the square root of the sum of squared entries. Its operator norm is the largest length of over column vectors of length one. The identity matrix leaves every -coordinate column unchanged; is a column of ones. The symbol forms a diagonal matrix from the listed numbers, and sums a square matrix’s diagonal entries.
A linear combination multiplies vectors by scalars and adds the results. A span is the set of all such combinations. A list is linearly independent when none of its vectors is a combination of the others. A basis is an independent list spanning a space. Its number of members is the space’s dimension. A matrix’s rank is the dimension of its column span, equal to the dimension of its row span. A nullspace consists of inputs the matrix maps to zero. These ideas explain both uniqueness and ambiguity in fitting.
The dot product of two equally sized vectors is the sum of their coordinatewise products. Two vectors are orthogonal when their dot product is zero. A set is orthonormal when its members are orthogonal and have length one. An orthogonal matrix is a square matrix whose transpose is its inverse. An eigenvector of a square matrix is a nonzero column whose image is a scalar multiple of that column; the multiplier is its eigenvalue. The earlier Eigen Times companion calculates these objects; the new lessons explain more of the numerical machinery for constructing them.
The natural logarithm is written or ; has base two. The exponential function is , with . A sum sign adds terms over its stated index. A hat denotes an estimate or reconstruction, and a bar denotes a mean. The notation asks which argument makes an objective smallest. A derivative measures local change of a function with respect to one input; the neural lesson develops its use from an explicit calculation. A gradient is the vector of derivatives with respect to several inputs. Local auxiliary symbols are defined before the formula using them.
A subscript names an entry rather than multiplying: is entry of a residual column . If the column has entries, the notation instructs us to square each entry and add all results. The squared Euclidean length is this sum; the length itself requires one more operation, the square root:
A residual can have negative entries even though its length cannot be negative. The sign records whether an individual prediction is above or below an observation; squaring removes that sign when measuring total error. Squared length and length have different units. If residual entries are measured in a unit of response, length has that unit and squared length has its square. Dividing squared length by instead gives mean squared error. Taking its square root gives root mean squared error. These quantities answer related questions but must not be exchanged under the single word “error.”
Products are read from their shapes. If has observation rows and predictor columns and is a coefficient column, then is an prediction column. Its th entry is , where temporarily indexes predictors. The transpose is , so contains one dot product with each predictor column. Reading those entries explicitly will explain the least-squares equations, rather than asking us to memorize a matrix formula.
An equality sign says two expressions have the same value. The sign marks an approximation, often a rounded display. An inequality such as includes equality. Brackets describe an interval including its endpoints; set braces describe a collection of elements. Set inclusion means every element of belongs to . A prime in the formal-concept lesson is a named set operation, while a prime on a scalar function in the fitting lesson means a derivative; each lesson states its local use.
When we differentiate a scalar function, we ask how much its output changes per small change of input. For a function of a scalar , its derivative is the limit of as the nonzero displacement approaches zero, when that limit exists. The symbol denotes a partial derivative: change one input while holding the others fixed. The fitting lesson develops this idea first for a square, and the network lesson builds a chain of such local changes.
The following symbols retain the meanings used by the history paper. New local algorithms declare their own temporary arrays explicitly. A locally subscripted neural quantity must not silently become a news or people matrix.
| Symbol | Meaning in the series |
|---|---|
| Observation count and observation index. | |
| Input-feature count and feature index. | |
| Vocabulary size and term index in text calculations. | |
| Article/input rows, reference mean row, centered rows. | |
| Covariance matrix, never the singular-value matrix. | |
| News-direction columns, their eigenvalues, retained direction count. | |
| Raw news scores, naming rotation, rotated scores. | |
| Positive marginal scale matrix and standardized news coordinates. | |
| A locally declared matrix and its singular-value factors. | |
| Rank or retained rank, stated locally. | |
| Design matrix, response column, fitted coefficient column in least squares. | |
| Nonnegative regularization penalty; distinct from eigenvalues. | |
| People count, person index, and person-article incidence in prerequisite lessons. | |
| Person profiles, people-direction loadings, people scores in prerequisites. | |
| Residual-energy and score-distance statistics in the earlier companions. | |
| A probability, with its event or outcome declared locally. |
A formula copied from a source using column observations may require a transpose before it can be compared with this book. The notation is a teaching convention, not a claim that every cited paper or production source file used these letters.
The earlier companions already calculate covariance, principal components, singular-value decompositions, and people profiles. The lessons here open the operations between those results. They use small disclosed inputs so that an answer can be checked by a second identity, not merely accepted because a library printed it. Observations remain rows. A coefficient vector or a column extracted while factoring a matrix is explicitly a column; that does not change the observation convention.
Least squares starts with a choice about error. Adrien-Marie Legendre’s 1805 formulation supplies the historical reference; the small examples below are modern teaching fixtures, not reconstructions of his astronomical observations. Legendre 1805
Let the three observed responses be , , and . Let index these observations, let be one fitted scalar, and let denote the sum of squared discrepancies. A completed-square form writes a quadratic as a nonnegative squared term plus a constant, making its minimum visible. Here it is
The minimum is therefore at two, and the remaining squared error is two. The notebook checks the identity at several candidate values. Squaring prevents opposite discrepancies from canceling. It also makes a large discrepancy more influential, so the criterion is a modeling decision rather than a universal meaning of accuracy.
Here the model predicts the same number for every response. At the candidate value two, subtract prediction from observation to obtain residuals . Their sum is zero, although two observations are missed. Their squared length is and their length is . A zero sum of signed errors therefore does not mean a perfect fit.
To obtain the completed square, expand each term:
The square is nonnegative and becomes zero exactly at two. This establishes the minimum without calculus. Differentiation reaches the same conclusion, but it is worth seeing what that operation does. At a candidate , change the parameter by a nonzero displacement . Subtract the old loss from the new one and divide by the displacement:
As tends to zero, the last term vanishes, leaving derivative . Below two this rate is negative, so increasing the prediction a little decreases loss. Above two it is positive. The derivative is zero at the bottom of the bowl. A zero derivative alone does not generally guarantee a minimum; the completed square supplies that guarantee here. The lab checks both the expanded polynomial and the derivative’s finite change formula against the original sum of residual squares.
For a linear fit, define a predictor matrix with observation rows and columns, a response column with entries, and a coefficient column with entries. Define prediction column and residual column . Let index a coefficient. Its residual entry is . The squared residual length is .
Hold all coefficients except fixed. Increasing that coefficient by changes residual by . The minus sign comes from subtracting the prediction. Expanding the changed square gives . Summing, subtracting the old loss, dividing by , and letting it approach zero gives
At a minimum, every one of these partial derivatives is zero. Each sum is an entry of . Substituting and moving terms gives the normal equations
The second equality says that the residual is perpendicular to every predictor column. It is a useful independent check. In the fixture, has rows and the response column is . The two fitted coefficients are and , with squared residual . Independent predictor columns make this minimizer unique. Dependent columns need a different discussion, which follows in the next lesson.
The first column of this design is all ones; its coefficient is the intercept, the fitted response when the second predictor is zero. The second coefficient is the slope, the prediction change per unit change of that predictor. The products can be calculated before solving:
Thus and . Subtracting the first equation from the second gives . Substitution gives . The fitted responses are , so the residual is . Squaring and adding gives . Its length is ; its mean squared error is . The residual’s dot product with the all-ones column is , and with the second predictor it is .
Why does perpendicularity minimize rather than merely stop the loss? Let be the fitted column and let be any -entry coefficient change. The new residual is . Expanding its squared length gives . The middle term is zero and the last term cannot be negative. Independent columns make it positive for every nonzero , establishing uniqueness.
Uniqueness does not guarantee stability. Let and be two centered columns, and set . Assemble and as the two columns of a new predictor matrix . Its columns are nearly the same. Fitting response gives coefficients ; changing the response to gives . A small response change causes a large coefficient change. Their fitted responses can nevertheless remain close.
The Anthropology companion’s ridge lesson already defines a squared-coefficient penalty. Let be the two-coordinate identity matrix. Here a fixed penalty demonstrates its motivation: repeat both fits with . These centered predictors have no fitted intercept in this comparison. The coefficient change is much smaller, although penalization changes the objective and introduces bias. Stability is not a substitute for held-out evaluation. Hoerl and Kennard 1970
Exercise. Why does the tiny perturbation not imply tiny coefficient changes? The two predictor columns nearly describe the same direction. Large, opposing coefficient adjustments can nearly cancel in the fitted response. Inspect both the response-change norm and the coefficient-change norm; the notebook computes them separately.
The word inverse can hide two different tasks: solving for a response that the model can reproduce and choosing among several coefficient vectors that reproduce it. A rank-deficient matrix can leave the first task possible while making the second ambiguous. A pseudoinverse supplies a declared choice. Gene Golub and William Kahan’s 1965 paper connects singular-value computation with this generalized inverse and least-squares applications. Golub and Kahan 1965
Let be a real matrix with rows, columns, and rank . Rank counts independent directions. Its compact singular-value decomposition uses orthonormal-column matrices of size and of size , together with a diagonal matrix of positive singular values. Define by taking the reciprocal of each positive diagonal entry. The pseudoinverse, denoted and sized , is
Only positive singular values are inverted. A thin factorization may contain additional zero singular values; those entries contribute zero to the pseudoinverse rather than an attempted division by zero. In floating-point arithmetic, deciding which small values count as zero requires a scale-aware tolerance. The example uses an exact rank deficiency and clearly separated positive scale, avoiding an ambiguous threshold experiment.
Take with rows and , and response column . Both and fit perfectly. So does every coefficient column whose entries sum to one. The nullspace consists of coefficient changes proportional to : multiplying such a change by gives zero. Denote the chosen coefficient column by . The pseudoinverse gives . Its squared length is , smaller than the squared length one of either endpoint solution.
Now change the response to . It no longer lies exactly along the single output direction. The shortest least-squares coefficient column is , giving fitted response and squared residual . The residual remains perpendicular to the columns of . “Shortest” describes the coefficient norm among all least-squares minimizers; it does not mean the smallest prediction error has become zero.
Write the sum of the two coefficients as . Multiplication by gives : the matrix can see the sum but cannot distinguish how it was divided between coefficients. For the exact response, . Parameterize all its solutions as , where is any real scalar. Their squared length is
It is smallest at . This is the geometric reason for equal coefficients; the pseudoinverse formula performs the same choice in general dimensions.
For the changed response the residual is . Its squared length expands to . The best sum is therefore and splitting it equally gives each coefficient . The actual residual is , with squared length . No choice of split can reduce that residual; adding a nullspace vector changes coefficients while leaving the fitted response fixed. Conversely, “shortest” does not mean choosing the zero coefficients, since their prediction fails the first requirement of having minimum residual error.
The laboratory constructs the inverse from the factors and checks four identities: , , and symmetry of both and . The two products project into different spaces and have different roles. Anthropology’s transpose-based return is a special case with orthonormal loading columns, not a general replacement for this construction.
Exercise. Double the first response column. Linearity doubles the chosen coefficients and quadruples their squared length. The notebook checks both changes. The result illustrates a mathematical convention for underdetermined fitting; it does not identify missing social evidence or reconstruct a discarded person attribute.
The existing notebooks use orthonormalization inside larger calculations. This lesson makes its construction visible. Jørgen Pedersen Gram and Erhard Schmidt supply the historical references; Schmidt explicitly credited Gram. We use finite matrices rather than reproducing their function-space settings. Gram 1883; Schmidt 1907
Let be a matrix with rows and linearly independent columns, where . Its observations are rows, but the algorithm constructs an orthogonal frame for its column space. Call its th column . Let denote an already constructed unit column, and let be the part of left after removing its projections onto earlier columns. In exact arithmetic the construction is
Independence ensures a nonzero denominator. Assemble the columns into , an orthonormal-column matrix. Define the matrix . Its entries below the diagonal are zero because each new input column has no component along a future constructed direction. Thus . The name QR describes these two factors; it does not identify this factorization with the residual statistic called elsewhere.
Take the columns and . The first unit column is . The coefficient removed from the second is , leaving , whose squared length is . The notebook prints these intermediate quantities, the final factors, and checks both orthonormality and reconstruction. Checking only one of those identities would be insufficient: a set of unit columns might reconstruct the wrong matrix, and a reconstruction could use nonorthogonal columns.
A dot product with a unit column returns a scalar coordinate along that column. Multiplying the coordinate by the unit column puts that component back in the original space. Here , so the vector removed is . Subtracting it from gives the stated remainder. Its squared length is , and its length is . Dividing the remainder by this positive length gives .
The dot product checks that the removed component has not been left behind. The decomposition explains the second column of : it stores reconstruction coefficients rather than another set of observation rows. The original squared length two splits into projection energy plus remainder energy . This is the Pythagorean identity for perpendicular components. It explains why residual energy measures what the selected direction failed to represent.
The executable procedure updates each remainder immediately after a projection. This is the modified Gram-Schmidt ordering. It repeats the projections once to reduce lost orthogonality from rounding. Exact formulas that agree algebraically need not have the same numerical behavior, so the order and the second pass are stated openly. The small tolerance detects a numerically empty remainder relative to the input matrix; it is not a general-purpose rank estimator for arbitrary data scales.
QR also gives a route to least squares without first forming a Gram matrix. For a response column with entries, compute the response coordinates , then solve the upper-triangular system with . The lab compares this answer with normal equations on the same well-conditioned fixture. Agreement here does not prove equal stability on a difficult matrix.
Exercise. Replace the second column by the first. The remainder vanishes, so there is no second direction to normalize. The code rejects that full-column QR request. A rank-revealing algorithm could return a smaller factorization instead; silently dividing by a tiny remainder would conceal the change in the problem.
Subtracting one projection at a time is not the only way to create zeros. A Householder transformation reflects an entire vector onto a coordinate direction while preserving lengths. Golub and Christian Reinsch’s published SVD procedures use such transformations as part of a larger numerical algorithm. The fixture here teaches the reflection and QR construction, not a complete replacement for their implementation. Golub and Reinsch 1970
Let be a nonzero column with entries, and let be the first coordinate column of length . Define to be one for a nonnegative first entry and minus one otherwise. Set the scalar and define the column . Let be the identity and define the reflection matrix by
The denominator is the squared length of a nonzero column. The sign choice makes the first component of add magnitudes rather than subtract nearly equal quantities. Because and , the matrix is symmetric and orthogonal. It sends to . A zero input needs no change, and the implementation returns the identity in that case.
For , the norm is five and the reflected column is . The notebook computes the matrix from the formula and checks the mapped vector, symmetry, and orthogonality. The minus sign is useful numerical bookkeeping, not a negative length. Applying the reflection twice returns the original vector.
The length calculation is . The sign rule gives , so and . Multiplying out the reflection gives
The first output entry is ; the second is . More generally, split a column into its component along and a component perpendicular to . The fraction extracts the first component. Subtracting twice that component reverses it while preserving the perpendicular part. Reversing a component preserves its squared length, which explains the word reflection and the orthogonality check.
To factor a matrix, embed each reflection in a larger identity so that already processed rows remain fixed. Let be the matrix whose rows are . Repeated left reflections make entries below the diagonal zero. Their accumulated product is a orthogonal matrix , and the resulting upper-trapezoidal matrix is . They satisfy . The full orthogonal factor has four columns, unlike the thin factor in the preceding lesson; the last row of the rectangular triangular factor is zero.
The checks verify the zeros, the factorization, and preservation of the Frobenius norm, which is the square root of the sum of squared matrix entries. Multiplication order matters: the sequence used to reduce the matrix is reversed by transposition when reconstructing it. Printing only a triangular-looking matrix would not test this bookkeeping.
Exercise. Use . The chosen first output becomes positive five, with the other entries zero. The lab checks this separately. Both choices represent the same length-preserving construction; neither assigns a stable semantic sign to a news direction.
The Gram identity is exact: squared singular values of a matrix are eigenvalues of its column-product matrix. That identity does not imply that forming the product is always a good numerical route. Golub and Kahan’s 1965 work and Golub and Reinsch’s 1970 procedures explain the importance of reducing the original matrix directly. Golub and Kahan 1965; Golub and Reinsch 1970
Let now be a matrix with rows and columns, with . Define as an orthogonal matrix and as an orthogonal matrix. We seek a matrix whose only potentially nonzero entries lie on its main diagonal and the diagonal immediately above it. Such a matrix is called upper bidiagonal. The reduction satisfies
First use a left reflection to remove entries below the current diagonal position. Then use a right reflection to remove entries farther right than the next position in that row. Embed each reflection so earlier zeros remain untouched. Repeating this sequence on the preceding lesson’s matrix yields a bidiagonal result. The notebook checks the zero pattern, both orthogonal factors, exact reconstruction within rounding tolerance, and preserved total squared size. It does not claim that reaching bidiagonal form completes a singular-value decomposition. Diagonalizing that reduced matrix requires an additional algorithm.
Why avoid an unnecessary squaring? Set , and let have rows and . Define and as its larger and smaller singular values. Define , the sum of their squares. The matrix maps a unit square to a parallelogram of area , also the product of its two singular stretches. A stable analytic calculation for this particular fixture is
The small value is approximately . Computing it from the product relation avoids subtracting nearly equal quantities. This formula is a special two-dimensional reference, not a general direct-SVD implementation. Python additionally compares it with its numerical SVD; OCaml computes the reference natively without using its Gram-based helper as the authority.
The spectral condition number, denoted , is the ratio for this nonsingular matrix. The exact Gram matrix has condition number . In binary64, the usual 64-bit floating-point format, its entry rounds to one for this fixture, erasing the small distinction before an eigensolver sees it. A sophisticated eigensolver cannot recover information already lost in its input.
The column-product matrix in this example is
Its trace is and its determinant is . The determinant of a two-by-two matrix with entries is ; here it equals the product of the two squared singular values. Solving the resulting quadratic gives the large squared singular value in the displayed formula. Instead of obtaining the small squared value by subtracting that large root from the trace, use the product relation to recover the small positive singular value.
Conditioning is a property of the problem: some input directions are stretched far more than others, so reconstructing their inputs can amplify small errors. Stability is a property of a computational procedure: how much extra error its rounding introduces. The Gram route both squares the ratio of scales and rounds the product entries before solving. The positive input entry still exists in the original matrix even when has rounded to one in the product. Working from the original matrix gives a direct algorithm a chance to retain that information; it does not promise exact arithmetic.
Exercise. Is the native Gram-based teaching SVD evidence that a direct SVD also loses this direction? No. It is evidence about that particular computational route. The notebook displays the positive analytic small value beside the rounded Gram eigenvalue and keeps those two claims separate. Production solver selection also requires documented precision, scale, and rank tolerances.
An eigensolver can be an inspectable sequence of geometric changes rather than one opaque call. Carl Gustav Jacob Jacobi’s 1846 paper is the historical point of reference. The following modern plane-rotation algorithm exposes the mechanism already used inside the native notebook helper. It is not a transcription of the original paper’s notation. Jacobi 1846
An angle measures a turn from a reference direction. Draw a circle of radius one centered at the origin of two perpendicular coordinates; this is the unit circle. Start at its point and turn through angle , with positive angles measured counterclockwise. The horizontal coordinate of the resulting point is called , the cosine. Its vertical coordinate is , the sine. Because the point remains one unit from the origin, their squared coordinates satisfy
A radian measures angle as the traveled arc length divided by the radius. Here is the circle constant, distinct from a probability: a full turn is radians, a half-turn is , and a right angle is . The numerical functions in both notebooks take angles in radians. The rotation below uses these circle coordinates as matrix entries. The identity above makes its columns unit length, and their opposite signed cross-products cancel to make the columns perpendicular. Its displayed sign convention describes a clockwise turn when applied to a column; the paired basis change below uses that convention consistently.
Let count coordinates and let be a real symmetric working matrix. Select two different coordinate indices and . Write , , and for the entries of their symmetric two-coordinate block. Let be a rotation angle and define its block, denoted , as
Embed this block in a identity matrix and call the resulting rotation . Define the next working matrix by . The new off-diagonal entry in the selected block is . The two-argument inverse tangent returns the direction of the point in the appropriate quadrant. Choose
This avoids an ambiguous division when is zero. If , skip that block. The choice removes the selected off-diagonal entry up to rounding. The trace is the sum of the diagonal entries. An orthogonal similarity transformation preserves eigenvalues, trace, and Frobenius norm.
For with rows , one rotation yields diagonal values one and three. The accumulated rotation columns are eigenvectors, so checking the diagonal alone is not enough: the laboratory also checks the original matrix multiplied by each recovered column against that column multiplied by its eigenvalue.
In this block and , so the rotation angle is one half of a right angle. Its columns are and . Multiplying the original matrix by the first column gives , while the second gives . This verifies eigenvalues one and three without relying on a printed diagonal.
The two multiplications in have different roles: the right multiplication applies the original map to the new basis columns, and the left multiplication expresses those outputs in the new basis. A diagonal result means each new basis direction maps into itself. For the larger matrix a rotation removes only the selected pair; other entries can move. That is why the lab tracks total off-diagonal energy over successive steps and does not mistake one vanished entry for a completed diagonalization.
Next use the symmetric matrix with rows . At each step select the largest remaining off-diagonal magnitude. Define off-diagonal energy as the sum of squared entries outside the diagonal. The code records its decrease, accumulates the rotations in the correct order, and stops when the largest off-diagonal entry is small relative to the matrix norm. A maximum iteration count prevents an unnoticed infinite loop. The final trace remains nine. Both languages compare the reconstructed matrix and eigenpair residuals with independent identities.
Repeated eigenvalues explain a limit on interpretation. Let denote the two-coordinate identity. The matrix remains unchanged under any orthogonal rotation. It has no preferred pair of directions inside that plane. A deterministic numerical sign or ordering convention makes outputs reproducible, but does not create unique mathematical directions.
Exercise. Run the repeated-eigenvalue fixture and rotate its axes. The notebook verifies that the matrix remains the same. This is why an axis can move while the subspace remains stable; the next lesson measures subspaces directly rather than demanding identical columns.
Two coordinate frames can disagree column by column while describing exactly the same plane. A successful Procrustes fit demonstrates one way to align paired coordinates; principal angles answer a related but different question about the spaces themselves. Åke Björck and Gene Golub studied their numerical computation in 1973. Björck and Golub 1973
Let count input coordinates and count retained directions. Let and be two matrices with orthonormal columns. Their rows must refer to the same input coordinates, in the same units. Let index the singular values of the cross-product , arranged from largest to smallest. The inverse cosine, written , returns the angle associated with a cosine. Define the principal angles as numbers between zero and a right angle satisfying
Each singular value lies between zero and one in exact arithmetic. The implementation clips tiny floating-point excursions to that interval before taking the inverse cosine, after checking the orthonormal-input contract. Clipping is not permission to feed arbitrary unnormalized matrices into the formula.
Use three input coordinates and retain two directions. Here denotes the circle constant, not a predicted probability. Let be the three coordinate unit columns, set the tilt angle , and choose and . The planes share their first direction, giving principal angle zero. Their second directions meet at , giving a second angle of thirty degrees. The code computes the cross-product, its singular values, and both angles rather than inserting the expected answer.
Now rotate the two columns inside each plane, using angles and respectively. The individual columns change, but the principal angles remain the same. Their projectors, defined as and , also remain unchanged by internal rotations. An independent check relates their squared Frobenius separation to the principal angles:
For this fixture the squared separation is . The factor two reflects that both directions of disagreement contribute to the projector difference. This check links the angle calculation to an independently constructed matrix quantity.
Here the cross-product is diagonal with entries and : the shared direction has dot product one with itself and zero with the tilted direction; the second original direction has dot product with that tilt. At , the latter is . Taking inverse cosines recovers angles zero and . A radian measures angle as arc length divided by radius; radians is a half-turn, making thirty degrees.
The projector maps any input column to its nearest column in the specified plane. Internal rotations change how coordinates are written inside that plane, but cannot change this projection. In the final two coordinates, the difference has entries . Squaring all four and adding gives , agreeing with . The projector calculation is independent of how an SVD routine chooses signs for its columns.
Exercise. Change first to zero and then to a right angle. The angle pairs become zero/zero and zero/right-angle. The notebook runs both cases. Identical spaces need not have identical displayed axes, while one shared direction does not imply that all retained directions agree.
This calculation compares geometry under a shared feature description. It cannot directly compare a term-loading matrix with an embedding-loading matrix whose row meanings differ. It also says nothing about whether the articles, people, or source populations used in the two fits were comparable. Those prerequisites need evidence before a small angle can support a continuity claim.
An archive can be summarized in blocks, but its variance is not simply the sum of the variances inside those blocks. The distances between their means matter too. Welford’s corrected-sum update and Tony Chan, Golub, and Randall LeVeque’s later analysis provide historical context for algorithms that keep deviations visible. Welford 1962; Chan et al. 1983
For a block of scalar observations, let be its positive count, its scalar mean, and its corrected sum of squares, the sum of over its values . This lesson uses unit weights. For two blocks labeled A and B, write their summaries as and . Define and the scalar mean difference . Their merged summary is
The final term is the variation between the block means. It is nonnegative and vanishes when those means agree. It is not an optional correction for a particular division of the data: omitting it changes the quantity being calculated. Empty blocks are handled by returning the other summary before applying the formula, because an empty block has no defined sample mean.
To see where that term comes from, take an observation in block A and write . Squaring produces the squared within-block deviation, twice the product of the two terms, and the squared mean shift. When summed over block A, the product term vanishes because . The block therefore contributes . Repeat for block B. Substituting and reduces the two mean-shift contributions to .
This derivation explains what the stored summary must retain: count tells us how heavily a block’s mean matters, the mean locates the block, and the corrected sum records spread around that location. Keeping only a variance would lose the count and location needed to combine blocks.
For the first block , the count is two, mean , and corrected sum . For the second block , the count is one, mean seven, and corrected sum zero. The mean difference is , and the between-block contribution is . The merged mean is four and corrected sum fourteen. Dividing by three gives the population-normalized variance; dividing by two gives sample variance seven. The choice of divisor is a reporting convention after the same centered accumulation.
The within-block arithmetic is and . The additional term is , making the merged sum . Check directly around the merged mean four: . Thus almost all the variation in this fixture is between blocks, exactly the part a sum of within-block variances would omit. The paired lab checks this direct route against summary merging.
Next add to every observation. In exact arithmetic this translation cannot change any deviation from the mean. Binary64 arithmetic can nevertheless lose the answer when it computes a raw sum of squared values and then subtracts count times squared mean. Both terms are enormous compared with fourteen. The lab displays that naive result, then checks a shifted calculation, Welford updates, and centered block merging against the original corrected sum. The exact rounded failure value is implementation-dependent, so it is not treated as a cross-language reference constant.
The existing companion’s raw weighted-moment accumulation teaches a useful algebraic identity and memory strategy. This lesson adds its numerical boundary and a centered alternative. A method that is more stable for a fixed population does not repair changing weights: if adding articles changes a day’s normalized fitting weights, the summaries must reflect the revised weights too.
Exercise. Merge and . Both means are two, so the between-block term is zero; their corrected sums add to ten. The code verifies this case and the empty-block identity. The exercise isolates why the correction is sometimes invisible in a convenient fixture but necessary in the general formula.
The news companion already shows why a chain of article links can join endpoints that are not themselves similar. A minimum spanning tree supplies a smaller representation of the same threshold connectivity. Florek and colleagues’ early connecting-tree work and Gower and Ross’s later single-linkage result provide the historical setting. The inspected story code uses pairwise comparisons and union-find, which maintains component identities as edges join groups; it does not construct a tree. Florek et al. 1951; Gower and Ross 1969
A graph has vertices and edges; a complete graph contains an edge for every pair of distinct vertices. A path is a sequence of edges meeting at successive vertices. A cycle is a closed path with no repeated intermediate vertex. A connected component consists of all vertices joined by paths. A spanning tree connects all vertices without a cycle, and a minimum spanning tree has the smallest total edge cost among such trees. Let count vertices, let and index them, and let denote a nonnegative symmetric edge cost between distinct vertices. This subscripted cost is unrelated to the input dimension . Let be a distance threshold. The full threshold graph keeps every edge whose cost is at most .
The equivalence says that removing from a minimum spanning tree every edge with cost greater than gives exactly the same connected components. It does not say the two graphs have the same edges. A complete graph can retain many redundant routes. If a tree path needed an edge more expensive than an available route joining its two sides, substituting a cheaper crossing edge would contradict the tree’s minimal total cost. This exchange argument explains why the tree preserves the relevant threshold connections.
Use four points at positions on a line. Costs are absolute differences. The six possible edge costs are , all distinct. A simple greedy construction sorts the edges and accepts one only if its endpoints are currently in different components. The accepted costs are , with total seven. The program computes these choices and verifies that exactly three edges connect all four points.
Initially each vertex is its own component. The edge between positions zero and one costs one and joins those two components. The edge between one and three costs two and joins position three to them. The edge between zero and three costs three, but both endpoints already have a path through position one; accepting it would create a cycle. Skip it. The edge between three and seven costs four and brings in the last vertex. All vertices now connect with three edges. The larger remaining costs cannot improve the greedy construction.
A path’s existence, rather than a bound on the distance between the endpoints, determines a component. Rejecting edges inside existing components prevents cycles; accepting edges between components grows connectivity. A tree on vertices has edges, so its compactness is structural, not an arbitrary stopping threshold.
Then cut both representations at thresholds . Their component counts are respectively four, three, two, two, and one. Equality includes an edge whose cost is exactly the threshold. At threshold two the first three points form one component, although the endpoints at zero and three are farther apart than two. The tree preserves chaining; it does not repair or conceal it.
For cosine-based article grouping, one can use cost one minus cosine, so a large similarity becomes a small cost. The input must first obey the nonzero-vector cosine convention, and the inequality must be translated consistently. No claim about event identity follows from this equivalence. It is a statement about a declared numerical graph.
Exercise. Why does moving the threshold from two to three leave the components unchanged? The newly admitted direct edge joins vertices already connected through the middle point. The notebook compares the full partitions, not merely their counts, so two different arrangements with the same number of groups could not pass by accident. Tied costs require a deterministic tree-selection policy, but need not make the threshold components ambiguous.
Factorizing a document table is not yet a complete retrieval example. A new query must use the training vocabulary, weighting rule, and coordinate map. The vector-space and latent-semantic-analysis papers provide the historical bridge; the precise TF-IDF recipe here is the one disclosed in the earlier companion, not a quotation of an original historical formula. Salton et al. 1975; Deerwester et al. 1990
Use six document rows and seven vocabulary columns, so and . The retained terms are troops, border, ceasefire, bank, rate, mortgage, and rise. The six documents, labeled d1 through d6 in order, are “troops shell border town,” “ceasefire troops border,” “bank raises rate,” “rate rise hits mortgage,” “troops ceasefire talks border,” and “bank mortgage rate rise.” Tokens outside the retained vocabulary do not receive columns. These are synthetic examples, not current news.
Let index retained terms, be the count matrix, and count documents containing term . Write for the natural logarithm. Define the training inverse-document-frequency (IDF) weight . A positive count gets damped frequency one plus its natural logarithm; an absent term gets zero. Multiply these values by the training IDF weights and normalize each nonzero row to obtain the document matrix . The earlier TF-IDF laboratory explains this recipe in detail; here its frozen application is the focus.
Let be the count row of a new query in those same seven columns. Applying the same weighting and normalization gives query row . Do not append the query to the training corpus and recompute document frequencies. That would change the fitted representation while pretending to perform only a lookup. The all-out-of-vocabulary query “radio quantum” has a zero row and returns no lexical direction, rather than dividing by zero.
Let be a positive retained rank no larger than the rank of . Define , of size , and , of size , as its first left and right singular columns. Let the diagonal matrix contain their positive singular values. Define the document-score matrix and the query-score row by
Compare the query with each document by cosine in this reduced space, checking for a zero reduced row as well. The lab ranks “ceasefire border” and “bank mortgage” in the original lexical space and at ranks two and three. It displays all scores and resolves display ties by document identifier. Rank choice changes the comparison space; the example makes no claim that one synthetic ordering measures real retrieval quality.
The document scores are , not just . Dividing both document and query coordinates by the corresponding singular values creates another metric. The notebook computes that alternative and exhibits changed similarities. It is a definable modeling choice, but should not be silently substituted for the stated projection.
For “ceasefire border,” set the ceasefire and border count entries to one and all others to zero. Damped frequency leaves each positive count at . Multiplication by each term’s stored IDF weight creates an unnormalized weighted row. Its length is the square root of the sum of those two squared weights. Divide every entry by that one length to obtain . Normalization changes the row’s scale without changing which terms it contains.
The th reduced coordinate is the dot product of with column of . Each direction can involve many vocabulary entries, so this step is a projection rather than selecting individual words. For a nonzero document row of , its reduced cosine with the query is . Both lengths belong to reduced coordinates: normalizing the original lexical row does not guarantee that its projection still has length one.
Why is the document table ? Multiply the retained approximation on the right by and use . The right factor cancels, leaving . Dividing by singular values reverses those coordinate scales; in a cosine calculation this generally reweights directions unequally and can change rankings. The choice of ruler must be part of the retrieval specification.
Exercise. Repeat “bank” in the query “bank bank mortgage.” The positive count becomes two, so logarithmic damping is now exercised rather than receiving only zero-or-one counts. IDF stays fixed. The lab recomputes the row and confirms that the transformation changes while the training document matrix remains untouched. Frozen preprocessing makes a lookup interpretable; it does not make vocabulary omissions harmless.
These lessons add a second kind of question to the geometry of observations. What does a probability promise? What follows from a recorded relation? How can a representation be learned? Small, fully specified examples let us inspect the arithmetic without treating a synthetic demonstration as evidence about the deployed newspapers. Every vector remains a row. Symbols for local neural computations carry explicit qualifiers so they do not replace the news and people notation.
Entropy describes uncertainty over specified possibilities. It does not measure the truth or social value of a news item. In Shannon’s communication theory, this separation lets us reason about distributions of messages without first resolving their meanings. The same mathematical distinction matters when an ontology interface proposes a question. Shannon 1948
Let count possible states, indexed locally by . Let be the probability of state , with nonnegative entries adding to one. Write for their row. Define entropy in bits, where is the base-two logarithm, and assign zero to the limiting expression :
The distributions and have entropies and approximately bits. A known state, represented by , has zero entropy. The outcome has become predictable, not necessarily desirable.
The logarithm asks what power produces a number. Since , : observing either equally likely outcome supplies one bit of information. Entropy averages this information over outcomes, weighting each by its probability. Thus the first calculation is . For the unequal probabilities it is . The rare outcome is more informative when it occurs, but contributes only one tenth of the average. This distinction between one outcome and the expectation will also matter when comparing questions.
The zero convention follows a limit: as a positive probability becomes smaller, its negative logarithm grows, but probability times that logarithm tends to zero. An impossible outcome contributes no expected information. We do not ask the computer to take a logarithm of zero and multiply the resulting infinity by zero; the lab omits zero-probability terms explicitly.
Now define a candidate question with answers indexed by . Let be the probability of answer if state is true. Each state’s answer probabilities sum to one. The probability of answer is . If , Bayes’ rule gives posterior state probabilities . A posterior is the distribution after observing an answer under this specified likelihood model. Define the question’s expected information gain by
Four equally likely concepts have entropy two bits. A deterministic question separating two concepts from two leaves one bit under either answer, gaining one bit. Separating one from three sometimes identifies the state completely, but its expected gain is only bits. The rare perfect answer cannot stand in for an average over answers.
For a noisy question, let the probabilities of “yes” in the four states be . “Yes” has probability and gives posterior . The opposite answer gives the reversed row. The notebook computes their entropy and the smaller expected gain. Answers with probability zero have no defined posterior and contribute no mass to the expectation.
To reproduce the update, multiply each prior probability by its conditional “yes” probability. The joint masses are , summing to . Dividing every mass by renormalizes the “yes” cases to a total of one and gives the stated posterior. This is Bayes’ rule as multiplication followed by normalization, rather than a change made because one answer sounds persuasive. For “no,” subtract each conditional “yes” probability from one before repeating the calculation.
Information gain is calculated before we know which answer will occur. Weight the entropy of each posterior by that answer’s probability, add, and subtract from prior entropy. For the deterministic one-versus-three split, the two answer probabilities are and . Their posterior entropies are zero and , giving expected remaining uncertainty . A question that sometimes identifies the state can therefore be worse on average than a question that always narrows it evenly.
Exercise. Would a question whose answer is always “yes” reduce uncertainty? Explanation. Its posterior equals its prior, so the gain is zero. A question’s usefulness depends on its likelihood model, not its wording alone. The ontology dependency contains information-based selection machinery, while the audited Anthropology picker uses text matching. This exercise establishes neither deployment nor reliable answer probabilities for a future interface.
A predicted probability becomes accountable when its event is specified and outcomes are observed. For example, the event might be that a reviewer, following a stated protocol, marks an article as substantially discussing databases. It is not an intrinsic probability that a person “is a database person.” A score summarizes forecasts against such outcomes; it does not repair an ambiguous target.
Let be the number of reviewed cases, index a case, its forecast probability, and its observed outcome. We use the modern one-category binary Brier score , where lower is better:
Forecast contributes when the event occurs and when it does not. For one of each outcome the mean is . Brier’s original multicategory convention sums errors over mutually exclusive categories. In a binary problem, including the complementary category doubles this value to . Neither arithmetic is wrong; comparisons require the same convention. Brier 1950
Another score is binary logarithmic loss. Define using the natural logarithm , with probabilities strictly between zero and one in this finite calculation:
The corresponding contributions for forecast are approximately and nats. A nat is the information unit obtained with natural logarithms. A zero probability assigned to an event that occurs incurs infinite log loss. Silently clipping probabilities changes the evaluated rule; a numerical implementation must declare any such convention.
To explain strict propriety, imagine a single event whose stipulated true occurrence probability is . Let be the binary random outcome, the reported forecast, and mean expectation over the two possible outcomes under that truth. Expanding the squared error gives
Only the first term depends on the report. Its unique minimum is at the truthful report, with expected score . Expected log loss also has its unique minimum there, approximately nats. The executable sweep checks both using reports from to . Strict propriety is a statement about expected optimization, not a guarantee that an estimated model is calibrated. Gneiting and Raftery 2007
The expectation is a weighted average over the two outcomes. With probability the event occurs and loss is ; with probability it does not and loss is . Expanding gives
The final term is unavoidable variability even under the correct probability. At it is . A truthful probabilistic forecast does not predict each outcome with certainty; its optimum expected loss need not be zero.
Before applying a derivative to a logarithm, we need its local rate rule. For a positive scalar input , the derivative of is . Two other rules used later are that the derivative of is , and the derivative of is when is nonzero. Here and is its inverse: . These rules describe limiting rates, in the same sense as the square derivative in the fitting lesson. Constants multiply derivatives, and the derivative of a sum is the sum of its terms’ derivatives.
For a composition, the chain rule multiplies the rate of the outer function by the rate of its inner input. In particular, the input changes at rate minus one when increases. Thus has derivative , while has derivative . The minus signs in log loss must be applied after these inner changes are accounted for. The network lesson develops this same multiplication along a longer path.
For log loss, differentiating the expected expression gives . Putting this over the positive denominator leaves numerator . The derivative is negative below the truth and positive above it, explaining the unique interior minimum checked by the sweep.
Exercise. Does one surprising outcome prove that a forecast of was dishonest? Explanation. No. The remaining outcome still has probability . Evaluate repeated forecasts under a defined sampling and review protocol. The proper-score principle explains a target for evaluation; it does not manufacture held-out data or justify displaying an arbitrary vector score as a percent confidence.
Calibration asks whether outcomes occur at the frequencies forecasts describe. Resolution asks whether forecasts distinguish groups with different event frequencies. These properties can pull apart. A forecaster that always predicts the overall base rate can be calibrated while telling us nothing about which particular case is more likely to be positive. Murphy’s decomposition makes this distinction quantitative. Murphy 1973
Use the binary outcomes and probabilities defined in the preceding lesson. Partition the cases into groups indexed locally by , each containing exactly the same forecast . Let count the group, its fraction of all cases, its observed positive fraction, and the overall positive fraction. These group weights are evaluation frequencies, not the article-fitting weights used for news covariance.
Define reliability error , resolution , and uncertainty by
Reliability error is smaller when group forecasts match their outcome frequencies. Resolution is larger when group frequencies differ from the base rate. For groups of identical forecasts the binary Brier score satisfies the exact finite-sample identity
Consider ten cases in two groups of five. The first outcomes are , with frequency ; the second are , with frequency . Overall prevalence is , so uncertainty is . A forecaster reporting and for these respective groups has reliability error zero and resolution . Its Brier score is .
Keep the outcomes fixed but change the reports to and . Resolution stays , because the two forecast groups are unchanged. Reliability error becomes , increasing the Brier score to . Finally, predict for every case. There is now one forecast group: reliability error is zero, resolution is zero, and the score is . Perfect aggregate reliability has not supplied discrimination.
Inside group , write . When we square and average over that group, the cross term vanishes because the average of is zero. The group’s mean squared error is therefore . The second term follows because binary outcomes satisfy .
Average over groups with weights . The first terms give reliability error. For the remaining terms, expand the definition of resolution to obtain , using . Consequently, . This proves the identity without assigning a separate empirical meaning to each case’s squared error.
For the well-calibrated fixture, each group contributes weight . Resolution is . Within the first group the direct score is ; the second group has the same value. These separate arithmetic routes make the identity inspectable. Merging all reports to also merges the groups, which is why the resolution then disappears.
Exercise. Why may we not substitute arbitrary wide probability bins into this identity without qualification? Explanation. Forecasts that vary inside a bin are no longer identical. Replacing them with one bin average can discard terms. Our code groups equal forecasts exactly, then verifies the identity directly against every case’s squared error. Reliability diagrams remain useful, but their binning, sample counts, and uncertainty must be disclosed. These fictional ten cases are an arithmetic fixture, not evidence that an Anthropology concept classifier has achieved any particular quality.
An incidence table records whether an object has an attribute. It can support a mathematics of exact shared properties, quite different from the covariance of numerical profiles. Formal concept analysis builds paired sets from such a table. Its word “concept” has a precise algebraic meaning here; it does not mean an eigenvector with an attractive label. Ganter and Wille 1999
Let be a finite set of objects and a finite set of attributes. These local sets are not article-background cells. Let be their binary incidence table, with one row per object and one column per attribute: entry is one when object has attribute . We deliberately do not call this table , which already denotes person–article association. For an object subset , define to be the attributes shared by every object in . For an attribute subset , define to be all objects possessing every attribute in . The prime denotes these two derivation operations, distinguished by their input sets:
Applying the derivations twice gives the object closure . It contains the original objects and is unchanged by another closure. A formal concept is a pair satisfying both and . Its object set is called its extent, and its attribute set its intent.
Use four fictional objects, numbered zero through three, with
attributes respectively {database, distributed},
{database}, {distributed, open}, and
{database, distributed, open}. Object zero’s shared
attributes are database and distributed. Both objects zero and three
have these attributes, so
.
Empty inputs require care: every attribute is shared by all objects in
an empty object set, because no counterexample exists. Thus the closure
of the empty extent here is
,
not the empty set.
The code enumerates all sixteen object subsets and finds six formal concepts. Order these concepts by extent inclusion. For two closed extents , their greatest lower bound is the largest closed extent contained in both, their intersection. Their least upper bound is the smallest closed extent containing both, the closure of their union. These operations are called meet and join. The notebook verifies the defining lower- and upper-bound conditions against every enumerated concept, rather than trusting the names.
Begin with the object set . Read across its row to collect database and distributed. Then read down those two attribute columns and keep only objects having both: objects zero and three. Reading across both of these rows again returns exactly database and distributed. We have reached the formal concept . Closure adds objects indistinguishable under the shared attributes; it does not invent a new recorded attribute or turn a similarity score into a fact.
“Every” matters in the definition. Selecting more objects can only reduce or preserve their shared attribute set. Conversely, requiring more attributes can only reduce or preserve the matching object set. Applying both operations therefore gives an object closure that contains the starting set, respects set inclusion, and no longer changes after another closure. These properties are called extensiveness, monotonicity, and idempotence, respectively.
A lower bound under extent inclusion is a concept whose extent is contained in both selected extents; the greatest such bound contains every other lower bound. An upper bound contains both, and the least upper bound is contained in every other upper bound. Raw union can fail to be closed, so the join needs closure after union. A lattice is an ordered collection in which every pair has these two bounds. The lab checks this universal condition over the finite list, including the empty-input convention rather than silently omitting it.
Exercise. Why does the most specific concept contain object three? Explanation. It has all three recorded attributes. Under extent inclusion, that singleton is below every other closed extent in this fixture. Real incomplete evidence may not justify treating an unrecorded attribute as absent. Our binary table is stipulated complete. Anthropology’s supported numerical people profiles neither supply that completeness nor imply that formal-concept inference is deployed.
An identifier, a text label, and a rule have different jobs. RDF describes statements as subject–predicate–object triples. The subject identifies the thing discussed; the predicate identifies a relation; the object is another identified thing or a literal value. A literal can be a number or text. Recording a tuple is easy. Deciding what follows from it requires declared semantics. Cyganiak et al. 2014
Consider a finite set
of triples. A triple
has subject
,
relation
,
and object
;
these letters are local tuple positions, not story scores or matrix
rank. We will use exactly two inference rules. The relation
subclass means class inclusion in this toy rule system, and
type assigns an individual to a class. For arbitrary
identifiers
and individual
,
the rules are
A premise is a statement required by a rule; a conclusion is the statement that rule licenses when its premises match. Provenance records where an assertion came from or which premises supported a conclusion. The arrow means that both premises license the displayed conclusion. These capital letters name classes locally; they are not the term table or concept count elsewhere in the book. Starting with asserted triples, repeatedly add rule conclusions until no new tuple appears. This fixed point is the closure under these two rules. Because our set of identifiers is finite and the procedure only adds triples, it terminates.
The fixture asserts five statements: Ada has type CEO; CEO is a subclass of Person; Person is a subclass of Agent; Databases has the broader concept Technology; and Ada has topic Databases. Closure adds three statements: CEO is a subclass of Agent, Ada has type Person, and Ada has type Agent. Each derived record retains its two immediate supporting premises. Tracing these links reaches the initial assertions.
Read the first rule as a pattern with a shared middle identifier. The premises CEO–subclass–Person and Person–subclass–Agent match because Person occupies that middle position, allowing CEO–subclass–Agent. The second rule matches Ada–type–CEO with CEO–subclass–Person to license Ada–type–Person. Applying it again with Person–subclass–Agent licenses Ada–type–Agent. Alternatively, the newly derived CEO–subclass–Agent provides another route to the same type conclusion. Adding an already present triple does not add new knowledge, so the implementation uses set membership rather than accumulating duplicate rows forever.
Repeated rule application is necessary because one conclusion can become another rule’s premise. A fixed point means a further full pass adds nothing; it does not mean that all possible facts about the represented world are known.
Nothing in either rule turns broader into
subclass, or hasTopic into type.
The code therefore does not derive that Databases is a subclass of
Technology or that Ada has type Databases. SKOS concept relations and
OWL class axioms serve different purposes; their shared graph-shaped
appearance is not permission to interchange them. This limited example
is not an OWL implementation or a SKOS conformance test. Miles and Bechhofer 2009; W3C OWL Working Group 2012
Exercise. If no statement about Bo appears, should the query “Bo has type Person” return false? Explanation. Under our open-world interpretation, it returns unknown. Lack of an assertion or derivation is not evidence of its negation. “Open world” is an interpretation of missing knowledge, not a guarantee that supplied assertions are correct. Provenance still matters: a valid derivation from an erroneous premise remains erroneous about the world. The newspapers’ numerical similarities cannot substitute for the premises of a sourced employment or financing relationship.
An embedding is useful partly because its coordinates can be learned rather than individually named in advance. Backpropagation efficiently applies the chain rule through a composed computation, allowing an error at the output to update parameters earlier in the network. The historical contribution of learned hidden representations should not be retold as the invention of differentiation itself. Rumelhart et al. 1986
Begin with one scalar input , equivalent to a one-coordinate observation row, and two scalar parameters . Define preactivation , hidden value , and prediction . Here is the hyperbolic tangent, a smooth nonlinear function with derivative . Let be a numerical target and let be half the squared prediction error:
The hidden value is an intermediate representation. Calling it hidden means it lies between input and output; it does not imply secrecy or a named semantic property. Denote output error by . The local symbol is distinct from the exponential constant.
The hyperbolic tangent can be written . Its output lies between minus one and one. Near zero it responds strongly to small changes; near either extreme its derivative becomes small. The nonlinearity matters: without it, multiplying by two scalar parameters would just be one linear coefficient .
First consider loss as a function of prediction alone. Increase prediction by while keeping the target fixed. The loss change is . Divide by and let it tend to zero. The local loss derivative is . The factor one half in the loss cancels the factor two from differentiating a square; it does not change the best-fitting parameters.
Next keep fixed and change . Prediction changes at rate , so the loss changes at rate . Changing takes a longer route: preactivation changes at rate , hidden value changes at rate with respect to preactivation, prediction changes at rate with respect to hidden value, and loss changes at rate with respect to prediction. Multiplying these local rates is the chain rule. Intermediate changes shrink along with the parameter change, so it is the first-order rates that multiply. In a network with several paths from one parameter to the loss, the contributions of those paths add. The resulting parameter derivatives are
Each factor in the second expression describes one dependency: output loss on prediction, prediction on hidden value, hidden value on preactivation, and preactivation on its parameter. Computing these local quantities once and reusing them is the central mechanism behind reverse accumulation in larger networks.
Set , , , and . The forward pass gives , prediction , and loss . The two derivatives are approximately and . A gradient-descent update subtracts a positive step size times each derivative. With step size , the new loss is approximately .
The forward arithmetic begins with , then evaluates , multiplies by for the prediction, and subtracts target to obtain . The local rates along the first-parameter path are this error, , , and . Their product is positive: both the output error and output weight are negative. Subtracting its step therefore decreases . The derivative for is negative, so that parameter increases. The new parameters are approximately . Recompute the entire forward path at the new parameters; retaining the old hidden value would not evaluate the updated network.
A derivative is a local prediction of change, not the full change over a finite step. The finite-difference check perturbs only one parameter at a time, runs the forward computation at both nearby settings, and divides the loss difference by the total parameter separation. Agreement with the analytic formula checks that dependencies and signs were handled correctly. It does not turn a chosen step size into a proof of improvement for every possible input.
The notebook also approximates each derivative by central finite differences. For a parameter displacement , this uses the loss difference at plus and minus , divided by . Choosing checks our formulas without treating finite differences as the intended training algorithm.
Exercise. Must an arbitrarily large descent step reduce the loss? Explanation. No. The derivative describes local behavior; a long step can overshoot or enter a differently curved region. Even a correctly decreasing training loss does not establish useful language representations. Our network has one invented example and no evaluation population. It teaches a differentiable dependency chain, not the training of the deployed embedding checkpoint.
A representation can be trained by rewarding observed word–context pairs while contrasting them with sampled alternatives. Negative sampling made this idea computationally useful in the word2vec lineage. It avoids treating every vocabulary entry as an explicitly normalized candidate at every update. Its objective is a training construction; a score between two learned rows is not a documented social relationship. Mikolov et al. 2013a; Mikolov et al. 2013b
Let be a center-word row. The local dimension counts this toy embedding’s coordinates. Let be an observed context row and one sampled negative context row of the same shape. Define positive and negative pair scores and . The logistic function is . For this one-positive, one-negative example, define loss
The first term rewards a large positive score for the observed pair. The second rewards a negative score for the sampled alternative. These two binary classifications do not constitute a probability distribution over the whole vocabulary. Larger training systems draw multiple negatives according to a specified sampling distribution; the choice of examples is part of the model.
Let be a scalar score, write , and name its positive denominator . A prime denotes a derivative with respect to . The exponential rule and the inner negative sign give . Applying the reciprocal rule to and multiplying by the denominator’s rate gives
For the last equality, one factor is and the other is . The derivative of the positive-pair loss now cancels an explicit probability factor:
For the negative-pair loss, put into that same positive-pair expression and multiply by the inner derivative minus one. The result is . The sigmoid identity turns it into . Thus both signs follow from declared scalar operations before any vector coordinates enter the calculation. The paired lab checks these local derivatives by central finite differences at the example’s pair scores.
Now a pair score is a dot product. If we change one coordinate of the center row while holding the context fixed, the score changes at the rate given by that matching context coordinate. Thus the gradient of the score with respect to the whole center row is the context row. Apply the scalar chain rule to each coordinate and add the two losses’ contributions.
Freeze both context rows and differentiate only the center row. The derivative, written in the same row shape as that parameter, is
Take , , , and . The scores are and , and the loss is approximately . The gradient is approximately . Subtracting times that gradient yields the center row . The positive score rises to , the negative score falls to , and loss falls to . Every value is recalculated by both notebook languages.
The implementation evaluates the equivalent stable softplus form: , where . A maximum-and-small-exponential rearrangement avoids overflowing large exponentials. Central finite differences independently check the center-row derivative.
Exercise. If the sampled negative changes, should this update remain identical? Explanation. Generally no. Its context row and score appear explicitly in the gradient. Negative sampling learns from the specified sampling experiment. Our single update illustrates that dependency; it cannot confer useful semantic structure on an entire vocabulary. The deployed newspaper uses a later pretrained sentence-embedding checkpoint, not these toy vectors or a newly trained word2vec model.
A token is a unit produced by a tokenizer, such as a word, subword, or punctuation unit. Transformer attention forms context-dependent combinations of token representations. This operation shares a word with the newspaper’s archive-attention score, but their inputs and purposes differ. Here weights depend on learned comparisons among tokens inside an input. Archive attention combines coverage and engagement to rank or summarize stories. Vaswani et al. 2017
A query describes what one token is looking for, a key describes how a token can be matched, and a value supplies the content to combine. These are roles of numerical rows; the toy fixture stipulates them directly.
Let count tokens and count query and key coordinates. Define query and key tables and , both . Define a value table of shape , where counts output coordinates. These qualified symbols are local: is not residual energy, and is not the news basis. Each row is one token’s representation.
Let be the table of pairwise scaled scores, called logits. For any finite row , let and index its positions and define . These are local token-position indices, not concepts. Applied separately to score rows, softmax gives the weight table . The output table is
The first two tables are square with one row and column per token; output has one row per token and columns. Every weight row is nonnegative and sums to one, so each output row is a weighted combination of value rows. These weights are not probabilities that a news claim is true.
For two tokens, use identity matrices for queries and keys, , and value rows and . The first logit row is . Its weights are approximately , yielding output . Swapping query position exchanges the weight order.
For query position and key position , the score is the dot product of those two rows, divided by . This is a conventional variance normalization, not a bound on every score. Under the simplifying assumption that query and key coordinates are mutually independent, each with zero mean and variance one, their coordinate products have variance one and their sum has variance . Dividing by its square root then gives score variance one. Actual learned coordinates need not satisfy those assumptions. For aligned all-ones rows, the unscaled dot product is , so the scaled score still grows as . The formula specifies the scale even in our tiny example; it does not promise that all inputs have comparable scores.
Softmax performs three operations on each score row: exponentiate each allowed score, add the resulting positive numbers, and divide each by their common sum. The output is a set of relative weights. To obtain output coordinate , multiply each value row’s coordinate by its weight and add. For the first query in the example the denominator is , since . Its output first coordinate is , and its second is . This last factor two explains why the second output coordinate is twice the second attention weight.
A mask identifies permitted keys for each query. A causal mask permits the first query to see only the first token and the second query to see both. The first output then becomes exactly . An excluded key receives zero softmax weight; merely setting its logit to zero would still give it positive weight. The implementation subtracts the largest permitted logit before exponentiating, which preserves the weights while avoiding overflow. It rejects an all-excluded row because no normalized distribution exists there.
Exercise. Why does adding to every permitted logit change no attention weight? Explanation. The common exponential factor cancels between numerator and denominator. Our stable implementation checks this identity without exponentiating directly. Real Transformers contain learned projections, multiple heads, position information, residual paths and other layers. This single-head example teaches the stated operation without claiming to reproduce an entire architecture.
A token’s representation can change when its surroundings change, while a search system needs one reusable row for a complete text. These are related but separate design decisions. BERT supplies contextual token representations; Sentence-BERT develops reusable sentence representations suited to comparison. A pooling rule alone does not recreate either training procedure. Devlin et al. 2019; Reimers and Gurevych 2019
Let count token coordinates and retain for the token count. Use an invented token table of shape . Set all three attention tables from the preceding lesson equal to this table, with . This makes a tiny contextual encoder with no learned projection. Call its output . Two-token inputs with rows and preserve the first input row but change its contextual output from approximately to . A representation now depends on a neighbor.
To illustrate a masked-token objective, replace the first input row with , retaining the second row . This replacement represents an invented mask token; it is not an attention-exclusion mask. The first contextual row becomes . Define a decoder with two rows and three vocabulary columns:
Multiplication by this table gives vocabulary logits . Their softmax probabilities are approximately . Taking vocabulary position one as the target, with positions counted from zero, the negative log target probability is . Reversing the context row to increases this target loss by exactly one nat. The notebook computes the change; it does not train BERT.
For sentence pooling, let mark which token positions contain real input rather than padding. This local token mask is unrelated to the number of people. If at least one position is retained, define mean row . Normalize a nonzero mean by its Euclidean length before cosine retrieval. Padding must be excluded from both attention keys and pooling. Appending a masked row therefore leaves our retained outputs and sentence row unchanged.
The two unmasked contexts above pool to perpendicular unit rows and . Their reusable cosine table is the identity. A simple contrastive objective chooses one candidate as positive. Let be the query’s cosine with candidate and let identify the positive. With a declared contrastive temperature of one, loss is . For similarities it is .
The pooling calculation has two distinct normalizations. First divide the sum of retained contextual rows by the number of retained tokens, obtaining a mean. Then divide that mean by its Euclidean length, obtaining a direction of unit length. For the first context the two attention rows exchange their weights, so each coordinate’s sum is one. Its mean is , of squared length . Normalizing yields . The changed context gives mean and unit row . Their dot product is ; each self-dot-product is one. Those calculations explain the identity cosine table without appealing to a semantic label.
For the contrastive calculation, exponentiating similarities produces weights . The probability assigned to the declared positive is . Its negative logarithm simplifies to . This probability is relative to the candidates supplied in that training comparison. Changing the candidate set changes its denominator even if the positive pair’s vectors stay fixed. A low contrastive loss and a calibrated probability of a factual event are therefore different claims.
Exercise. Does averaging contextual rows guarantee useful semantic retrieval? Explanation. No. The training objective, input construction, pooling, and evaluation population all matter. This fixture proves dependencies and invariances only; it supplies neither pretrained language knowledge nor evidence about a real archive’s search quality.
A classifier’s probabilities can be systematically too extreme even when its ordering is useful. Temperature scaling adjusts the conversion from a score to a probability using a separate calibration set. It is a small fitted model on top of another model, not an assertion that a sigmoid already guarantees calibration. Guo and colleagues studied this approach empirically; a successful result in one evaluation does not establish success in every population. Guo et al. 2017
For the binary case, let be case ’s scalar logit, meaning the score supplied to the sigmoid. Let be the temperature parameter. It is qualified to distinguish it from regression coefficients and is not the raw news-coordinate table . The adjusted probability is
A temperature above one moves probabilities towards ; a temperature below one makes them more extreme. Positive scaling preserves the sign of each logit and their ordering. Therefore it preserves the binary decision at probability threshold , though decisions at other thresholds can change. Choose the parameter by minimizing a declared loss on calibration cases excluded from the original parameter fitting. A separate test set is needed for an honest subsequent evaluation.
Reuse a deliberately simple fixture: one hundred calibration cases all receive probability , with fifty-five positive outcomes. Their common logit is . The observed frequency is . For a constant probability, binary log loss is minimized at this frequency. Solving the sigmoid equation gives the attainable positive-temperature optimum
The adjusted probability is . The calibration-set Brier score falls from to , and both versions still classify every case positive at threshold . The notebook checks the log loss at this optimum against smaller and larger temperatures. This analytic shortcut depends on the constant-logit fixture; a general calibration set requires numerical optimization of its actual objective.
To solve the sigmoid equation explicitly, start with a desired probability . Rearrange to and take natural logarithms. The ratio is called the odds; its logarithm is the log odds or logit. Thus . Substituting and gives the stated temperature.
For a common probability and observed positive fraction , the average log loss is . The preceding proper-score derivative, now with in place of the stipulated truth, vanishes at . This is a fitted sample optimum; using the same labels to choose and to evaluate that optimum does not supply an independent test. The separate populations below make that distinction visible.
The Brier arithmetic is before scaling and afterward. Threshold decisions have not changed, yet the distances between probabilities and outcomes have. A decision and its reported probability answer different questions.
Now evaluate two separate synthetic test populations. In the first, forty of one hundred cases are positive. Brier score changes from to . In the second, ninety-five are positive. It changes from to , becoming worse. We selected neither temperature nor reported result by looking at these test labels. The contrast shows how a shift in population can undermine a calibration rule fitted elsewhere.
Exercise. Could a positive temperature turn this constant positive logit into probability ? Explanation. No. Division by a positive number preserves its sign, so its sigmoid remains above . Temperature scaling has limited expressive power; it cannot repair every error. No calibrated ontology probabilities or deployed temperature-scaling layer are established for the audited sites. These examples specify how such a proposal could be evaluated, not a confidence claim about a person.
The embedding pipeline inspected in the History uses a quantized ONNX port. Quantization represents selected model numbers with fewer possible stored values. This can reduce storage and alter the computation available to an inference engine, but the exact scheme, operators, scales and resulting quality must be checked for the particular artifact. We teach one explicit scheme without claiming that it reproduces the deployed port. Qdrant n.d.
Let be a nonzero table of model weights. These qualified dimensions count this layer’s input and output coordinates, indexed by and respectively; is not the document TF-IDF table. Let be a scale, the integer table, and the reconstructed real-valued table. Our symmetric scheme uses integers from to :
Here rounding selects the nearest integer, with exact ties away from zero. Our maximum-based scale avoids saturation, which would clip an out-of-range value to an endpoint. Under these conditions, every reconstructed entry differs from its original by at most . The all-zero weight table requires a separate trivial convention because this formula would produce scale zero.
Use weight rows and . The scale is , and the stored rows are and . For input row , the original output is . Multiplying by the reconstructed table changes it slightly; the executable example prints every reconstructed weight and both output rows.
Define the error table and let index an output coordinate. The triangle inequality bounds the output error by
For this input, the bound is . The actual maximum error is smaller, approximately . A local layer bound does not establish unchanged neighbors or rankings after a full nonlinear model.
The weight is divided by scale , giving . Rounding gives integer , and reconstruction gives rather than exactly . Neighboring represented weights are separated by one scale unit, so rounding to the nearest one moves an entry by at most half a scale unit. This is the source of the entrywise bound; clipping an out-of-range value would need a separate bound.
The original first output coordinate is . Using integer weights and the common scale, the reconstructed output row is : each numerator comes from the same weighted column sum. For example, the first is . The rounding errors need not all have the same sign, so they can partly cancel in an output. To obtain a bound that does not depend on such cancellation, replace each signed contribution by its absolute value. The input’s sum of absolute coordinates is , hence the stated bound . The lab separately checks each reconstructed entry, the exact integer numerators, and the output bound.
Exercise. Are six eight-bit integers automatically an eightfold reduction in this example? Explanation. The six original float64 values occupy forty-eight number bytes. Six integer bytes plus one float64 scale occupy fourteen, before any file or operator metadata. Ignoring the scale overstates the saving on a tiny table. The OCaml lesson stores exact integer-valued floats for transparent arithmetic, then calculates the specified int8 payload size; it does not claim that its in-memory array itself has shrunk. Deployment requires the actual representation and runtime to realize the budget.
A successful notebook run establishes that its stated checks passed on the given inputs in that environment. It does not establish that an entire archive has the same properties. To make a larger claim, first state what must remain fixed: the representation, training sample, reference period, fitted transformation, and evaluation rule. The earlier companions explain how these choices affect news and people coordinates.
For any calculation, write down the input, the operation, the output, and the property that checks it. For a least-squares fit, the input includes both a design and a response, the operation chooses coefficients minimizing squared residual length, and the output includes the fitted response as well as those coefficients. Residual perpendicularity checks optimality, while the pseudoinverse’s coefficient length answers a separate question when solutions are not unique. Naming all four parts prevents an accurate scalar result from standing in for an explanation.
For a learned representation, trace the dependencies before editing a parameter. A change can pass through a score, normalization, weighted combination, pooling, and loss. The derivative describes local changes along those paths. A notebook that exposes each intermediate value makes it possible to locate an unexpected result, rather than treating a changed final score as an unexplained success or failure. The paired languages provide two implementations of the same declared experiment; the algebraic identities and reference arithmetic supply checks that remain meaningful even if both implementations share a misconception.
A retrieval experiment. Start with the frozen query lesson. Change the query while preserving its vocabulary and fitted directions. Compare a direct query projection with a coordinate rescaled by singular values. Explain which inner product or distance is being used before calling either result more similar. Then repeat the comparison after deliberately changing the reference model. Separate the effect of a changed query from the effect of a changed ruler.
A concept recommendation. Use the entropy lesson to select an informative question, the formal-context lesson to inspect which observed attributes travel together, and the semantic-rule lesson to state which inferences your rules actually license. An informative question can be useful even when the answer does not establish an ontology relation. Write the evidence needed to convert a numerical suggestion into a maintained factual assertion.
A representation experiment. Change one token row in the attention example, then change the pooling mask. Describe the route through attention weights, contextual rows, and the final sentence row. Compare the result with a contrastive objective and explain why a good training loss does not by itself prove good retrieval on a different domain.
A probability experiment. Compare a score’s ordering with its calibration. Use the proper-score and temperature lessons to explain why preserving the ordering of two examples need not preserve their probability errors. Separate the data used to choose a temperature from the data used to report its quality. The grouped Murphy identity is exact for its stated group predictions; it is not a guarantee that coarse bins diagnose every calibration error.
These are open exercises rather than additional claimed numerical results. The supplied code computes the fixed examples and exposes editable inputs. For each extension, state your prediction first, retain checks that still apply, and explain why any fixed-answer assertion must change. This makes the notebook a laboratory rather than a transcript to execute without inspection.
These are the previously published lessons used by this companion.
The linked section opens the existing public reader. Each laboratory
title identifies a worked lesson in both frozen notebooks. The study
bundle contains those notebooks under prerequisites/, with
index.html linking directly to every exact code cell. These
links do not assume a public notebook host.
eigentimes-math-python-1.0.0-fe8e8983.ipynb; edition
1.0.0-fe8e8983; 38 lessons.eigentimes-math-ocaml-1.0.0-fe8e8983.ipynb; edition
1.0.0-fe8e8983; 38 lessons.anthropology-math-python-1.0.0-ac612429.ipynb; edition
1.0.0-ac612429; 31 lessons.anthropology-math-ocaml-1.0.0-ac612429.ipynb; edition
1.0.0-ac612429; 31 lessons.The SHA-256 file checksums and source revisions are recorded in
research/prerequisites.json. Copies in the bundle retain
their original bytes and saved outputs. A cell is identified by its
stable ID, not by a browser line number. Their original numerical
examples remain separate from the new examples in this companion.
Counting words (Eigen companion).
python-009-3c9c6f21; ocaml:
ocaml-009-281cca42.TF‑IDF (Eigen companion).
python-012-266676b6; ocaml:
ocaml-012-9692ef44.A matrix times a vector (Eigen companion).
python-022-b5a4f5f9; ocaml:
ocaml-022-9f6c6b73.A matrix times a matrix (Eigen companion).
python-025-c97f6079; ocaml:
ocaml-025-4bfd0ce5.Transpose, symmetry, orthogonality (Eigen companion).
python-028-5c0acc3a; ocaml:
ocaml-028-624f1c6e.Centering and covariance (Eigen companion).
python-032-d484a6f1; ocaml:
ocaml-032-066421dc.Sums over blocks (Eigen companion).
python-106-80aca9d5; ocaml:
ocaml-106-ba6a6f5f.Eigenvectors: the axes of the cloud (Eigen companion).
python-035-a0ba2058; ocaml:
ocaml-035-21a2897f.How many axes to keep (Eigen companion).
python-038-e6fbb926; ocaml:
ocaml-038-fb9bc110.What the SVD is (Eigen companion).
python-042-ed38dce5; ocaml:
ocaml-042-5e9db47b.Truncation (Eigen companion).
python-045-2abd3ce6; ocaml:
ocaml-045-d7817043.Latent semantic analysis (Eigen companion).
python-048-6a0f53cd; ocaml:
ocaml-048-2469b260.Varimax (Eigen companion).
python-055-aa42178b; ocaml:
ocaml-055-3c51fef7.Projection, reconstruction, residual (Eigen companion).
python-059-59c63a4e; ocaml:
ocaml-059-7103155e.T² and Q (Eigen companion).
python-062-7269271c; ocaml:
ocaml-062-3020ef3d.Dominant axis, spectrum, profile (Eigen companion).
python-065-56ae5aa4; ocaml:
ocaml-065-580763be.±2σ on the spectrum (Eigen companion).
python-083-4e445193; ocaml:
ocaml-083-70c876ed.Stories: single linkage on a day’s articles (Eigen companion).
python-069-2eadc8bd; ocaml:
ocaml-069-5956b46f.python-071-f64e3672; ocaml:
ocaml-071-cd4001b5.Kuhn–Munkres (Eigen companion).
python-096-9d00ff86; ocaml:
ocaml-096-5744026e.Why nightly refinement is stable (Eigen companion).
python-099-afe09a1d; ocaml:
ocaml-099-4a9edd3d.Aligning maps needs anchors and room to test them (Anthropology companion).
python-105-254689a8; ocaml:
ocaml-105-d281bb97.Start with associations and keep the background unique (Anthropology companion).
python-026-ab58740d; ocaml:
ocaml-026-7253129f.Give supported cells equal weight (Anthropology companion).
python-032-dcd3e6cf; ocaml:
ocaml-032-f8ecc71b.Keep direction and report support separately (Anthropology companion).
python-035-7fcec816; ocaml:
ocaml-035-21840e49.Covariance asks which coordinates move together (Anthropology companion).
python-040-f8e223ea; ocaml:
ocaml-040-e97c7b18.Keep two patterns in the toy world (Anthropology companion).
python-043-ddb92c1d; ocaml:
ocaml-043-9540f874.Locate a person on the patterns (Anthropology companion).
python-046-572ba87d; ocaml:
ocaml-046-c4ff85bb.People can be ranked in two distinct ways (Anthropology companion).
python-052-3d3c0f2a; ocaml:
ocaml-052-a5e3ae77.python-054-356e18a7; ocaml:
ocaml-054-350b7ced.The return operation (Anthropology companion).
python-058-1bbf0100; ocaml:
ocaml-058-05bede39.An even smaller exact example (Anthropology companion).
python-061-b0bc45da; ocaml:
ocaml-061-c703ebaa.Why the singular value decomposition gives two views (Anthropology companion).
python-065-0162dc38; ocaml:
ocaml-065-3e5e4db6.Reuse the ruler (Anthropology companion).
Separate movement from a change of mixture (Anthropology companion).
python-076-159139e0; ocaml:
ocaml-076-28c3e1f4.Three clocks and a model version (Anthropology companion).
A tag, an axis, and a hierarchy (Anthropology companion).
A small, reviewed bridge comes first (Anthropology companion).
python-093-40c08441; ocaml:
ocaml-093-bc7d1dbf.Learning a concept score, one multiplication at a time (Anthropology companion).
python-096-cac2408e; ocaml:
ocaml-096-60fe0a2f.A score becomes a probability only after another question (Anthropology companion).
python-099-da7456b3; ocaml:
ocaml-099-3c21f499.Definitions here are reminders. The linked lesson introduces each specialized term before its calculation. The notation guide declares the common symbols; a local symbol is defined in its own lesson.
Attention. A weighted combination of value rows, with weights determined by query–key compatibility and normalization. Study.
Backpropagation. Application of the chain rule from a loss backward through intermediate computations to obtain parameter derivatives. Study.
Basis. An independent list of vectors spanning a space. Study.
Bayesian posterior. A probability distribution updated from a prior using a stated likelihood and observed evidence. Study.
Bidiagonal matrix. A matrix with possible nonzero entries only on the main diagonal and the diagonal immediately above or below it. Study.
Binary64. A floating-point format using 64 bits per value; it rounds many real numbers and arithmetic results. Study.
Brier score. Squared probability error; the binary convention used here records one outcome, whereas the original category-summed convention counts both binary outcomes. Study.
Calibration. Agreement between predicted probabilities and observed frequencies under a stated evaluation grouping or distribution. Study.
Cancellation. Loss of significant digits when nearly equal finite-precision quantities are subtracted. Study.
Chain rule. A derivative rule that multiplies local derivatives along a composed path and sums contributions from multiple paths. Study.
Closure. An operation that is extensive, monotone, and idempotent: it contains its input, preserves inclusion, and stops changing after one application. Study.
Collinearity. Linear dependence, or near dependence, among predictor columns; it makes fitted coefficients nonunique or sensitive. Study.
Compact SVD. A singular-value decomposition retaining only positive singular values and their corresponding columns. Study.
Completed square. A quadratic rewritten as a scaled square plus a constant; when the scale is positive, the constant is its minimum. Study.
Conclusion. A statement licensed by an inference rule when its required premises match. Study.
Condition number. A measure of worst-case relative output sensitivity to relative input changes for a specified problem and norm. Study.
Conditioning. The sensitivity of a mathematical problem to changes in its inputs, separate from the rounding behavior of a chosen algorithm. Study.
Contextual representation. A token representation influenced by surrounding tokens rather than by token identity alone. Study.
Contrastive loss. An objective comparing a designated positive pair against alternatives; it is not by itself a probability calibration guarantee. Study.
Corrected sum of squares. The sum of squared deviations about a stated mean; a centered summary can preserve it for stable merging. Study.
Covariance. A matrix summarizing how centered coordinate products vary together, under declared weights and normalization. Study.
Cross entropy. Expected negative log probability under a target distribution; for a single observed class it reduces to negative log probability of that class. Study.
Cycle. A closed graph path with no repeated intermediate vertex; trees contain no cycles. Study.
Derivative. The limiting output change per input change for a scalar function, when that limit exists. Study.
Determinant. For a square matrix, the signed volume scaling; in two dimensions, ad minus bc for rows (a,b) and (c,d). Study.
Dot product. The sum of coordinatewise products of two equally sized vectors. Study.
Eigenvalue. The scalar by which a square matrix multiplies one of its nonzero eigenvector directions. Study.
Eigenvector. A nonzero direction mapped to a scalar multiple of itself by a square matrix. Study.
Embedding. A numerical representation placing objects in a common coordinate space for specified computations. Study.
Entropy. Expected information content of an outcome under a stated probability distribution and logarithm base. Study.
Euclidean norm. The length of a vector: square each coordinate, add, and take a square root. Study.
Expectation. An average over possible outcomes weighted by their probabilities under a declared distribution. Study.
Expected information gain. The average reduction in uncertainty from observing an answer, averaging over the possible answers before selecting one. Study.
Exponential function. The positive function exp(x)=e raised to x, inverse to the natural logarithm and equal to its own derivative. Study.
Extensiveness. The property that a closure contains its original input set. Study.
Extent. The objects belonging to a formal concept in a binary object–attribute context. Study.
Finite precision. Representation of numbers using finitely many bits, with rounding of many exact mathematical operations. Study.
Fixed point. A state unchanged by another application of an operation; rule closure stops when another pass adds no statements. Study.
Fold-in. Mapping a new query through a previously fitted lexical representation without refitting its vocabulary or directions. Study.
Formal concept. An extent and intent that determine each other through the incidence relation of a formal context. Study.
Frobenius norm. The square root of the sum of squares of all matrix entries. Study.
Frozen transform. A learned transformation whose vocabulary, means, scales, and directions are held fixed while new observations are mapped. Study.
Gradient. The vector of partial derivatives of a scalar objective with respect to its parameters. Study.
Gradient descent. A parameter update subtracting a positive step size times the loss gradient; its local motivation does not guarantee improvement for every finite step. Study.
Gram matrix. A matrix of pairwise inner products; forming it can square a matrix’s spectral condition number. Study.
Gram–Schmidt. Construction of orthonormal columns by subtracting projections onto previously constructed columns and normalizing the remainder. Study.
Householder reflection. An orthogonal transformation reflecting across a hyperplane, used to zero selected coordinates without changing lengths. Study.
Hyperbolic tangent. A smooth nonlinear function taking values between minus one and one, with derivative one minus the square of its output. Study.
Idempotence. The property that applying an operation twice has the same result as applying it once. Study.
Incidence. A declared relationship recording which object has which attribute, or which person is associated with which article. Study.
Information content. The negative logarithm of an event’s probability, with the logarithm base determining the unit. Study.
Intent. The attributes shared by all objects in a formal concept’s extent. Study.
Intercept. The coefficient of an all-ones predictor column, representing the fitted response when other predictors are zero. Study.
Invariant. A property preserved by an operation when its assumptions hold. Study.
Jacobi rotation. A plane rotation chosen to remove an off-diagonal entry of a real symmetric matrix during iterative diagonalization. Study.
Lattice. A partially ordered set in which each pair has a greatest lower bound and a least upper bound. Study.
Least squares. Fitting parameters by minimizing a sum of squared residuals. Study.
Likelihood. The probability of observed evidence as a function of a hypothesized state or model. Study.
Log score. Negative log probability assigned to the observed outcome, when interpreted as a loss to minimize. Study.
Logarithm. The inverse of exponentiation; natural logarithms use base e and base-two logarithms measure information in bits. Study.
Logistic sigmoid. The function mapping a real score to a number between zero and one; its output still needs an appropriate probabilistic model and calibration evidence. Study.
Logit. A scalar score supplied to probability normalization; for a binary sigmoid probability, its inverse is the log odds. Study.
Mask. A rule excluding selected attention positions before the remaining weights are normalized. Study.
Mean squared error. The average of squared residual entries, obtained by dividing squared residual length by the number of entries. Study.
Minimum spanning tree. A tree connecting all vertices of a weighted undirected graph with minimum total edge weight. Study.
Minimum-norm solution. Among all solutions with the same best fit, the coefficient vector with the smallest Euclidean length. Study.
Monotonicity of closure. Preservation of inclusion: enlarging an input set cannot shrink its closure. Study.
Negative sampling. A training objective that contrasts an observed pair with sampled alternative pairs; it is distinct from full-vocabulary softmax likelihood. Study.
Normal equations. The least-squares stationarity equations obtained by setting the residual’s inner products with the design columns to zero. Study.
Nullspace. The set of input vectors that a matrix maps to zero. Study.
Numerical stability. The behavior of a computational procedure under finite-precision rounding, distinct from sensitivity inherent in its mathematical problem. Study.
Odds. For an event probability strictly between zero and one, its probability divided by the probability of its complement. Study.
Off-diagonal energy. The sum of squared entries outside the diagonal of a matrix; Jacobi rotations reduce it for the symmetric diagonalization problem. Study.
Ontology. An organized collection of concepts and relations with stated meanings and rules; similarity alone does not establish those relations. Study.
Open-world assumption. The absence of a fact does not establish its negation. Study.
Orthogonal matrix. A square matrix whose transpose is its inverse; it preserves Euclidean lengths and inner products. Study.
Orthonormal frame. A list of mutually orthogonal unit vectors, represented here as columns. Study.
Partial derivative. A local rate of change with respect to one input while holding the others fixed. Study.
Path. A sequence of graph edges meeting at successive vertices; existence of a path defines connectivity. Study.
Pooling. Combining token representations into one row, using a declared operation such as a mask-aware mean. Study.
Premise. A statement required by an inference rule before that rule licenses a conclusion. Study.
Principal angles. Angles between subspaces defined through singular values of the cross-product of their orthonormal frames. Study.
Prior. The probability distribution over hypotheses before incorporating the observation being considered. Study.
Projection. Mapping onto a subspace; an orthogonal projection leaves a perpendicular residual. Study.
Projector distance. A stated matrix norm of the difference between two orthogonal projection matrices, used to compare subspaces without naming individual axes. Study.
Proper scoring rule. A forecast score whose expected loss is minimized by reporting the true probability distribution. Study.
Provenance. A record of the source of an assertion or of the supporting premises for a derived conclusion. Study.
Pseudoinverse. The Moore–Penrose generalized inverse, yielding the minimum-norm least-squares solution even when a matrix is rank deficient. Study.
Pythagorean identity. The squared length of a sum of perpendicular vectors equals the sum of their squared lengths. Study.
QR factorization. Representation of a matrix as orthonormal columns times an upper triangular or upper-trapezoidal factor, with dimensions stated for full or thin forms. Study.
Quantization. Mapping a range of real numbers to a finite set of represented levels, usually trading storage or computation against error. Study.
Radian. An angle measured as circular arc length divided by radius; pi radians is a half-turn. Study.
Rank. The dimension of a matrix’s column span, equal to the dimension of its row span. Study.
Reciprocal rule. The derivative of one divided by a nonzero scalar input is minus one divided by that input squared; a composed denominator also requires the chain rule. Study.
Reconstruction. An approximation mapped back into the original coordinate system from a compressed representation. Study.
Reorthogonalization. Repeating projection subtraction to reduce numerical loss of orthogonality. Study.
Residual. The difference between an observed quantity and its fitted or reconstructed value. Study.
Residual length. The nonnegative Euclidean length of the vector of observed-minus-fitted discrepancies; it has the units of the response. Study.
Resolution. In the Murphy decomposition, the separation of group event frequencies from the overall event frequency. Study.
Root mean squared error. The square root of the average squared residual; unlike squared error, it has the response unit. Study.
Rounding error. Discrepancy caused by finite-precision representation or arithmetic, separate from model approximation or factual error. Study.
Semantic triple. A subject–predicate–object assertion, whose interpretation depends on the vocabulary and entailment rules used. Study.
Sine and cosine. The vertical and horizontal coordinates, respectively, of a unit-circle point after a specified counterclockwise turn from (1,0). Study.
Single linkage. Clustering by connected components after retaining pairwise edges below a distance threshold. Study.
Singular value. A nonnegative scale in the singular-value decomposition of a matrix. Study.
Slope. A fitted response change per unit change of a predictor when other predictors are held fixed. Study.
Softmax. The transformation of finite scores into positive normalized weights by exponentiating and dividing by their sum. Study.
Softplus. The function ln(1+exp(x)), evaluated with an algebraic rearrangement when needed to avoid overflow. Study.
Span. The set of all linear combinations of a list of vectors. Study.
Squared residual length. The sum of squared residual entries; squaring removes cancellation of opposite signs but changes the physical units. Study.
Temperature scaling. Dividing fixed model logits by a positive fitted scalar before probability normalization; fitting needs separate labeled data. Study.
Token. A unit produced by a text tokenizer, such as a word, subword, or punctuation unit. Study.
Transpose. The operation exchanging the rows and columns of a matrix. Study.
Uncertainty. In the Murphy decomposition, the overall binary event variance; elsewhere the term must specify its probability model and measure. Study.
Union-find. A data structure maintaining disjoint sets under unions and representative lookup, useful for graph connected components. Study.
Unit circle. The circle of radius one centered at the origin of two perpendicular coordinates, used to define sine and cosine. Study.
Original publication dates are given below. A linked title leads to its publisher, institutional archive, author copy, or model card; a separate scan link is supplied when useful. Biographical links identify authors and do not replace primary evidence. The source bundle includes the access and attribution ledger, machine-readable records, and a BibTeX file.
Åke Björck; Gene H. Golub. (1973). Numerical Methods for Computing Angles Between Linear Subspaces. Mathematics of Computation 27(123), 579–594. Primary copy.
Glenn W. Brier. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review 78(1), 1–3.
Tony F. Chan; Gene H. Golub; Randall J. LeVeque. (1983). Algorithms for Computing the Sample Variance: Analysis and Recommendations. The American Statistician 37(3), 242–247. Primary copy.
Richard Cyganiak (editor); David Wood (editor); Markus Lanthaler (editor). (2014). RDF 1.1 Concepts and Abstract Syntax. W3C Recommendation, 25 February 2014.
Scott Deerwester; Susan T. Dumais; George W. Furnas; Thomas K. Landauer; Richard Harshman. (1990). Indexing by latent semantic analysis. Journal of the American Society for Information Science 41(6), 391–407. Primary copy.
Jacob Devlin; Ming-Wei Chang; Kenton Lee; Kristina Toutanova. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT 2019, 4171–4186.
K. Florek; J. Łukaszewicz; J. Perkal; Hugo Steinhaus; S. Zubrzycki. (1951). Sur la liaison et la division des points d’un ensemble fini. Colloquium Mathematicum 2(3–4), 282–285. Primary copy.
Bernhard Ganter; Rudolf Wille. (1999). Formal Concept Analysis: Mathematical Foundations. Springer, first English edition.
Tilmann Gneiting; Adrian E. Raftery. (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association 102(477), 359–378.
Gene H. Golub; William Kahan. (1965). Calculating the Singular Values and Pseudo-Inverse of a Matrix. Journal of the Society for Industrial and Applied Mathematics, Series B: Numerical Analysis 2(2), 205–224.
G. H. Golub; C. Reinsch. (1970). Singular value decomposition and least squares solutions. Numerische Mathematik 14(5), 403–420. Primary copy.
J. C. Gower; G. J. S. Ross. (1969). Minimum Spanning Trees and Single Linkage Cluster Analysis. Journal of the Royal Statistical Society. Series C (Applied Statistics) 18(1), 54–64.
Jørgen Pedersen Gram. (1883). Ueber die Entwickelung reeller Functionen in Reihen mittelst der Methode der kleinsten Quadrate. Journal für die reine und angewandte Mathematik 94, 41–73.
Chuan Guo; Geoff Pleiss; Yu Sun; Kilian Q. Weinberger. (2017). On Calibration of Modern Neural Networks. ICML 2017, PMLR 70, 1321–1330.
Arthur E. Hoerl; Robert W. Kennard. (1970). Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics 12(1), 55–67. Primary copy.
Carl Gustav Jacob Jacobi. (1846). Über ein leichtes Verfahren die in der Theorie der Säcularstörungen vorkommenden Gleichungen numerisch aufzulösen. Journal für die reine und angewandte Mathematik 30, 51–94.
Adrien-Marie Legendre. (1805). Nouvelles méthodes pour la détermination des orbites des comètes. Paris: Firmin Didot, Appendix, 72–80.
Tomas Mikolov; Kai Chen; Greg Corrado; Jeffrey Dean. (2013a). Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781; submitted 16 January 2013.
Tomas Mikolov; Ilya Sutskever; Kai Chen; Greg S. Corrado; Jeff Dean. (2013b). Distributed Representations of Words and Phrases and their Compositionality. Advances in Neural Information Processing Systems 26.
Alistair Miles (editor); Sean Bechhofer (editor). (2009). SKOS Simple Knowledge Organization System Reference. W3C Recommendation, 18 August 2009.
Allan H. Murphy. (1973). A New Vector Partition of the Probability Score. Journal of Applied Meteorology 12(4), 595–600.
Qdrant. (n.d.). Qdrant/bge-small-en-v1.5-onnx-Q. Official quantized ONNX model card; exact artifact publication date not independently fixed. Model card, accessed 3 October 2026.
Nils Reimers; Iryna Gurevych. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. EMNLP-IJCNLP 2019, 3982–3992.
David E. Rumelhart; Geoffrey E. Hinton; Ronald J. Williams. (1986). Learning representations by back-propagating errors. Nature 323, 533–536.
G. Salton; A. Wong; C. S. Yang. (1975). A Vector Space Model for Automatic Indexing. Communications of the ACM 18(11), 613–620. Primary copy.
Erhard Schmidt. (1907). Zur Theorie der linearen und nichtlinearen Integralgleichungen. I. Teil: Entwicklung willkürlicher Funktionen nach Systemen vorgeschriebener. Mathematische Annalen 63, 433–476.
Claude E. Shannon. (1948). A Mathematical Theory of Communication. Bell System Technical Journal 27(3,4), 379–423,623–656. Primary copy.
Ashish Vaswani; Noam Shazeer; Niki Parmar; Jakob Uszkoreit; Llion Jones; Aidan N. Gomez; Łukasz Kaiser; Illia Polosukhin. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems 30.
W3C OWL Working Group. (2012). OWL 2 Web Ontology Language Document Overview (Second Edition). W3C Recommendation, 11 December 2012.
B. P. Welford. (1962). Note on a Method for Calculating Corrected Sums of Squares and Products. Technometrics 4(3), 419–420.
This index points to substantive definitions and discussions, rather than every occurrence. PDF references are page numbers; digital references are section links. Both are generated from the same index inventory.
Attention. Token attention is a normalized matrix computation.
Backpropagation. A hidden representation learns through the chain rule.
Basis. Notation and prerequisites before calculation.
Bayesian posterior. Uncertainty and the question worth asking.
Bidiagonal matrix. Reduce both sides without squaring the difficult scales.
Binary64. Reduce both sides without squaring the difficult scales.
Brier score. Probability forecasts and honest expected scores.
Calibration. Calibration can be perfect while resolution is absent.
Cancellation. Merge centered summaries without losing between-block variation.
Chain rule. A hidden representation learns through the chain rule.
Closure. An incidence table can generate a concept lattice.
Collinearity. Fit an error before trusting a coefficient.
Compact SVD. Recover the shortest solution when a fit is not unique.
Completed square. Fit an error before trusting a coefficient.
Conclusion. A recorded triple and a licensed inference are different.
Condition number. Reduce both sides without squaring the difficult scales.
Conditioning. Reduce both sides without squaring the difficult scales.
Contextual representation. Contextual token rows become reusable sentence rows.
Contrastive loss. Contextual token rows become reusable sentence rows.
Corrected sum of squares. Merge centered summaries without losing between-block variation.
Covariance. The shared series convention.
Cross entropy. Probability forecasts and honest expected scores.
Cycle. Find the same clusters in a tree and a complete graph.
Derivative. Notation and prerequisites before calculation.
Determinant. Reduce both sides without squaring the difficult scales.
Dot product. Notation and prerequisites before calculation.
Eigenvalue. Notation and prerequisites before calculation.
Eigenvector. Notation and prerequisites before calculation.
Embedding. A word vector learns by contrasting contexts.
Entropy. Uncertainty and the question worth asking.
Euclidean norm. Notation and prerequisites before calculation.
Expectation. Probability forecasts and honest expected scores.
Expected information gain. Uncertainty and the question worth asking.
Exponential function. Probability forecasts and honest expected scores.
Extensiveness. An incidence table can generate a concept lattice.
Extent. An incidence table can generate a concept lattice.
Finite precision. What a teaching computation establishes.
Fixed point. A recorded triple and a licensed inference are different.
Fold-in. Fold a new query into a frozen lexical model.
Formal concept. An incidence table can generate a concept lattice.
Frobenius norm. Notation and prerequisites before calculation.
Frozen transform. Fold a new query into a frozen lexical model.
Gradient. Notation and prerequisites before calculation.
Gradient descent. A hidden representation learns through the chain rule.
Gram matrix. Reduce both sides without squaring the difficult scales.
Gram–Schmidt. Build an orthogonal frame one subtraction at a time.
Householder reflection. Reflect a column and obtain a QR factorization.
Hyperbolic tangent. A hidden representation learns through the chain rule.
Idempotence. An incidence table can generate a concept lattice.
Incidence. An incidence table can generate a concept lattice.
Information content. Uncertainty and the question worth asking.
Intent. An incidence table can generate a concept lattice.
Intercept. Fit an error before trusting a coefficient.
Invariant. A route through the three companions.
Jacobi rotation. Watch Jacobi rotations remove off-diagonal entries.
Lattice. An incidence table can generate a concept lattice.
Least squares. Fit an error before trusting a coefficient.
Likelihood. Uncertainty and the question worth asking.
Log score. Probability forecasts and honest expected scores.
Logarithm. Notation and prerequisites before calculation.
Logistic sigmoid. A word vector learns by contrasting contexts.
Logit. A positive temperature changes confidence without changing the decision.
Mask. Token attention is a normalized matrix computation.
Mean squared error. Notation and prerequisites before calculation.
Minimum spanning tree. Find the same clusters in a tree and a complete graph.
Minimum-norm solution. Recover the shortest solution when a fit is not unique.
Monotonicity of closure. An incidence table can generate a concept lattice.
Negative sampling. A word vector learns by contrasting contexts.
Normal equations. Fit an error before trusting a coefficient.
Nullspace. Notation and prerequisites before calculation.
Numerical stability. Reduce both sides without squaring the difficult scales.
Odds. A positive temperature changes confidence without changing the decision.
Off-diagonal energy. Watch Jacobi rotations remove off-diagonal entries.
Ontology. A recorded triple and a licensed inference are different.
Open-world assumption. A recorded triple and a licensed inference are different.
Orthogonal matrix. Notation and prerequisites before calculation.
Orthonormal frame. Build an orthogonal frame one subtraction at a time.
Partial derivative. Notation and prerequisites before calculation.
Path. Find the same clusters in a tree and a complete graph.
Pooling. Contextual token rows become reusable sentence rows.
Premise. A recorded triple and a licensed inference are different.
Principal angles. Compare whole subspaces with principal angles.
Prior. Uncertainty and the question worth asking.
Projection. Build an orthogonal frame one subtraction at a time.
Projector distance. Compare whole subspaces with principal angles.
Proper scoring rule. Probability forecasts and honest expected scores.
Provenance. A recorded triple and a licensed inference are different.
Pseudoinverse. Recover the shortest solution when a fit is not unique.
Pythagorean identity. Build an orthogonal frame one subtraction at a time.
QR factorization. Reflect a column and obtain a QR factorization.
Quantization. Quantization exchanges precision for a storage budget.
Radian. Watch Jacobi rotations remove off-diagonal entries.
Rank. Notation and prerequisites before calculation.
Reciprocal rule. Probability forecasts and honest expected scores.
Reconstruction. Recover the shortest solution when a fit is not unique.
Reorthogonalization. Build an orthogonal frame one subtraction at a time.
Residual. Fit an error before trusting a coefficient.
Residual length. Fit an error before trusting a coefficient.
Resolution. Calibration can be perfect while resolution is absent.
Root mean squared error. Notation and prerequisites before calculation.
Rounding error. What a teaching computation establishes.
Semantic triple. A recorded triple and a licensed inference are different.
Sine and cosine. Watch Jacobi rotations remove off-diagonal entries.
Single linkage. Find the same clusters in a tree and a complete graph.
Singular value. Recover the shortest solution when a fit is not unique.
Slope. Fit an error before trusting a coefficient.
Softmax. Token attention is a normalized matrix computation.
Softplus. A word vector learns by contrasting contexts.
Span. Notation and prerequisites before calculation.
Squared residual length. Fit an error before trusting a coefficient.
Temperature scaling. A positive temperature changes confidence without changing the decision.
Token. Token attention is a normalized matrix computation.
Transpose. Notation and prerequisites before calculation.
Uncertainty. Calibration can be perfect while resolution is absent.
Union-find. Find the same clusters in a tree and a complete graph.
Unit circle. Watch Jacobi rotations remove off-diagonal entries.
Names link to verified Wikipedia biographies where available. Publication references remain in the bibliography.
Sean Bechhofer. publication.
Åke Björck. Compare whole subspaces with principal angles, Bibliography, publication.
Glenn W. Brier. Bibliography, publication.
Tony F. Chan. Merge centered summaries without losing between-block variation, Bibliography, publication.
Ming-Wei Chang. publication.
Kai Chen. publication, publication.
Greg Corrado. Bibliography, publication, publication.
Richard Cyganiak. publication.
Jeff Dean. Bibliography, publication, publication.
Scott Deerwester. Bibliography, publication.
Jacob Devlin. publication.
Susan T. Dumais. Bibliography, publication.
K. Florek. publication.
George W. Furnas. Bibliography, publication.
Bernhard Ganter. Bibliography, publication.
Tilmann Gneiting. publication.
Gene H. Golub. Recover the shortest solution when a fit is not unique, Compare whole subspaces with principal angles, Bibliography, publication, publication, publication, publication.
Aidan N. Gomez. Bibliography, publication.
J. C. Gower. publication.
Jørgen Pedersen Gram. Build an orthogonal frame one subtraction at a time, Bibliography, publication.
Chuan Guo. publication.
Iryna Gurevych. Bibliography, publication.
Richard A. Harshman. Bibliography, publication.
Geoffrey E. Hinton. Bibliography, publication.
Arthur E. Hoerl. publication.
Carl Gustav Jacob Jacobi. Watch Jacobi rotations remove off-diagonal entries, Bibliography, publication.
Llion Jones. Bibliography, publication.
William Kahan. Recover the shortest solution when a fit is not unique, Bibliography, publication.
Łukasz Kaiser. Bibliography, publication.
Robert W. Kennard. publication.
Thomas K. Landauer. Bibliography, publication.
Markus Lanthaler. publication.
Kenton Lee. publication.
Adrien-Marie Legendre. Fit an error before trusting a coefficient, Bibliography, publication.
Randall J. LeVeque. Merge centered summaries without losing between-block variation, Bibliography, publication.
Tomáš Mikolov. Bibliography, publication, publication.
Alistair Miles. publication.
Allan H. Murphy. publication.
Niki Parmar. publication.
Julian Perkal. Bibliography, publication.
Geoff Pleiss. publication.
Illia Polosukhin. Bibliography, publication.
Adrian E. Raftery. Bibliography, publication.
Nils Reimers. publication.
Christian Reinsch. Reflect a column and obtain a QR factorization, Bibliography, publication.
G. J. S. Ross. publication.
David E. Rumelhart. Bibliography, publication.
Gerard Salton. Bibliography, publication.
Erhard Schmidt. Build an orthogonal frame one subtraction at a time, Bibliography, publication.
Claude Shannon. Bibliography, publication.
Noam Shazeer. Bibliography, publication.
Hugo Steinhaus. Bibliography, publication.
Yu Sun. publication.
Ilya Sutskever. Bibliography, publication.
Kristina Toutanova. publication.
Jakob Uszkoreit. publication.
Ashish Vaswani. Bibliography, publication.
Kilian Q. Weinberger. publication.
B. P. Welford. publication.
Rudolf Wille. Bibliography, publication.
Ronald J. Williams. Bibliography, publication.
A. Wong. publication.
David Wood. publication.
C. S. Yang. publication.