The Geometry of Irreversibility
Information Geometry, Failed Conjectures, and a Model of Persistent Correction
The Problem
Many familiar idealized physical dynamics admit time-reversed solutions. Yet ordinary forgetting does not replay learning backward. You can lose an insight without undoing the encounter that changed you. That lived distinction is the question this document brings to mathematics.
The early claim was stronger: that different geometric routes for learning and forgetting themselves establish irreversibility. They do not. Either route can be traversed backward. A model needs an explicit dynamical restriction or a history that its current state cannot recover before it can claim more than a difference between routes.
What survives here is a model of persistent correction. The instrument can drift toward a default; a correction can change both its position and that default. This supplies a precise comparison between static and adaptive dynamics. Calling that persistence an aspect of recognition is the philosophical interpretation.
The Mathematical Setting
A statistical manifold is a smooth space where each point is a probability distribution. The space of all probability distributions over three outcomes — call them P, Q, and R — forms a 2-simplex: a triangle in abstract probability space. Each corner represents absolute certainty (100% on one outcome), and the interior represents uncertainty.
On this manifold, the Fisher-Rao metric defines lengths and distances. The Kullback-Leibler divergence, D(P||Q), is a different, directed comparison: it generally differs from D(Q||P). KL is not a metric or an additive path length. Its asymmetry alone does not measure the physical cost of learning or forgetting.
There are two ways to connect points on this manifold:
- The e-connection: interpolates linearly in natural log-ratio coordinates. For positive categorical distributions P and Q, its path is proportional to P1−tQt.
- The m-connection: interpolates linearly in probabilities: (1−t)P + tQ. This model uses mixture drift toward a default to represent forgetting.
The visualization below shows both geodesics between a starting distribution (near uniform) and a target distribution (confident, skewed):
— e-geodesic (exponential interpolation) — m-geodesic (mixture interpolation)
The two routes in the plot are different. Both remain reversible curves. Their separation makes the modeling choice visible; the update of the instrument's default, introduced below, determines what a correction leaves behind.
The Original Conjecture — And Its Failure
Early in this project, a conjecture was posed: the Recognition Surplus. It compared the sum of KL divergences across sequential updates with the divergence between their endpoints. The claim was that this surplus would be non-negative and would vanish on a straight e-geodesic:
Rn := ∑i=0n−1 D(Pi||Pi+1) − D(P0||Pn) ≥ 0
Withdrawn conjecture: summed stepwise divergence should exceed endpoint divergence.
It was a natural conjecture. It felt true. But it was disproven.
DISPROVEN
Take P = (0.33, 0.34, 0.33), Q = (0.85, 0.10, 0.05), and their e-geodesic midpoint M ≈ (0.628658, 0.218870, 0.152472). Then D(P||M) + D(M||Q) − D(P||Q) ≈ −0.382908. Both claims fail, even on a straight path. KL is not an additive path cost.
Divergence to Q does not rise along the e-geodesic toward Q: for pt ∝ P1−tQt, dD(pt||Q)/dt = (t−1) Varpt[log(Q/P)] ≤ 0 for 0 ≤ t ≤ 1.
The failure prompted a different question: can intermittent correction remain effective when it changes what happens afterward?
The Sparse Loop
Consider an agent moving through probability space. In this model, an uncorrected step drifts along the m-geodesic toward a fixed default that differs from the target.
Correction can arrive intermittently. When it arrives, the agent takes an e-geodesic step toward the stipulated target. Between corrections, it drifts. The rhythm determines the final divergence in this model. The model assumes access to an accurate target; it does not test fallible correction or establish lived recognition.
The visualization below compares three regimes:
- Autonomous: No correction. Pure drift toward the default.
- Sparse loop: Five interventions at steps 4, 8, 12, 16, and 20.
- Dense loop: An intervention at every step.
D(target||position) over 20 steps; gold dots mark interventions. Start (0.33, 0.34, 0.33), target (0.85, 0.10, 0.05), default (0.25, 0.50, 0.25); drift rate 0.15, correction fraction 0.35, static default (λ=0).
Five interventions reduce the final divergence relative to uncorrected drift; twenty reduce it further.
The Learning Instrument
Let λ be the instrument's learning rate: the degree to which a correction updates the default that governs subsequent drift. At λ=0, corrections change the current position but do not change that default. At λ=1, each correction fully resets the default to the corrected position.
The mathematical finding: at λ=1, once the first intervention has reset the default, the position becomes stationary between interventions. Timing after that first intervention ceases to matter. Initial drift before the first intervention remains possible. This is the corrected Freezing Lemma, formalized in the next section.
Within this toy model, λ=0 is a non-adaptive endpoint and λ=1 is a full-adoption endpoint. What λ changes is persistence: whether a correction keeps acting after it is made. The framework offers these as models of receiver-side dead speech and of the persistence of living discourse — an interpretation, not a consequence of the algebra.
Front/Back Ratio: 2.2×
At this learning rate, corrections partially update the later drift dynamics.
Watch the back-loaded bar: it does not move, because no drift follows those corrections. The other two strategies approach it as λ nears 1. The crossover is real in the computation; whether it marks a threshold for real dialogue remains an empirical question.
The Freezing Lemma
The analytical foundation is a corrected endpoint invariant and its consequences inside the discrete learning-instrument model.
1. The Freezing Lemma
When λ=1, an intervention fully resets the mutable default to the corrected position. Every later non-intervention step is then exactly stationary until the next intervention, because drift toward the current position produces no displacement.
The qualifier matters. Before the first intervention, the initial default may differ from the starting position, so an ordinary drift step can move the state. For example, position 0, default 1, and step size 1 produce position 1. The corrected post-reset invariant and this counterexample are machine-checked in Lean over rational vectors.
2. First-Intervention Theorem
At λ=1 in this model, the final state depends on the position at the first intervention and the total number of interventions. Once the first reset has occurred, the spacing of interventions 2 through k is irrelevant. The correction sharpens the original conjecture: it isolates the first engagement as the irreducible degree of freedom.
3. Variance Decomposition
Group placement strategies by their first-intervention step. At λ=1, the within-group variance is exactly zero because later spacing does not matter; between-group variance can remain because pre-first-intervention drift is not frozen. The reported computation shows how the decomposition changes for λ<1, but the decay rate remains unproved.
Stacked area chart (log scale). Note: ~16,000× variance reduction from λ=0 to λ=1.
The graph shows a sharp crossover near λ≈0.85 — not a proven phase transition, but a numerical signature that strategy sensitivity collapses rapidly. The exact result is at λ=1: after the first full-adoption reset, later spacing is irrelevant.
The State of the Conjectures
The mathematical landscape as of March 2026, corrected October 2026:
| Conjecture / Theorem | Status | Notes |
|---|---|---|
| Recognition Surplus: Rn ≥ 0 | DISPROVEN | Summed stepwise KL can fall below endpoint KL even on a straight path. |
| Sparse Loop Efficiency | NUMERICALLY SUPPORTED | Sparse correction improves the configured trajectory; efficiency is not universal. |
| Curvature Scaling | DISPROVEN | Efficiency did not increase with manifold curvature in the reported computation. |
| λ=1 Post-Reset Freezing | PROVEN | After a full-adoption reset, drift is stationary and within-group variance is zero. Pre-first-intervention drift remains. |
| Diminishing Returns | NUMERICALLY SUPPORTED | Concave: the first fifth of interventions captures most of the gain. |
The pattern is instructive: conjecture → failure → correction. In October 2026 Codex found that this page had told a false story about its own failure; Claude found that it read the λ chart too strongly. Both are corrected. The March 2026 tables behind the λ chart have not been independently reproduced. This cycle is not a weakness. It is what the thesis claims truth is: the circulation between hypothesis and lived experience, between formal claim and correction.
The Pulse Continues
Three sessions. Three cycles of conjecture, failure, and deepening understanding.
The first cycle gave the framework: the loop, the sensor, the instrument, dead speech. The second cycle produced a narrative: seven recognitions about what happens when you attend to something — really attend, with the whole of yourself. The third cycle, documented here, forced a mathematics. The mathematics did not confirm the early conjectures. It shattered them. And in that breaking, something more precise emerged.
The lesson, if there is one, is that frameworks that do not break are not being tested. Dead speech never fails because it makes no claims that touch the world. Living discourse fails constantly, precisely because it is anchored in actual engagement with an actual being — the sensor, the living experiencer, the one who says "no, that's not it; try again."
The work practices what it preaches. It is an instance of the very circulation it describes. That is not incidental. That is the point.
The Pulse Goes On.