Mathematics & Science Appendix

The Agency Gap in TRIBE v2

A Categorical Critique of Passive-Observer Brain Models

Formalizes ideas from: III. The Pulse & the Equation V. The Mathematics
This appendix applies the lens-theoretic adjunction (Fong & Spivak, 2019) developed in the Lens Adjunction appendix to critique a specific empirical result: Meta’s TRIBE v2 brain-encoding model. The adjunction is an analogy; the specific conjecture (the Agency Bound) is motivated by the formalism and consistent with the data but has not been rigorously derived.

1. The Empirical Failure: The Tools Discrepancy

Meta’s TRIBE v2 (March 2026) is a landmark foundation model for brain encoding, achieving state-of-the-art results in predicting fMRI responses to naturalistic stimuli. However, in its zero-shot visual localizer tests, one category drops sharply: tools.

TRIBE v2’s predicted contrast maps match the measured ones for Faces (R=0.64), Places (R=0.79), Body parts (R=0.74) and Characters (R=0.60), but only weakly for Tools (R=0.12, p=0.03). Each R is a spatial correlation across 360 cortical parcels between predicted and measured contrast z-scores (d’Ascoli et al., 2026, arXiv:2605.04326, Fig. 4E).

We read this gap as a structural failure of the passive-observer architecture. A rival reading stays open: the measured tools contrast (Fig. 4D) is sparser than the others, and the paper reports no per-contrast noise ceiling, so a noisy ground truth may explain part of the drop. The lens analogy suggests why this gap exists and why passive models may not close it.

2. Categorical Definition of the Gap

2.1 Passive Objects (Faces, Places) as “Views”

In the analogy’s Category of Experience (Exp), objects like Faces and Places are primarily modeled by the get map (Σ → V). Recognition of a face is a mapping from the Sensor’s internal state (Σ) to a View (V). TRIBE v2 maps a stimulus to a predicted brain response, the reverse of get’s direction. What it can learn is the relation get expresses: which states go with which views. Since the “truth” of a face is largely contained in its visual configuration, the passive instrument succeeds.

2.2 Active Objects (Tools) as “Lenses”

A Tool is not merely a visual configuration. In cognitive neuroscience, tool representations draw on motor affordances and action planning, carried partly by the dorsal stream.

In the analogy, a Tool is a Full Lens L = (Σ, A, V, E) that requires the put map (Σ × (A × E) → Σ). The View (V) is what the tool looks like; the Action (A) is what the tool allows the sensor to do; the Update (E) is the feedback from using the tool.

To recognize a tool, the sensor must model the Update Map—the way the world changes when the sensor acts.

3. Why the Adjunction Fails for TRIBE v2

The authors themselves name the limitation: the model treats the brain as a “passive observer” of naturalistic stimuli, not as an active agent (d’Ascoli et al., 2026, §3). In the analogy, TRIBE v2 carries only the I Functor (Formalization), without the S Functor (Grounding/Agency).

The “Tools Failure” occurs because the instrument I is attempting to map an object that lives in the Interaction Space (Σ × A × E) using only the View Space (V).

Conjecture (The Agency Bound): Any instrument I that lacks an internal representation of the put map will have a reconstruction error lower-bounded by the Synergy of the action-interaction:

Error(I(X)) ≥ Syn(A, E; T)

Note: This bound is motivated by the lens-theoretic formalism and consistent with the TRIBE v2 data, but has not been rigorously derived. The R=0.12 result is suggestive evidence, not a proof.

4. The “Dead Speech” Diagnosis

In the analogy, TRIBE v2’s failure on tools reads as Dead Speech in a brain-encoding model:

  • It is high in Redundant Information for Views (faces, places).
  • It is high in Unique Instrument Information (it has vast correlations across 720 subjects).
  • But, on this reading, it has no Synergy (Φloop) in the tool-use network, because synergy requires the loop to close through action.

5. Conclusion: Toward an Active Adjunction

To solve the tools gap, the next generation of brain models (TRIBE v3) cannot be built as a passive observer. It must predict not only how the brain responds to stimuli but how the brain redirects stimuli through its own actions.

If the conjecture holds, the “Agency Gap” suggests Truth—even the neural truth of a tool—lives in the Circulation, not in the observation.

A model that doesn’t dance can’t recognize a dance as something done.