The cost
of finding out.
A learner must spend part of its opportunity to act on finding out which action works.
A successful warning can make the event it warned against disappear. To interpret that success, the learner needs to know what happens under both actions. Here it must acquire that distinction within the same 256-unit horizon in which its decision will matter.
← Three threads notebook · Argument and exact derivation · Claim-to-evidence ledger
The same success can
have different causes.
All three factories produce a 20% defect risk with prevention. A record containing only those outcomes has exactly the same distribution in every world. More of that record cannot reveal how much prevention helped.
| Scenario | Without prevention | With prevention | Defect reduction | Lower-loss action |
|---|---|---|---|---|
| Strong reduction | 80% | 20% | 60 percentage points | Prevention |
| No reduction | 20% | 20% | 0 percentage points | No prevention |
| Marginal reduction | 55% | 20% | 35 percentage points | Prevention |
A defect costs 1 synthetic loss unit; prevention adds 0.30 per unit. Prevention is preferred when its reduction in defect probability exceeds 0.30. The no-reduction scenario has identical defect risk under both actions, so prevention is worse by its cost. These are separate scenarios, with no prior or average across them.
Testing the alternative
uses part of the future.
The design assigns an equal number of audit units to each action. From the two observed defect counts, the learner estimates the reduction and compares it with the cost. It then commits to one action for the remaining units. It receives one outcome per unit, never both possible outcomes for the same unit.
These controls belong to you as analyst. The learner is not given the chosen world's risks. Audit size is fixed before observing any outcomes.
| Full-horizon accounting | Expected loss units |
|---|
The oracle knows both action risks without spending units learning. Its lower loss is a privileged reference, not an available learning policy. Every learner protocol uses the same 256 units, with all defects and preventive-action costs included.
The lowest-loss audit on this grid is a hindsight comparison
The calculation knows the generating world. A learner with unknown risks cannot select this audit size for free. This is a comparison of predetermined designs, not a globally optimal or adaptive stopping rule.

Scroll sideways on a narrow screen to read the full figure.
In the strong-reduction world, increasing the audit from 20 to 128 units lowers wrong-action probability from about 3.21% to 0.00464%. Yet expected excess loss rises from about 5.28 to 19.20 units. The extra audit is more precise, but it assigns many more units to the worse action while learning.
A precise estimate can
still lead to the worse action.
In the marginal-reduction world, prevention's true benefit exceeds its cost by only 0.05 per unit. Even after auditing 128 units, this learner chooses the worse action about 28.29% of the time. The cost of that mistake is also smaller than in the strong-reduction world.

Scroll sideways on a narrow screen to read the full figure.
The strong and no-reduction worlds have the same estimator variance at each audit size, but different finite sampling distributions around the decision threshold. Exact ties favor prevention. These discrete effects explain why a larger audit does not lower decision error at every grid step.
The horizontal coordinate is sampling standard deviation; the vertical coordinate is exact wrong-action probability across repeated audits in the specified world. Neither is posterior confidence. Points come from complete finite-binomial enumeration, not Monte Carlo or empirical observations.
What did the return
cost us to make answerable?
The response can bear the trace of the prediction that prompted an action. Learning that causal history can require changing the action, and the change can consume opportunities where the answer matters. The “handshake” has an explicit causal route here: a prediction informs action, and action changes outcomes.
The model cannot decide what a defect or an intervention ought to cost. Its preferred policy is conditional on the objective we gave it. Changing those weights can change the preferred action; it does not change the action-specific defect risks. The living inquiry must still be able to question whether that objective expresses what matters.
Circulatory Epistemology's claim concerns truth becoming actual through living recognition between sensor and instrument. This model contains neither lived experience nor recognition. It makes one proposed scientific requirement precise: the return must distinguish the alternative being claimed. It gives no result that seeing creates observer-independent reality.
Explore-then-commit learning is established decision theory. A distinct contribution from the philosophy would have to improve how an inquiry identifies a consequential ambiguity, chooses a meaningful test, or revises the objective. These calculations make that challenge concrete; they do not establish that improvement.
The conditions that make these numbers possible
Independent exchangeable units; stationary Bernoulli risks under each action; no interference or carryover; the same action risks in audit and deployment; a known preventive cost; and a fixed audit followed by one committed action. The model measures neither elapsed time nor financial costs. It does not imply that all inquiry should intervene or that experimentation always pays.