# The Warning That Changes Its Evidence

A factory sees a high defect rate and adds a preventive inspection. The next batch improves. It retrains its estimate on that improved batch, concludes the danger has passed, and removes the inspection. Defects return. The system learns again. Nothing in the factory's underlying response has changed, yet the evidence alternates because the prediction governs the action that produces the next evidence.

This is a small instance of [performative prediction](https://proceedings.mlr.press/v119/perdomo20a.html): a prediction enters a decision and thereby changes the distribution on which later predictions are evaluated. Perdomo and colleagues' threshold example shows that repeated risk minimization can alternate when the induced distribution changes discontinuously at a threshold. The factory model uses the same elementary structure but is not their example. It separates the predictive rate from the action so the causal mistake is visible.

## The exact cycle

Let the defect probability under action \(a\in\{0,1\}\) be

$$
p(a)=b-ea,
$$

where \(a=1\) is the preventive action. The controller stores \(q_t\), the defect rate observed under the previous deployed action. It chooses

$$
a_t=\mathbf{1}[q_t\ge \tau]
$$

and then naively sets \(q_{t+1}=p(a_t)\). This is an exact population update: there is no sampling noise, batch dependence or mechanism drift.

For \(b=0.8\), \(e=0.6\), \(\tau=0.5\), and \(q_0=0.8\), the sequence is

$$
0.8\xrightarrow{a=1}0.2\xrightarrow{a=0}0.8\xrightarrow{a=1}0.2\cdots.
$$

The long-run defect probability is \(0.5\), compared with \(0.2\) if prevention is maintained. The controller's mistake is not that \(0.2\) is inaccurate. It is exactly the risk under action 1. The mistake is using that on-policy quantity as though it were the risk that would persist after switching to action 0. A calibrated estimate of the manifested outcome can therefore be the wrong input to a counterfactual decision.

The boundary is asymmetric because the rule takes action 1 on equality. Write \(L=b-e\) and \(H=b\). A two-cycle occurs exactly when

$$
L<\tau\le H.
$$

At \(\tau=L\), the low rate still triggers prevention, so the dynamics settle at action 1. At \(\tau=H\), the high rate triggers prevention and the low rate removes it, so the cycle remains. Below the interval the policy eventually holds action 1; above it the policy eventually holds action 0. The initial \(q_0\) selects the first action and, inside the cycle region, its phase. If \(e=0\), then \(L=H\) and no cycle exists because the action does not change the outcome.

These are exact decimal boundaries, not tolerance tests. The executable model reads finite numeric inputs by their displayed decimal values, converts them to rational numbers, and carries exact rational arithmetic through the response, trajectory and phase decisions before converting results to JSON numbers. Thus \(0.3-0.2=0.1\) at the lower endpoint, where equality must retain prevention. The boundary decisions do not depend on a decimal context or on the display precision of the JSON.

## Prevention is not free, and counterfactuals are not free

Defect reduction alone does not determine the preferred action. With preventive cost \(c\), the one-round objective is

$$
\ell(a)=p(a)+ca.
$$

An oracle that knows both action-specific risks chooses prevention when \(p(1)+c\le p(0)\), or equivalently when \(c\le e\); the implementation compares cost directly with effect and favors prevention on a tie. At costs \(0,0.2,0.6,0.8\), the oracle actions are respectively \(1,1,1,0\). The last choice accepts more defects because preventing them costs more under the stated objective. No moral conclusion follows from that arithmetic: the loss function is a modeling choice, and real defects and interventions need not share a common unit.

The oracle is a reference, not a fair learning competitor. It receives \(b\) and \(e\) for free. A learner that observes only action 1 sees \(p(1)\); it does not thereby identify \(p(0)\). Action variation, causal knowledge, or another valid source of counterfactual information is required. The accompanying observational-equivalence case makes this explicit with a separate, positive-response mechanism: a world with \(p(a)=0.8\) and a world with \(p(a)=0.2+0.6a\) both produce \(0.8\) whenever action 1 is deployed, while making different predictions under action 0. The second mechanism increases defects under action 1 and is not the preventive factory model. Equal on-policy evidence does not identify the response curve.

The [executable model](performative.py), [tests](test_performative.py), and [plotting data](data/performative.json) record these assumptions and boundary cases. They establish the finite arithmetic they compute. They do not show that intervention is inherently beneficial or harmful, validate a causal estimate from unvaried actions, or demonstrate that Circulatory Epistemology improves prediction.

The remaining CE question is narrower. A live loop can notice that its own action has entered the path that produces its evidence. Can a claim-specific return to the world then choose the observation or intervention that separates tracking a target from changing it, and can it do so better than ordinary causal analysis, control theory, or active experimental design? Calling the oscillation recognition adds nothing. The work begins only where the framework proposes a discipline that can be compared with those existing methods.
