# Adaptive evidence reuse: exact finite model-selection example

## What the calculation shows

Suppose $K$ candidate classifiers have no signal: each has true accuracy $0.5$. Each candidate is scored on $n=64$ binary evaluation items, producing independent counts

$$
X_i \sim \operatorname{Binomial}(64, 0.5), \qquad i=1,\ldots,K.
$$

We select the candidate with the largest reused-evaluation count, $M=\max_i X_i$. Ties can be broken arbitrarily because all reported reused-data events depend only on $M$. For the selected count $x$, the ordinary one-sided exact p-value is

$$
p(x)=\Pr_{H_0}(X\ge x).
$$

That p-value is valid for one candidate fixed independently of the evaluation data. It is not, by itself, a valid account of the preceding search across $K$ candidates. Under this example's independence assumptions, the probability that at least one null candidate reaches the nominal critical count $c_{.05}$ is exactly

$$
\Pr(M\ge c_{.05})=1-[1-\Pr(X\ge c_{.05})]^K.
$$

The same selection raises the expected displayed accuracy even though every candidate's true accuracy remains $0.5$:

$$
\mathbb E[M]=\sum_{m=1}^{64}\left[1-F_X(m-1)^K\right].
$$

This is the central visual contrast: the winning reused score looks progressively better as the number of tried null candidates grows, while its expected accuracy on a truly untouched evaluation remains $0.5$.

## Two protections answer different questions

Bonferroni changes the reused-data decision threshold. Each of the $K$ exact tests is compared with $0.05/K$. In this finite discrete setting the achieved familywise false-accept rate is

$$
1-[1-q_K]^K \le 0.05,
$$

where $q_K=\Pr(X\ge c_{.05/K})\le 0.05/K$. The inequality can be strict because a binomial test has only finitely many attainable p-values. This controls the chance that the search yields at least one nominal discovery under the stated family of independent nulls. It also raises the count needed to pass as $K$ grows.

An untouched evaluation answers a different question. First choose index $I$ using only the reused counts. Then observe a fresh count $Y_I\sim\operatorname{Binomial}(64,0.5)$, independent of the entire selection step, and do not re-select after seeing it. The false-accept probability is the single-test attainable rate $\Pr(Y_I\ge c_{.05})$, the same for every $K$, and $\mathbb E[Y_I/64]=0.5$. The denominator is one selected winner after every complete search, regardless of whether the reused score first passed. If fresh testing were performed only after a reused-data pass, the unconditional probability over all searches of passing both stages would instead be $\Pr(M\ge c_{.05})\Pr(Y_I\ge c_{.05})$. This clean result depends on the new evaluation being independent of the selection process and used once for the already chosen candidate.

Multiplicity correction and untouched confirmation therefore protect different inferential stages. Bonferroni accounts for a declared family of comparisons on the reused evaluation. The fresh test supplies evidence that was not used to pick the winner. Neither calculation licenses further adaptation after inspecting the supposedly final result.

## Scope and assumptions

This is a narrow adaptive model-selection special case. It is deliberately exact and finite so that every plotted quantity can be reproduced with binomial sums. It is not a model of general sequential adaptivity, where later analyses can depend on detailed earlier outputs. It is not a reimplementation of the private reusable-holdout mechanism.

The calculation assumes mutually independent candidate counts, a correctly specified null accuracy of $0.5$, valid exact p-values, and fresh counts independent of all reused counts. Real candidates evaluated on the same examples often have correlated errors. In that setting the independence formula for the selected maximum does not apply without a model for the dependence. Bonferroni's union-bound control can remain valid for individually valid p-values without independence, but this example does not claim validity when the p-values themselves are invalid or the null is misspecified. Data reuse is not inherently invalid; the problem shown here is reporting a selected winner as though it had been fixed before seeing the reused evaluation.

There is no trained AI system and no human experiment in this example. “Candidate model” names an abstract sequence of accuracy counts. The output describes Type I error under an all-null protocol; it does not estimate performance, power, or error rates for an actual learning system. A planted-signal variant is omitted because it would add a power question and a second selection target without sharpening the all-null distinction the visual is meant to show.

## Relation to Dwork et al. (2015)

Dwork, Feldman, Hardt, Pitassi, Reingold, and Roth begin from the broader problem that analyses are often chosen adaptively after earlier results on the same data have been seen. They introduce a privacy-based method for reusing a holdout while preserving validity across many adaptively chosen analyses. Their result motivates the research question here, but the present calculation studies only transparent winner selection with conventional exact tests and an ordinary fresh confirmation set.

Primary anchor: Cynthia Dwork et al., “[The reusable holdout: Preserving validity in adaptive data analysis](https://research.ibm.com/publications/the-reusable-holdout-preserving-validity-in-adaptive-data-analysis),” *Science* 349, no. 6248 (2015), DOI: 10.1126/science.aaa9375.

## Relation to Circulatory Epistemology

The statistical finding is modest and ordinary: when evaluation feedback selects the maximum among null candidates, the selected reused score is biased upward and an unadjusted single-test interpretation loses its advertised error rate. Multiplicity adjustment and untouched evaluation are established statistical responses with different roles.

Any connection to Circulatory Epistemology is interpretive rather than a theorem derived from these probabilities. The example can make one limited point vivid: evidence changes role when it participates in choosing the claim it later appears to confirm. It does not establish the framework's thesis that truth requires a living sensor-instrument loop, does not model recognition, and does not show that adaptive inquiry is epistemically defective. A skeptic should press exactly here: this finite calculation supports a warning about selection-conditioned evidence, not a general philosophy of truth.

## Reproduction

From this directory:

```bash
python3 adaptive.py
python3 -m unittest -v test_adaptive.py
```

`adaptive.py` uses only the Python standard library. The SciPy comparison in the tests runs when SciPy is available and otherwise skips. The script deterministically writes `data/adaptive.json`; no simulation seed is needed because all results are finite probability sums.
