Circulatory Epistemology · Evidence reuse experiment 02

A real signal
can lose the search.

More candidates can raise the winning score while making the genuinely better candidate less likely to win.

One candidate really has 65% accuracy. Its competitors have 50%. They share some of the same easy and difficult cases. This experiment asks whether the search recovers the real advantage—and what would justify accepting the winner.

01 / WHO WINS?

The best-looking result
need not be the best candidate.

Every search scores all candidates on 64 shared items and selects the highest score. Ties are broken uniformly. Increasing the number of candidates adds null competitors; there is still exactly one real signal. The selector is not told which candidate it is.

Probability the real signal wins selection, over all searches
Expected winning score on the reused selection items
Selected candidate's true accuracy, averaged across searches

The last quantity averages 65% when the signal wins and 50% when a null wins. It is a property of the specified generating model, not a performance estimate supplied to the selector.

Line chart of the signal's chance of being selected, using every search as the denominator. At K=2, recovery is 95.76%, 96.09%, 97.07% for delta 0.00, 0.10, and 0.20; At K=100, recovery is 46.93%, 48.74%, 54.92% for delta 0.00, 0.10, and 0.20. One p=0.65 signal competes with K-1 p=0.50 nulls on the same 64 items, with uniform tie-breaking. These finite-law rates, numerically evaluated, do not establish a general effect of dependence outside the stated shared-difficulty mechanism.

Scroll sideways on a narrow screen to read the full figure.

Selecting from more candidates can hide the planted signal
PNG · SVG · PDF

With independent candidates, moving from 2 to 100 candidates raises the expected winning score from 65.15% to 67.86%, while the selected candidate's average true accuracy falls from 64.36% to 57.04%. The search has more opportunities to promote a lucky null.

Why dependence helps selection in these particular scenarios

Each item is independently easy or difficult. Conditional success probability is p + δ on an easy item and p − δ on a difficult one. All candidates experience the same difficulty sign, while their remaining outcomes are conditionally independent. Increasing δ preserves the marginal abilities of 65% and 50%.

At 100 candidates, raising δ from 0 to 0.20 increases signal selection from 46.93% to 54.92%. Under this additive mechanism, the common difficulty component partly cancels when candidates are compared. This is not a general claim that correlation improves selection. Other dependence structures can have different consequences.

02 / WHAT GETS ACCEPTED?

Selecting a winner
does not settle its claim.

Compare three rules after the same selection step. Nominal reuse tests the winning score at a nominal 5% cutoff. Bonferroni reuse adjusts that cutoff for the number of candidates. Fresh confirmation tests the one selected winner on 64 new independent items, after every search, with no requirement to pass reused testing first.

Every percentage below uses all searches as its denominator. Accepted signal, accepted null and no acceptance partition those searches. “No acceptance” still includes a selected winner; it means the applicable test did not pass.

Acceptance ruleAccepted signalAccepted nullNo acceptanceUnique itemsCandidate-item evaluations

Fresh confirmation uses 128 unique items in total; the reuse rules use 64. Its additional evidence is part of the comparison. No candidate is reselected after the fresh result.

Any null crossing is a different event

A null can cross the threshold while the real signal scores higher and wins. The chance that any null crosses is therefore different from the chance that the selected null is accepted.

RuleAny null crossesSelected null accepted

The fresh familywise entry is not applicable: only the selected winner is tested on the fresh batch, not every unselected null. The ordinary single-null test has an attainable 2.9971% rejection probability here; discreteness makes it lower than the nominal 5%.

Nine stacked bars at K=100 partition all searches. delta 0.00, nominal reuse: accepted signal 46.76%, accepted null 51.83%, no acceptance 1.42%; delta 0.00, Bonferroni reuse: accepted signal 15.16%, accepted null 2.70%, no acceptance 82.13%; delta 0.00, fresh confirmation: accepted signal 33.41%, accepted null 1.59%, no acceptance 65.00%; delta 0.10, nominal reuse: accepted signal 48.18%, accepted null 48.38%, no acceptance 3.44%; delta 0.10, Bonferroni reuse: accepted signal 15.14%, accepted null 2.59%, no acceptance 82.27%; delta 0.10, fresh confirmation: accepted signal 34.70%, accepted null 1.54%, no acceptance 63.77%; delta 0.20, nominal reuse: accepted signal 52.23%, accepted null 37.54%, no acceptance 10.23%; delta 0.20, Bonferroni reuse: accepted signal 15.05%, accepted null 2.16%, no acceptance 82.79%; delta 0.20, fresh confirmation: accepted signal 39.10%, accepted null 1.35%, no acceptance 59.55%. Reuse uses 64 unique items. Fresh confirmation uses 128 unique items and tests only the selected winner on the fresh batch. Any-null threshold crossing is a separate reuse quantity and is not applicable to fresh confirmation; the evidence budgets are not equal.

Scroll sideways on a narrow screen to read the full figure.

Selection and acceptance answer different questions at K = 100
PNG · SVG · PDF

At 100 candidates and δ = 0.20, nominal reuse accepts a null in 37.54% of searches and the signal in 52.23%. Bonferroni reduces accepted nulls to 2.16%, while accepting the signal in 15.05%. Fresh confirmation accepts a null in 1.35% and confirms the signal in 39.10%, using the extra 64 items.

These are numerically evaluated finite laws, not simulated or empirical frequencies. Bonferroni's union bound needs valid marginal null tests, but does not require independent candidates. Fresh confirmation needs independence from the selection batch. This experiment neither covers arbitrary sequential adaptivity nor implements the private reusable-holdout algorithm.

03 / THE SURPRISE CAN STILL MATTER

What redirects the inquiry
may not settle the new question.

The unexpected winner can be the real signal. Nothing in this calculation makes surprise worthless. The difficulty is that some of the apparent confirmation has already been used to choose what deserves attention. Its history remains part of the evidence.

For Circulatory Epistemology, this sharpens a proposed discipline of the return: remember what the evidence helped select, and what the next encounter could actually test. The force of recognition can open an inquiry while leaving its explanation to be examined.

The framework claims that truth becomes actual through living recognition between sensor and instrument. These probabilities model neither experience nor recognition. They challenge a shortcut from a compelling selected pattern to a justified general claim. A distinct CE contribution would have to improve how a living inquiry handles that passage beyond the statistical practices already available.

Where this model stops

One fixed 65% signal; exchangeable 50% nulls; independent items; a specified shared difficulty sign; conditional independence across candidates; uniform selection ties; valid binomial null tests; one independent fresh confirmation batch. This excludes persistent null ability differences, unknown or misspecified dependence, trained candidate generation and reuse of the final confirmation results. It establishes neither a best rule at equal total budget nor a measured advantage for CE.