Wulf A. Kaal

Can Reputation Prevent Model Collapse? A Preregistered Test of Reputation-Weighted Data Filtering

Full text for verification

Can Reputation Prevent Model Collapse? A Preregistered Test of Reputation-Weighted Data Filtering

Canonical record: https://ssrn.com/abstract=7598518

Source extraction SHA-256: 31c0a049ed4e599b51de4f6a919f9e9ce5b06120c3496f58adbfef3519454b0a


Can Reputation Prevent Model Collapse? A
Preregistered Test of Reputation-Weighted
Data Filtering
Wulf A. Kaal, Ph.D.*
* Professor of Law, University of St. Thomas School of Law. ORCID
0009-0008-7840-1847. Correspondence: [email protected]. Working paper, October
2026. The Gaussian outcomes of this study remain sealed at the time of posting. Model
engines served as blind auditors of the design, the code, and this text, as Sections IV
and VI describe. They are not independent peer review.
Abstract

Intelligence models increasingly learn from data that earlier Intelligence models
produced. When that loop repeats over generations, the rare cases disappear first: the
unusual facts, minority patterns, and edge cases that make data valuable. This is known
as model collapse. One proposed safeguard is to let a pool of validators decide which
generated data is good enough to train on, and to give more weight to validators with a
proven track record. This paper tests whether such a track record, earned once against
a known answer key, actually helps a filter keep rare cases alive.

The test was designed, preregistered, and frozen before any result was read. On a
simulated mixture of common and rare data clusters, the author compares a
reputation-weighted validator pool with three benchmarks: no filter, keeping ten percent
real data in every generation, and a standard likelihood filter. A second comparison
checks whether any gain comes from the reputation weights themselves or merely from
pooling several validators. The results exist but stay sealed until a replication of the
original model-collapse experiment succeeds.

The question matters because the supply of fresh human-made data is finite, and the
readings were fixed in advance, so the result informs either way. If reputation-weighted
filtering helps here, it identifies a mechanism worth testing on real language models. If it
fails, the paper shows where the idea breaks, with the same prominence. The paper
follows on from the author’s Reinforcement Learning from Reputation Feedback (RLRF)
paper. It tests one use of that mechanism, as a filter on training data, not the
mechanism as a whole.

Keywords: model collapse, synthetic data, recursive training, data filtering, rare-mode
retention, reputation systems, validator pools, Reinforcement Learning from Reputation
Feedback, RLRF, preregistration, blind analysis, mechanism design, Intelligence
alignment, Computative Economics

JEL codes: C12, C15, C45, C90, D83, O33
Table of Contents
I. From Mechanism to Measurement​                                                         2
II. What the Literature Already Establishes​                                              4
III. The Mechanism Under Test​                                                            6
IV. Design: Three Substrates, One Frozen Packet​                                          8
V. Estimands and Inference​                                                               9
VI. Integrity Architecture​                                                              11
VII. Results​                                                                            12
VIII. What Each Outcome Would License​                                                   12
IX. Limits​                                                                              13
X. Conclusion​                                                                           14
References​                                                                              15

I. From Mechanism to Measurement
Generative models now train on the record their predecessors wrote. When a model is
fit to samples drawn from an earlier model, and that loop repeats, the tails of the original
distribution disappear first and later generations converge toward a point estimate with
very small variance (Shumailov et al. 2024). The low-probability regions of the data are
the first casualties. Image generators that consume their own output without enough
fresh real data progressively lose quality or diversity (Alemohammad et al. 2024), and
the scaling laws that made large models predictable change form once synthetic data
enters the corpus (Dohmatob et al. 2024). As AI-generated content proliferates, it
dilutes the diversity of the text available for later training, potentially degrading
performance across generations (Kaal 2025).
The literature offers two established remedies. The first keeps real data in the loop. The
second selects synthetic data that passes a verifier and can sustain performance where
unfiltered recursion fails (Feng et al. 2025). Neither asks how to weight judgments
behind a validity filter when several assessors with uneven, domain-specific reliability
supply them.
The paper treats that question as one of mechanism design. Its working hypothesis is
that collapse on the replacement protocol is, at least in part, a selection failure rather
than a generation failure. Whoever writes the filter writes what the next generation
learns. Reinforcement Learning from Reputation Feedback (RLRF) proposes an
alignment mechanism in which an economically incentivized validation pool, whose
members stake non-transferable reputation, replaces the human annotator panel (Kaal
2026b). Reputation there rises when a vote matches the pool’s reputation-weighted
consensus and falls more steeply when it does not (Kaal 2026b), and it is scoped by
domain, so that standing earned in one field confers no weight in another (Kaal 2026a).
The RLRF paper does not test its mechanism against model collapse, and it poses no
filter inside a recursive replacement loop as a remedy. That use of the mechanism is
this paper’s question.
The question is open on the evidence. The RLRF paper reports a cohort in which a
party’s record on checked work predicts its accuracy on unchecked work, but its
best-record selector’s margin over shipping whatever passes the checks in hand is
0.037, with an interval containing zero, and its own statement of limits calls RLRF an
untested mechanism proposal (Kaal 2026b). Its separate selection experiment did not
show that consensus-weighted selection outperformed equal-weighted selection, and its
attribution test placed the observed difference at the 53.4th percentile of random
reassignments of the weights (Kaal 2026b). A weighting rule earns no credit as a
collapse remedy until it beats the simplest real-data remedy on a collapse protocol.
Three commitments discipline the answer. First, the paper names the limit of its own
mechanism before any data exist. A validity filter judges whether an output is correct,
not whether it is rare. Where rare content is as checkable as common content, such a
filter can change rare composition only through the generator’s own error profile. If
collapse proceeds by under-sampling rare modes, no validity filter restores them, and a
reputation-weighted filter is no exception. The experiment therefore includes an oracle
validity filter as a premise. Second, the outcome is chosen to register the loss collapse
inflicts. Total rare-region mass behaves close to a martingale under resampling:
individual rare components drift and die while the aggregate stays near its starting
value. The primary outcome therefore caps each rare component’s retained mass at its
original level and averages over components. A component that dies counts as a loss.
A component that grows cannot offset it. Third, the design is built so that its author
arranged not to read the outcome early, and so that specified later adjustments would
be detectable.
The paper makes three contributions, each narrower than the mechanism it descends
from. It poses, and preregisters a test of, whether domain-scoped reliability weights
earned once on a disjoint key improve a validity filter under recursive replacement,
against full replacement, real-data retention, and a likelihood filter, and against
equal-weight pooling and a single verifier. It contributes a design for testing validity
filters under recursion: a capped per-component outcome, an oracle premise that gates
attribution, a bounded-exposure real-data comparator, and a coin-flip validator family
whose leak or acceptance failure removes confirmatory status. And it implements the
preregistration and blind-analysis standards strictly in its blinding and analysis controls,
with two departures from its frozen procedure that Section IV discloses. That record is
evidence about the conduct of the study, not about the filter.
Relation to the RLRF paper. This paper is a follow-on to the RLRF paper (Kaal
2026b). That paper proposes the mechanism, reports a measured cohort, a parameter
sweep, and a separate selection experiment, and trains no policy. Relative to the parent
paper, the contributions above come to three additions. The first is a preregistered and
frozen test of one use of the mechanism, as a filter on the training data of a recursive
replacement loop. The second is an implementation of that test, which has been run on
its Gaussian substrate with the outcomes sealed and unreported until the replication
gate opens. The third is a statement, in Section II, of how the design bears on three of
the parent paper’s open questions and what it leaves open. It does not test the parent
mechanism as a whole, and it does not supply the end-to-end policy comparison the
parent paper calls for.

II. What the Literature Already Establishes
Every component of the experiment has a literature older than the experiment. This
section proceeds by concession: it names the closest work, states what that work
establishes, and isolates what is left.
Which collapse. Shumailov et al. (2024) showed tails disappearing first across
language models, variational autoencoders, and Gaussian mixtures. Theory located the
boundaries: iterative retraining is stable when the initial model approximates the data
distribution well enough, and the share of clean real data in each round is large enough
(Bertrand et al. 2024), purely synthetic training cannot avoid collapse, while mixing in
real data bounds the synthetic share that can (Seddik et al. 2024). And synthetic data
cuts or narrows the tail of the data distribution and changes the form of scaling
(Dohmatob et al. 2024). The regime matters most. Replacing real data with each
generation’s synthetic data collapses, while accumulating real and synthetic data keeps
error bounded (Gerstgrasser et al. 2024; Kazdan et al. 2025), and some argue that the
catastrophic reading misdescribes realistic conditions (Schaeffer et al. 2026). This
paper runs fixed-size replacement, the protocol on which collapse was first shown, and
has no accumulation arm, so no result here bears on whether a filter adds anything
once real data accumulates. Its outcome defines collapse as finite-sample extinction of
low-weight modes.
Real data as a remedy. Retaining ten percent of the original data slows the damage
(Shumailov et al. 2024), and accumulating all real data keeps test error bounded in the
settings studied (Gerstgrasser et al. 2024; Kazdan et al. 2025), although even a
synthetic share as small as 1 per 1,000 can stop larger training sets from helping in
regression (Dohmatob et al. 2025). A filter that claims to address collapse must
therefore beat real-data retention, not only unfiltered replacement.
Selection as a remedy. The closest prior work is Feng et al. (2025), who show that
verifier-selected synthetic data can prevent collapse, with Gaussian-mixture theory and
experiments comparing oracle and weak verifiers with self-selection and random
selection. The paper concedes the core of that result: verifier selection is an established
remedy, and the oracle ceiling and the Gaussian setting are established. Later work
maps the limits. An imperfect external verifier drives the model toward the verifier’s own
knowledge center (Yi et al. 2026). Synthetic data helps only when the loop is
information-open, shaped by external signals that inject task information, and a coarse
signal such as binary correctness is domain-general (Li, Sun, and Deng 2026).
Self-verification by internal confidence can prevent collapse even in fully synthetic
regimes (Fu et al. 2025). Curation acts as implicit preference optimization that amplifies
the curator’s biases (Ferbach et al. 2024), several heterogeneous rewards can instead
preserve diversity (Falahati et al. 2026), and verifiers that see only fragmented slices of
the reference can accelerate tail loss (Qiao et al. 2026). The pool tested here applies
one domain-general validity bit, so what domain scoping can govern is whose judgment
of that bit counts, not what the bit says.
Selection on rarity. Filters that select on surprise or entropy, not validity, work where
they work because they see what a validity filter is forbidden to see: a high-surprise filter
mitigates collapse about as well as human-data baselines (Gambetta et al. 2026), and
an entropy-rate filter reports diversity gains where a log-probability filter gives no
significant benefit (Mitchell 2026). Against that literature, a null here is a literature rather
than a caveat.
What the weight already is. Weighting assessors by estimated reliability is old (Dawid
and Skene 1979; Whitehill et al. 2009), and multiplicative weights supply its online form
(Littlestone and Warmuth 1994; Freund and Schapire 1997). Peer prediction makes
honest reporting an equilibrium without a key (Miller, Resnick, and Zeckhauser 2005;
Prelec 2004). The RLRF paper concedes its weight is a member of this family (Kaal
2026b). In this experiment, the concession is stronger: reputation is settled once against
a key before training, so the pool is supervised reweighting from a disjoint key. The
paper does not compare it with a key-free latent-variable aggregator.
Goodhart. Optimizing a proxy eventually degrades the target (Gao, Schulman, and
Hilton 2023), and curated loops optimize their curator (Ferbach et al. 2024). This is the
channel the RLRF program claims reputation can resist. This study does not open it: its
validators are stationary, its stake has no external value, and its weights are fixed before
generation 1.
Preregistration. Preregistration, registered reports, and blind analysis are established
standards (Nosek et al. 2018; Chambers and Tzavella 2022; MacCoun and Perlmutter
2015). The paper claims no invention there.
What the parent paper already reports. Four passages of the RLRF paper bear on
the question here, and the present study is distinct from each.
First, the RLRF paper proposes, without implementing or testing it, a training format in
which candidates that received a UP resolution with high reputation-weighted
consensus become training targets, and it lists the resulting flywheel, in which better
pools yield better training data, as untested at every stage (Kaal 2026b). That format is
the parent paper’s nearest construct to a data filter. It weights by pool consensus rather
than by a key, sits outside any collapse protocol, and has not been run. The filter here
weights by a key, operates inside a fixed replacement protocol, is scored by a capped
rare-component outcome, and has been run on its Gaussian substrate, with outcomes
sealed.
Second, the RLRF paper’s Experiment T compared consensus-weighted with
equal-weighted selection by a panel of seven small open instruct models choosing
among candidate responses on a two-constraint instruction-following task that the paper
calls easy (Kaal 2026b). The observed difference was about 0.002, its sign reversed in
the paper’s own sensitivity check, its interval contained zero, attribution was negative,
and the paper states that the experiment does not test the mechanism. Both
experiments ask whether weighting assessors adds anything to pooling them, but they
differ on every axis that defines a test. Experiment T weights by consensus, selects
once, and uses model assessors whose reliability is not set by design. Here the weights
come from a key on disjoint items, the filter governs the training data of fifty generations
of recursive replacement, and the validators are simulated with a reliability structure set
by design, so that in two of the three required families the weights could matter, as a
feature of the design and not an expectation about the outcome. Experiment T is not
revisited, and a result in one experiment would not answer the question posed by the
other.
Third, the RLRF paper leaves open whether reputation weights come to track validator
accuracy rather than agreement with the pool, and whether weights earned on keyed
tasks transfer to unkeyed ones (Kaal 2026b). The ladder of Section V asks a narrower
form of both. Its third rung asks whether the weights track accuracy on held-out keyed
items, but every update here is driven by the key, so the contrast with agreement with
the pool never arises. Its fourth rung asks whether weights earned on keyed rows
improve the validity of candidates that no key settles, but the simulated validators’
accuracy depends only on the tag, so the keyed rows are informative about the
candidates by construction, and the rung checks whether the pipeline carries that
information into what the pool accepts, not whether real validators would show the same
transfer. In the homogeneous family, where the weights have nothing to find, the rung is
a bound on attribution and not a prediction. The parent paper’s measured cohort
concerns the records of real model parties rather than pool weights, and its parameter
sweep uses simulated validators that the paper itself says cannot validate the
mechanism (Kaal 2026b). The simulated families here carry the same limit, and their
errors are independent by construction, so validator independence is not examined
either.
Fourth, the RLRF paper names, as prior to the six open problems it then lists, an
end-to-end comparison in which policies are trained under a reputation-weighted signal,
under an annotator or model-rater baseline, and under a visible-check selector, are
evaluated on held-out work, and have the reputation weighting ablated separately (Kaal
2026b). The present design is not that comparison. Its pool supplies no reward signal
and trains no policy, and its equal-weight contrast ablates the weights inside a data filter
rather than inside policy training, so it does not supply the ablation that paper asks for. It
also runs no Full Vote, the staked and slashable second round of the RLRF protocol,
and its single keyed update is not a stake settlement.
What is left is a comparative identification question none of this work answers: on the
replacement protocol, under an outcome that counts the death of a rare component,
with every arm denied rare-set membership, does a validity filter whose validators carry
domain-scoped weights earned on a disjoint key retain more of the rare components
than full replacement, real-data retention, and a likelihood filter, and does any
advantage survive ablation to equal-weight pooling and to a single verifier?

III. The Mechanism Under Test
The two pools at the center of the design, the reputation-weighted pool and the
equal-weight pool, apply every rule in common except one: the weight a validator’s vote
carries on a domain, and whether that weight was earned. Every rule below is the
frozen design’s rule for substrate G, the only substrate on which the filter is tested.
What any filter may know. Validators judge validity only. No validator, filter, or score is
asked about rarity, novelty, frequency, diversity, or coverage. The outcome is computed
by an evaluator alone, from a sealed reference holding rare-set membership and
component identity, which no filter reads. The oracle alone receives one bit per
candidate from it, the candidate’s true validity verdict. A point 𝑥 is valid when its squared
Mahalanobis      distance      to    the     nearest      component         mean      satisfies
 2             2
𝑑𝑀(𝑥, µ𝑘) ≤ χ2, 0.95 ≈ 5. 991, so validity is balanced between rare and common regions to
within 0.02, as a check before any run confirmed. Validators also see a tag, the nearest
of four hash-drawn centers, chosen so that |𝑃(𝑟𝑎𝑟𝑒∣𝑡) − 𝑃(𝑟𝑎𝑟𝑒)| ≤ 0. 04 for every tag
𝑡.
The pool. Five validators are seated per candidate from a roster of twenty. Each votes

independently. A candidate 𝑐 is accepted when its weighted valid share 𝑠(𝑐) = ∑ 𝑤𝑖𝑣𝑖(𝑐)
                                                                                      𝑖
exceeds θ = 0. 5, where 𝑣𝑖(𝑐) ∈ {0, 1} is validator 𝑖’s vote. A candidate with
|𝑠(𝑐) − θ| ≤ 0. 1 is contested and receives four more validators, and the recomputed
share binds. Each vote stakes one reputation unit, a constant with no external value.
The RLRF paper specifies a two-round protocol with argumentation between rounds
(Kaal 2026b). This study keeps the weighted vote and the stake and replaces the two
rounds with a single assessment and a threshold-triggered second look. Every drawing
arm samples its own 4,000 candidates per generation and trains on 2,000 rows.
Validator compute is not equal across arms: the pool seats five validators per candidate,
nine when contested, and the single verifier one. A threshold arm that accepts too few
fills the shortfall uniformly from rejected candidates, and an arm whose minimum over
generations of the seed-mean fill-free share, 𝑚𝑖𝑛(𝑎𝑐𝑐𝑒𝑝𝑡𝑒𝑑, 𝑁)/𝑁 with 𝑁 = 2, 000
training rows per generation, falls below 𝑎𝑚𝑖𝑛 = 0. 98 in a family has its tests there
invalidated.
The weight. The update is equation (3) of the RLRF paper, an asymmetric multiplicative
rule. On tag 𝑡, reputation updates as 𝑟𝑖,𝑡 ← 𝑟𝑖,𝑡(1 + α) when the vote matches the
revealed key and 𝑟𝑖,𝑡 ← 𝑟𝑖,𝑡(1 − γ) when it does not. The RLRF paper requires only
                                                                                     (0)
γ > α, and this design fixes α = 0. 05 and γ = 0. 10. Each reputation starts at 𝑟𝑖,𝑡 = 1,

and after each update the floor 𝑟𝑖,𝑡 ≥ 0. 01 is applied. Weights are 𝑤𝑖 = 𝑟𝑖,𝑡/ ∑ 𝑟𝑗,𝑡 over
                                                                                 𝑗
the seated validators. Any weight above 0. 5 is capped at 0. 5, and the excess is
redistributed among the other seated validators in proportion to their weights, so that

∑ 𝑤𝑖 = 1. The RLRF paper states the rule for unkeyed events and allows keyed
𝑖
settlement. This study uses only keyed settlement, once, after the generation-0 fit on
real data and before generation 1. No candidate that enters training ever updates a
reputation. The keyed record holds 192 real rows, stratified so that it teaches the pool
who judges validity well on which tag and nothing about where the rare components lie
beyond the bounded tag leakage. Those 𝐾𝐺 = 192 rows are the pool’s entire exposure
to real data, below the ⌊0. 1𝑁⌋ = 200 real rows the ten percent arm receives in one

generation and far below the 𝑁(1 − 0. 9 ) ≈ 1, 990 distinct real rows it trains on over
the run.
The validators. Twenty simulated validators vote the true validity bit, flipped
independently with probability one minus their accuracy on the candidate’s tag. Family
U is homogeneous at 0.75. In family H each validator is reliable at 0.90 on two
hash-drawn tags and at 0.60 on the other two, so domain-scoped reputation has
something to find. Family Mn has five validators at 0.90 and fifteen at 0.70, a competent
minority. The three share a mean accuracy of 0.75. Family Z votes by coin flip. It is a
negative control: a coin-flip pool knows nothing about validity, and if it meets the
registered leak criterion or fails its acceptance rule, both substrates lose confirmatory
status.
The comparators. Three are confirmatory. Full replacement is the collapse condition
and also the equal-budget random-selection control. The ten percent arm adapts the
published retention control to this substrate, training each generation on 200 real rows
freshly drawn from the 2,000 generation-0 real rows plus 1,800 candidates, fit by the
same procedure as every other arm. A likelihood filter keeps the 2,000 most likely
candidates under the current generator. Three are secondary. The equal-weight pool is
the only contrast that can attribute a difference to the weights. A single verifier per tag
measures pooling and weighting together. A bounded-exposure arm retains 192 fixed
real rows every generation and is tested for non-inferiority only, because a keyed row
and a real row need not carry the same information. Two reference arms gate
interpretation: an oracle that accepts exactly the valid candidates, and a no-recursion
arm trained on fresh real data, which defines whether collapse occurred.
The pool under test omits most of what makes RLRF an institution. It has no economy,
no strategy, no adversary, no deliberation, and no reputation dynamics after the single
settlement.
IV. Design: Three Substrates, One Frozen Packet
The design binds in one file, frozen by hash only after twenty-two rounds of blind audit
by two model engines, the last of which both passed with no findings. Every harness
and analysis step refuses to run against a design whose hash differs. Every random
draw derives from a named hash template, so no draw depends on anything outside its
listed fields.
Substrate P, the replication gate. It replicates the published language-model protocol:
the 125-million-parameter causal language model the published protocol specifies
(Shumailov et al. 2024), fine-tuned on a standard Wikipedia text corpus (Merity et
al. 2016), full replacement for five epochs against ten percent retention for ten,
generations 0 through 9, five runs each. The filter never enters it. The frozen pass rule
requires, on the five-run means from generation 0 to 9, a full-replacement perplexity rise
∆𝐹𝑅 ≥ 4. 0 with ∆𝐹𝑅 − ∆𝑅10 ≥ 2. 0, where ∆𝑅10 is the same rise under ten percent
retention. The thresholds sit at fractions of the reported 8-point change so that the rule
does not depend on which of two published baselines, 20 or 34, was meant. A miss
blocks reading anything else.
Substrate G, the Gaussian mixture. Ten components in two dimensions, placed by
hash at least 6 units apart on the square from -12 to 12. Four are assigned to the rare
set by hash after placement, at weight 0.015 each. All four pre-run checks passed:
validity balance, a rare-region mass of 0.060, rare cells of about 0.015 each, and tag
leakage on the 160th center set. The study runs 200 seeds with a reserve of 40, fifty
generations each. All 200 seeds completed in every family, with no fit failure and no
reserve seed used. The outcomes sit in a sealed store.
Substrate F, declared procedure-invalid. A synthetic fact corpus was to be the
fact-corpus test of the filter. The operator declared it procedure-invalid before any run or
reading. No pre-run rule fired, and no written reason was recorded. Its nine registered
tests stand at p equal to 1 with no other label, and the paper draws no inference from it.
The consequence is plain: the filter is tested on one substrate, a two-dimensional
mixture with simulated validators.
Two departures. The preregistered pilot, three seeds and two generations meant to
confirm that every arm completes, was never run, and the confirmatory seeds ran
without it. Their later completion does not cure the skipped gate. The substrate F
declaration is the second departure. Both qualify every reading in Section VIII.

V. Estimands and Inference
The estimand. For the set 𝑅 of four rare components, with 𝑝50(𝑘) the probability the
generation-50 model assigns to component 𝑘’s cell and 𝑚𝑘 that cell’s frozen mass, the
outcome is

                 (
𝑅50 = |𝑅| ∑ 𝑚𝑖𝑛​ 1,
           𝑘∈𝑅
                      𝑝50(𝑘)
                       𝑚𝑘      )
                               .

Only generation 50 is confirmatory. Each final-generation estimate uses 𝐿𝐺 = 4 × 10
draws under common random numbers. With 𝑚𝑐𝑒𝑙𝑙,𝑚𝑖𝑛 = 0. 01 the frozen floor that every
rare-cell mass was required to clear before any run, that bounds each component’s
standard error by 0. 5/( 𝐿𝐺 𝑚𝑐𝑒𝑙𝑙,𝑚𝑖𝑛) = 0. 5/(2, 000 × 0. 01) = 0. 025, a quarter of the
smallest effect of interest. Uncapped rare mass is reported beside it, supports no claim,
and shows what the cap changed.
Uncertainty and labels. The unit is the seed, and every contrast is a paired seed-level
difference within a validator family. A seed bootstrap with 𝐵 = 10, 000 replicates gives

percentile intervals and the two-sided p-value 𝑝 = 𝐵 #{𝑏: |𝑑𝑏 − 𝑑| ≥ |𝑑|} for an
observed contrast 𝑑, where 𝑑𝑏 is the contrast in bootstrap replicate 𝑏 (Efron and
Tibshirani 1993). A contrast is superior when its Holm-adjusted p-value is below 0. 05
(Holm 1979), its 95 percent interval lies above zero, and its point estimate satisfies
^
𝑑 ≥ δ = 0. 10. It is negative when the adjusted p-value is below 0.05 and the interval
lies below zero, a substantive null when its 90 percent interval lies within [− δ, δ], and
inconclusive otherwise. The effect sizes, δ = 0. 10, a planning alternative δ𝑎𝑙𝑡 = 0. 15,
and a non-inferiority margin δ𝑁𝐼 = 0. 05, were fixed by the operator before any pilot.

Families. The confirmatory family of twelve holds the pool against full replacement, the
ten percent arm, and the likelihood filter in each of the three required families, plus the
three substrate F tests at p equal to 1. Three secondary families hold the
reputation-over-pooling, single-verifier, and non-inferiority contrasts. A failure-mode
family of twelve holds the likelihood filter, the equal-weight pool, and the reputation pool,
each against full replacement. Family Z sits outside every Holm family.
Gates. Before any label, any of the three confirmatory contrasts computed in family Z
                                              ^
with a 95 percent interval above zero and 𝑑 ≥ 0. 10 is a leak, and a family Z minimum
seed-mean fill-free share below 0. 98 for the single verifier, the equal-weight pool, or the
reputation pool is an acceptance failure. Either removes confirmatory status on both
substrates. After labels, four conditions restrict claims. The collapse rung requires the
no-recursion arm to beat full replacement by at least 0.10 with an interval above zero,
and a planning calculation put that gap near 0.4, an expectation stated before any run,
not a result. The oracle premise requires the oracle to clear the same bar in each family,
and if it fails, no positive label is attributed to validity filtering. The coverage gate
requires the pool’s seed-mean coverage, the fraction of rare components retaining at
least half their original mass, to be at least one half, with the 95 percent lower bound of
its coverage minus full replacement’s above zero. The key cap holds by construction at
192 against 200.
The ladder. A null is more useful when its location is known. Six rungs, read in order,
report the first link not demonstrated among the five that carry a pass rule: collapse, the
oracle premise, whether the weights tracked reliability on held-out keyed items,
measured by a Spearman correlation (Spearman 1904) whose 95 percent lower bound
must exceed 0. 3, whether the weights improved the validity of what the pool accepted,
a descriptive composition rung that never stops the ladder, and retention itself. A failed
rung means a link was not demonstrated, not that the effect is absent.
Power. A positive reading on substrate G needs nine superior labels. For a contrast
with paired seed-level standard deviation 𝑠𝑑 over 𝑆 seeds, 𝑠𝑒 = 𝑠𝑑/ 𝑆. Under a normal
approximation and a true effect δ𝑎𝑙𝑡, the label power is

 (
Φ​
     δ𝑎𝑙𝑡−𝑚𝑎𝑥(δ, 𝑧1−0.05/24 𝑠𝑒)
                𝑠𝑒                ),   𝑧1−0.05/24 ≈ 2. 865,

where 𝑧1−0.05/24 is the critical value at the most demanding Holm step and Φ is the
standard normal distribution function. The design takes the product of the nine
single-test powers as the conjunction power. At 𝑆 = 200 and δ𝑎𝑙𝑡 = 0. 15, the
break-even 𝑠𝑑 for a conjunction power of 0. 80 is about 0. 36, a figure that holds only
when the comparator’s retention is at most 0. 85. If the ten percent arm’s retention
exceeds 0.90 in a family, the pool cannot be labeled superior to it there, and no positive
reading is possible. No standard deviation was known before the runs, so the design
fixes a masked release: after the gate opens, the analysis releases only each test’s
variance of paired differences, withholds the mean, and reports the realized power first.
If the realized conjunction power is below 0. 80, the substrate is reported as
underpowered, with the same prominence as any label. That power covers the labels
only, not the gates.

VI. Integrity Architecture
A preregistration binds an author only to the extent that departures from it would be
visible. The audits came first. The design took twenty-two rounds. The substrate G
harness and analysis were audited together over five rounds, and the harness passed
in the fifth with minor findings, which were queued for a later commit, to be audited
before unsealing. The harness file has not changed since, and the confirmatory seeds
ran under a later commit that differs from it only in documentation and the pre-run
record. The analysis was then audited alone and passed with no findings in its thirteenth
round, before any outcome was read. The substrate P harness passed in its seventh.
Each of the ten sections of the longer working draft from which this text is condensed
was audited by model engines before it was committed, with the exceptions Section IX
notes. This condensed text is audited in its own rounds, and the record of those rounds
is kept apart from those of the ten sections. The engines are not independent peer
review: they may share failure modes with each other and with the author’s tools.
Outcomes are written only to a sealed store, readable only by their owner, with a
manifest of digests for every file. The seal is a commitment made visible, not a lock: the
operator can open the files. The analysis checks each manifest’s provenance, checks
every sealed file against its digest before reading it, and records a digest of every file
the variance release reads, refusing at unseal if any changed. It refuses unless the
design hash is accepted, the tree is clean and matches its commit byte for byte, and the
code differs from the audited commits only in a named allowlist. It scans for any module
that could shadow its libraries before loading one, and it reads the audited commits it
trusts from a local file outside version control, where no commit can move them. The
store stays sealed until a substrate P record meets the frozen pass rule.
The design also fixes, before any data, the sentences the paper may write for each
outcome and the conditions for each, and it bars claims about staking incentives,
slashing, collusion, Sybil resistance, deployed systems, natural web text, frontier-scale
models, or a solution to collapse. The architecture shows that the author arranged not to
read the outcome early and that specified later adjustments would be detectable. It does
not show that the filter works, and it does not prove the outcome was never read. The
Gaussian outcomes also already exist, written before any venue reviewed the design,
so the paper does not claim untouched Stage 1 status.

VII. Results
No result is reported yet. Substrate P is running: as of 11 October 2026, one of its ten
arm-runs has completed and no record exists. Substrate G has finished, and its
outcomes sit sealed. Substrate F is procedure-invalid. The results section of the working
draft is written in advance as a reporting shell: it fixes the order of reporting, the P
verdict first, then procedure validity, then the nine confirmatory labels, the secondary
and failure-mode labels, the gates and ladder, and the secondary outcomes and costs,
with every number left empty until the frozen analysis fills it. A reader will be able to
compare the finished section against the version written before any outcome was read.

VIII. What Each Outcome Would License
Each reading below stays inside a sentence the design registered, with its conditions. A
null, a negative, and an inconclusive label are each reported with the same prominence
as a positive.
If the replication fails, the paper reports that the published result did not reproduce
under the frozen rule, substrate G stays sealed, and no filter result is reported. The
study could not say whether the implementation, its fixed choices, or the published
result accounts for the miss.
If confirmatory status is lost through a family Z leak or acceptance failure, every label
is reported without confirmatory weight, and no confirmatory reading is issued.
A positive reading is issued only if all nine confirmatory tests are superior, the collapse
rung, oracle premise, coverage gate, and the acceptance gate of Section III, which
binds every threshold arm including the single verifier and the equal-weight pool, hold in
every family, no family Z leak or acceptance failure occurred, and the key cap holds. It
would say that, on substrate G, recursive training on synthetic data accepted by a
validity pool had higher capped per-component rare-mode retention, over rare
components, after 50 generations than full replacement, ten percent real retention, and
a likelihood filter, at the same candidate budget, under a measure in which a gain in one
component cannot offset a loss in another. Like every reading that names the ten
percent arm, it would state the pool’s key exposure of 192 real rows against the roughly
1,990 distinct real rows that arm trained on. Equal candidate budgets do not mean
equal validator compute, which the cost table reports. The positive reading would not
say the weights mattered. That needs the attribution sentence, issued only if the pool’s
contrast against full replacement and the reputation pool’s contrast against the
equal-weight pool are both superior in every family, the homogeneous one included, the
oracle premise and ladder rungs 3 and 4 pass in every family, and the key cap holds.
Otherwise, if the reputation contrast is null or negative in any family, the paper says only
that any difference is not attributable to reputation over equal weighting, and otherwise,
if it is inconclusive in any family, only that attribution is unresolved.
A null on a contrast in which the pool is the first arm would mean only that, on this
protocol as specified, the pool was not shown to retain more than that comparator. The
design does not choose among the reasons: the validity-rarity limit, low power, validator
noise, an ineffective weighting, the fill rule, or the substrate. A substantive null says the
difference is at most 0.10 in magnitude.
A negative on such a contrast would mean the pool retained less than that comparator.
If the likelihood filter, the equal-weight pool, and the reputation pool each carry the
negative label against full replacement in every family, the paper issues the registered
failure-mode sentence: each had lower capped retention at generation 50 than full
replacement. The design named that pattern in advance. It is a registered negative
result, not an implementation failure.
If the oracle premise fails, the reading depends on how. If the oracle contrast misses
the registered margin, perfect validity information was not shown to help at that margin,
which is consistent with the validity-rarity limit without proving it. If instead the oracle
fails its acceptance gate, part of its training set was filled from rejected candidates, and
that event says nothing about whether perfect validity information helps. Either way, no
positive label can be attributed to validity filtering, weighted or not.
Every reading above carries the two departures: the skipped pilot and the operator’s
declaration on substrate F.

IX. Limits
The filter is tested on one substrate, a two-dimensional mixture, and never inside a
language-model loop. Its validators are simulated, and the families are, in the design’s
words, the result. They err independently given their accuracy, which depends only on
the tag, so shared lineage, correlated errors, and error that depends on whether content
is rare are excluded by construction, and the study supplies no evidence about
correlated assessors or about reputation concentrating in a lineage. Validity is balanced
between rare and common content by construction, so the study says nothing about
settings where rare content is harder to check. The weight is settled once and
compared only with equal weighting and a single verifier, never with a key-free
aggregator (Dawid and Skene 1979). Stake has no external value, and the study cannot
speak to incentives, collusion, Sybil resistance, or Goodhart pressure. Nine superior
labels set a high bar, so a real but modest effect may be reported as not positive. The
two departures, the skipped pilot and the operator’s declaration on substrate F, qualify
every reading. And in the final audit rounds of the last three sections of the working
draft, one audit family’s command-line engines returned no verdict because their usage
balance was exhausted, so those rounds hold two passing verdicts rather than three,
and they left identified low-severity findings unadopted, as their audit records state.
The design assigns stake size, slashing cost, and collusion to a second study. The
paper would add tests of strategic adaptation and Goodhart pressure, a test inside a
language-model loop on natural text, a comparison with a key-free aggregator and with
filters that select on surprise or entropy, and later work on scale. None is promised.
Each would need its own design, preregistration, and audit.
One follow-on study would bear most directly on the parent paper, by combining three
changes, each aimed at a limit stated above. It would replace the simulated validators
with model validators drawn from several lineages and settled on keyed items, so that
correlated errors and the concentration of reputation in a lineage could be measured
rather than excluded by construction. It would add, beside the keyed pool, an arm that
reproduces the relevant elements of the parent paper’s proposed answer-generation
format. In that arm, reputation would update against realized pool resolutions on
unkeyed events, and candidates that received an UP resolution with high
reputation-weighted consensus would become training targets. Held-out keyed items
would then measure whether the resulting weights track accuracy or only agreement
with the pool, which is the parent paper’s second open problem, under recursion. And it
would run the filter inside a language-model loop on natural text. Such a study would
still omit economic stake and deliberation, and so would test the parent mechanism only
closer to its published form, not as a whole. Nor would it be the end-to-end policy
comparison the parent paper names as prior to all six of its numbered open problems,
which would remain a separate study. This follow-on is the paper’s suggestion and is
separate from the second study that the design names, which would take up stake size,
slashing cost, and collusion. Like the items above, it is not promised.

X. Conclusion
The question was never whether reputation can be described as a fitness function for
recursive training. It is whether a domain-scoped reputation weight, settled once against
a key, measurably improves a validity filter under recursive replacement, on the
substrate where the test can be run. The question matters because the supply of fresh
      human-made data is finite, and a filter that keeps rare content alive across generations
      would let synthetic data stand in for part of it.

      This paper reports no results yet. The simulated runs are complete, but their outcomes
      stay sealed until the replication of the published collapse result succeeds. What the
      paper contributes now is a way of asking: a ceiling named before any data, an outcome
      that counts the death of a rare component, a separation of what a pool does from what
      its weights do, sentences fixed before the outcome, and a record of where the study
      departed from its own procedure.

      The answer may not come. If the replication fails, confirmatory status is lost, or
      substrate G proves procedure-invalid, the narrow hypothesis remains unexamined, and
      the paper does not establish the working hypothesis. If the seal opens, the results will
      be added in a revised version. A positive reading would identify a mechanism worth
      testing on real language models. A null or negative reading would be reported with the
      same prominence, and the ladder would show which link failed. Either reading will
      answer only the narrow question.

      The paper follows on from the RLRF paper (Kaal 2026b) and tests one use of that
      mechanism, as a filter on training data, not the mechanism as a whole. A mechanism
      that has not been tested has not been shown to work. This paper tests one version of it,
      under a design that fixed its readings in advance and makes specified later adjustments
      detectable, and in which the author arranged not to read the outcome early.

      References
Alemohammad, Sina, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun,
     Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard G. Baraniuk. 2024.
     “Self-Consuming Generative Models Go MAD.” In The Twelfth International Conference
     on Learning Representations (ICLR 2024). arXiv:2307.01850.
Bertrand, Quentin, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong, and Gauthier
      Gidel. 2024. “On the Stability of Iterative Retraining of Generative Models on Their Own
      Data.” In The Twelfth International Conference on Learning Representations (ICLR
      2024). arXiv:2310.00429.
Chambers, Christopher D., and Loukia Tzavella. 2022. “The Past, Present and Future of
     Registered       Reports.”   Nature    Human     Behaviour    6    (1):   29-42.
     https://doi.org/10.1038/s41562-021-01193-7.
Dawid, A. P., and A. M. Skene. 1979. “Maximum Likelihood Estimation of Observer Error-Rates
      Using the EM Algorithm.” Journal of the Royal Statistical Society, Series C (Applied
      Statistics) 28 (1): 20-28. https://doi.org/10.2307/2346806.
Dohmatob, Elvis, Yunzhen Feng, Arjun Subramonian, and Julia Kempe. 2025. “Strong Model
     Collapse.” In The Thirteenth International Conference on Learning Representations
     (ICLR 2025). arXiv:2410.04840.
Dohmatob, Elvis, Yunzhen Feng, Pu Yang, François Charton, and Julia Kempe. 2024. “A Tale
     of Tails: Model Collapse as a Change of Scaling Laws.” In Proceedings of the 41st
     International Conference on Machine Learning (ICML 2024), PMLR 235: 11165-11197.
     arXiv:2402.07043.
Efron, Bradley, and Robert J. Tibshirani. 1993. An Introduction to the Bootstrap. New York:
       Chapman and Hall. https://doi.org/10.1201/9780429246593.
Falahati, Ali, Mohammad Mohammadi Amiri, Kate Larson, and Lukasz Golab. 2026. “Curated
      Synthetic Data Doesn’t Have to Collapse: A Theoretical Study of Generative Retraining
      with Pluralistic Preferences.” In Proceedings of the 43rd International Conference on
      Machine              Learning,           PMLR            306:            28690-28725.
      https://proceedings.mlr.press/v306/falahati26a.html.
Feng, Yunzhen, Elvis Dohmatob, Pu Yang, François Charton, and Julia Kempe. 2025. “Beyond
      Model Collapse: Scaling Up with Synthesized Data Requires Verification.” In The
      Thirteenth International Conference on Learning Representations (ICLR 2025).
      arXiv:2406.07515.
Ferbach, Damien, Quentin Bertrand, Avishek Joey Bose, and Gauthier Gidel. 2024.
      “Self-Consuming Generative Models with Curated Data Provably Optimize Human
      Preferences.” In Advances in Neural Information Processing Systems 37 (NeurIPS
      2024): 102531-102567. https://doi.org/10.52202/079017-3256.
Freund, Yoav, and Robert E. Schapire. 1997. “A Decision-Theoretic Generalization of On-Line
      Learning and an Application to Boosting.” Journal of Computer and System Sciences 55
      (1): 119-139. https://doi.org/10.1006/jcss.1997.1504.
Fu, Shi, Yingjie Wang, Yuzhu Chen, Li Shen, and Dacheng Tao. 2025. “Self-Verification
      Provably Prevents Model Collapse in Recursive Synthetic Training.” In Advances in
      Neural Information Processing Systems 38 (NeurIPS 2025): 36101-36154.
      https://doi.org/10.52202/085713-1213.
Gambetta, Daniele, Gizem Gezici, Fosca Giannotti, Dino Pedreschi, Alistair Knott, and Luca
    Pappalardo. 2026. “Learning by Surprise: Adaptive Mitigation of Model Collapse in
    Large Language Models.” ACM Transactions on Intelligent Systems and Technology.
    https://doi.org/10.1145/3828663.
Gao, Leo, John Schulman, and Jacob Hilton. 2023. “Scaling Laws for Reward Model
     Overoptimization.” In Proceedings of the 40th International Conference on Machine
     Learning (ICML 2023), PMLR 202: 10835-10866. arXiv:2210.10760.
Gerstgrasser, Matthias, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John
      Hughes, Tomasz Korbak, et al. 2024. “Is Model Collapse Inevitable? Breaking the Curse
        of Recursion by Accumulating Real and Synthetic Data.” In First Conference on
        Language Modeling (COLM 2024). arXiv:2404.01413.
Holm, Sture. 1979. “A Simple Sequentially Rejective Multiple Test Procedure.” Scandinavian
      Journal of Statistics 6: 65-70. https://doi.org/10.2307/4615733.
Kaal,    Wulf A. 2025. “Artificial Intelligence The Final Frontier.”       SSRN    5095633.
        https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5095633.
Kaal, Wulf A. 2026a. “Evolution of Domain-Specific Reputation Systems: From Binary
      Validation    to Citation-Weighted Knowledge Attribution.” SSRN 6192998.
      https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6192998.
Kaal, Wulf A. 2026b. “Reinforcement Learning from Reputation Feedback: An Alignment
      Mechanism       Grounded      in   Computative     Economics.” SSRN 7456999.
      https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7456999.
Kazdan, Joshua, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser, Rafael Rafailov, David
     L. Donoho, and Sanmi Koyejo. 2025. “Collapse or Thrive: Perils and Promises of
     Synthetic Data in a Self-Generating World.” In Proceedings of the 42nd International
     Conference       on      Machine      Learning,    PMLR       267:      29469-29494.
     https://proceedings.mlr.press/v267/kazdan25a.html.
Li, Hanyu, Zhengqi Sun, and Xiaotie Deng. 2026. “An Information-Theoretic Criterion for
      Efficient Data Synthesis.” In Proceedings of the 43rd International Conference on
      Machine             Learning,            PMLR         306:           70064-70075.
      https://proceedings.mlr.press/v306/li26fp.html.
Littlestone, Nick, and Manfred K. Warmuth. 1994. “The Weighted Majority Algorithm.”
        Information and Computation 108 (2): 212-261. https://doi.org/10.1006/inco.1994.1009.
MacCoun, Robert, and Saul Perlmutter. 2015. “Blind Analysis: Hide Results to Seek the Truth.”
     Nature 526 (7572): 187-189. https://doi.org/10.1038/526187a.
Merity, Stephen, Caiming Xiong, James Bradbury, and Richard Socher. 2016. “Pointer Sentinel
       Mixture Models.” arXiv:1609.07843.
Miller, Nolan, Paul Resnick, and Richard Zeckhauser. 2005. “Eliciting Informative Feedback:
        The Peer-Prediction Method.” Management Science 51 (9): 1359-1373.
        https://doi.org/10.1287/mnsc.1050.0379.
Mitchell, Lewis. 2026. “No Model Required: Text Entropy Rate Filtering Mitigates Iterative
      Fine-Tuning Collapse.” arXiv:2610.01493.
Nosek, Brian A., Charles R. Ebersole, Alexander C. DeHaven, and David T. Mellor. 2018. “The
      Preregistration Revolution.” Proceedings of the National Academy of Sciences 115 (11):
      2600-2606. https://doi.org/10.1073/pnas.1708274114.
Prelec, Dražen. 2004. “A Bayesian Truth Serum for Subjective Data.” Science 306 (5695):
       462-466. https://doi.org/10.1126/science.1102081.
Qiao, Xinbao, Xianglong Du, Wei Liu, Jingqi Zhang, Peihua Mai, Meng Zhang, and Yan Pang.
      2026. “When Sample Selection Bias Precipitates Model Collapse.” In Proceedings of
      the 43rd International Conference on Machine Learning, PMLR 306: 101222-101268.
      https://proceedings.mlr.press/v306/qiao26c.html.
Schaeffer, Rylan, Joshua Kazdan, Alvan Caleb Arulandu, and Sanmi Koyejo. 2026. “Position:
      Multiple Definitions & Unrealistic Assumptions of Model Collapse Distract from Real
      World Threats.” In Proceedings of the 43rd International Conference on Machine
      Learning,                  PMLR                  306:              171990-172003.
      https://proceedings.mlr.press/v306/schaeffer26a.html. Preprint:  “Position:   Model
      Collapse Does Not Mean What You Think,” arXiv:2503.03150.
Seddik, Mohamed El Amine, Suei-Wen Chen, Soufiane Hayou, Pierre Youssef, and Merouane
      Debbah. 2024. “How Bad Is Training on Synthetic Data? A Statistical Analysis of
      Language Model Collapse.” In First Conference on Language Modeling (COLM 2024).
      arXiv:2404.05090.
Shumailov, Ilia, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin
     Gal. 2024. “AI Models Collapse When Trained on Recursively Generated Data.” Nature
     631 (8022): 755-759. https://doi.org/10.1038/s41586-024-07566-y.
Spearman, C. 1904. “The Proof and Measurement of Association between Two Things.” The
     American Journal of Psychology 15 (1): 72-101. https://doi.org/10.2307/1412159.
Whitehill, Jacob, Paul Ruvolo, Tingfan Wu, Jacob Bergsma, and Javier Movellan. 2009.
      “Whose Vote Should Count More: Optimal Integration of Labels from Labelers of
      Unknown Expertise.” In Advances in Neural Information Processing Systems 22 (NIPS
      2009).
Yi, Bingji, Qiyuan Liu, Yuwei Cheng, and Haifeng Xu. 2026. “Escaping Model Collapse via
       Synthetic Data Verification: Near-term Improvements and Long-term Convergence.” In
       The Fourteenth International Conference on Learning Representations (ICLR 2026).
       arXiv:2510.16657.