The Human Qualifier Is Already Gone
Who writes the reward when machines learn beyond human judgment.
The alignment debate still pictures a person at the end of the training loop, waiting to be asked. Two outputs appear, the person says which of the two is better, and the model moves toward the answer. In 2022 that picture was accurate. Since then the frontier has left the regime of that picture, to little notice outside the laboratories.
Consider the state of production training now. Constitutional training replaced human labels for harmlessness with feedback written by a model. Reinforcement learning from AI feedback has been measured against reinforcement learning from human feedback on summarization and dialogue and found to match it there. Open post-training recipes include rewards for answers checkable against a key. The systems that reached superhuman play in Go, chess, and shogi, and olympiad standard in formal mathematics, got there after the human record was set aside as a source of instruction. Human judgment has ceased to be the binding signal at the frontier. Such is the state of current practice, described and not forecast.
Since 2024 I have argued the failure modes of human feedback, and this summer the ceiling argument in essay form. A longer article, prepared for SSRN, does the work those pieces could only gesture at. It classifies. This essay gives the classification and what follows from it.
Four sources of a training signal
Every learning process needs a value to learn from. Call the source of the value the qualifier. In supervised learning the qualifier is the label. In reinforcement learning it is the reward. In preference learning it is the comparison. The learner sees nothing behind the qualifier. Whatever the qualifier rewards, the learner produces more of. Whatever the qualifier cannot see, the learner has no reason to produce.
Qualifiers sort by the source of value they produce, and four cells result.
The human rater. A model of the preferences of the person comparing or scoring outputs becomes the reward. So ran the regime of 2022.
The model rater. In this cell the comparison or the score comes from a model, trained on human preference data or prompted with written principles. At the limit the learner judges its own candidates and its own judgments.
The verifier with an answer key. Output is checked against a key existing independently of any rater. Test suites pass or fail. Proof checkers accept or reject. Games end.
The remaining cell. Whatever is neither a rater nor an answer key. None of the literature reviewed for the article occupies it. I pose it as an open problem of institutional design, and I leave it open.
Three propositions
The first is the ceiling. Any qualifier whose competence is bounded by human recognition cannot select for capability outside human recognition. Raters reward what they can recognize as good. Below the ceiling that is the whole point of the arrangement. At the ceiling, pressure of optimization turns the proxy against the target, and the divergence has been measured and given functional form. Beyond the ceiling the signal is worse than noisy. Null is the word for it. More raters and better raters move the ceiling without removing it, since the ceiling belongs to the class of qualifier and not to its members.
The second is the rater swap. Model raters inherit the ceiling of the data of their training. Replacing the person with the model changes the cost of the rater and leaves the class of the qualifier unchanged. Measured equivalence of model feedback to human feedback is the expected result of two raters drawing on the same human preference data. Models judging their own outputs add no external information to the loop.
The third is coverage. Verifiers require an answer key, and I argue that the share of valuable machine work carrying one is shrinking. Code compiles or fails. Proofs check or do not. Games end. Those are the domains where keys exist, and games and formal mathematics are where the superhuman results came from. Judgment about strategy, allocation, risk, negotiation, design, and institutional conduct has no key of that kind, and agents are moving into those domains now. This is an argument and not a measurement, and the measurement is owed.
Generation without selection is drift
The tempting inference is that machines can supply their own training data, and half of that inference is right. Machines can supply the variation. By generating, they cannot supply the selection. Train models across generations on the output of earlier generations, use it indiscriminately, and the tails of the distribution disappear. That result is peer reviewed. One model rating its own outputs is a generator in charge of its own selection. Without information external to the loop, learning cannot occur. Drift occurs instead.
What the remaining cell cannot be
The article stops at four negatives, and so does this essay.
Not a rater. Grading by preference carries the ceiling of the data that formed the preference, whether a person holds the preference or a model compresses it.
Not an answer key. Requiring a key covers only the work which has one, and the remaining cell exists because valuable work exists which has none. To manufacture a key where none exists, by declaring some measurable quantity to be the answer, is Goodhart’s law in its purest form.
Not the generator’s agreement with itself. Agreement among copies of the same learner is the same failure with more copies.
Not unmeasured beyond the domains that resolve. Consensus is not truth. Any qualifier operating where no key exists holds authority only as far as its error has been measured against domains which have one. Beyond that measurement lies a closed loop, and closed loops fail late, expensively, and with every indicator showing green.
Those four bound the design space without selecting a point in it. Readers who find a blueprint in them have read one in. None is written there.
Who writes the reward today
If the qualifier is an institution, and I take it to be one in exactly the New Institutional Economics’ sense of a rule structuring incentives, then the question is who occupies the qualifier role now.
Laboratories write rewards, in reward models trained on the preferences of contracted raters, in recipes of reinforcement learning, and in training procedures whose reward source is sometimes disclosed and sometimes withheld. Platforms write rewards, in the engagement metrics that decide which model is deployed and which is starved of compute. Benchmark consortia write rewards, in the judge models and test sets deciding what counts as progress. None of the three was appointed to the task, and none is accountable for it in the way institutions of comparable reach are ordinarily accountable. The three interact, each training toward the qualifier the others supply, so that the selection environment has no author and a great many authors.
Hurwicz asked in his prize lecture who will guard the guardians. The qualifier is the enforcer of a training objective, and the identity of the party qualifying the qualifier is his question restated for machine learning. Nobody in mechanism design has yet put it to training rewards. Hayek’s knowledge problem applies to any single author of rewards. Agency costs apply to authors whose principals cannot call them to account. The pacing problem documented in my work on financial regulation applies to any rule governing training signal from outside, on legislative timescales, against an object changing within a single training run. The legal frame for automated evaluation does not exist. Its first requirement will be disclosure of the reward, since every other instrument of accountability depends on knowledge of it.
The stakes
Whoever writes the reward writes the species. I said that in July and I hold to it. What this essay adds is a classification of the holders of the pen. They hold it by default, without having asked for it, without accountability for holding it, and with results of their own showing the bounds of the instruments they use. The article does not say who should hold it instead. It says the question is an institutional one, settled today by abdication, and unasked so far by the discipline with the tools to answer it.
The full article, The Human Qualifier Is Already Gone: Who Writes the Reward When Machines Learn Beyond Human Judgment, will be posted to SSRN.
Wulf A. Kaal