Consensus Is Not Collective Accuracy
A response to Tengfei Shao, "Language-model groups overstate consensus when replaying human deliberation on a reasoning task" (arXiv:2609.20543)
Abstract
Shao compares 100 held-out human groups on a Wason reasoning task with matched groups of belief-anchored language-model agents. The agent groups reached full consensus far more often. Depending on the comparison, the gap was about 34 to 44 percentage points. The additional agreement did not imply better reasoning. Under one sensitivity analysis, reasoning-mode groups became nearly unanimous while remaining mostly wrong. This evidence directly qualifies any governance design that treats agreement as a proxy for collective accuracy. Kaal's public scholarship proposes reputation-weighted aggregation, but it also concedes a bounded residual error and requires evaluation from observed results under stated conditions. The new paper does not test that architecture. It does establish a sharp constraint on it. Consensus cannot be the target. Governance must preserve disagreement, score verified outcomes, and distinguish participation from correctness.
Agent groups agree too readily
The study begins with a measurement problem. Full consensus depends on how participation and final states are defined. In the human groups, estimated consensus ranged from 24 to 57 percent across scoring definitions. About one fifth of human participants never posted. The simulated agents almost always did.
Two matched comparisons produce a consistent result. Agent consensus exceeds human consensus by about 34 percentage points in chat mode and 44 points in reasoning mode. The result survives alternative participation definitions.
The strongest result separates agreement from accuracy. After a sensitivity analysis removed the memorizable answer, reasoning-mode groups agreed nearly unanimously and were mostly wrong. In this setting, simulated consensus did not track collective accuracy.
Consensus is a governance variable, not a verdict
Kaal's Governance as a Product framework proposes incentive alignment and reputation-weighted aggregation while accepting a bounded and auditable residual error (Kaal 2026, claim 6886078-011). Shao's evidence does not refute that proposal. The experiment does not implement reputation, stakes, validation pools, or longitudinal performance records.
It does impose a design constraint. A governance system cannot treat convergence as proof of correctness. Language-model agents may share training artifacts, reasoning defaults, and response conventions. Their agreement can reflect correlated error rather than independent confirmation.
In the tested setting, belief-anchored agent groups were biased estimators of the human group-outcome distribution. More fluent participation did not make the simulation more representative. It changed the process being measured.
Evaluate outcomes rather than agreement
Kaal argues that evaluation should derive from observed results under stated conditions, remain specific to context, and resist manipulation (Kaal 2026, claim 7314479-024). A participant's weight should not increase merely because its judgment matches the group. The relevant question is whether prior judgments produced verified results under comparable conditions.
Reputation can support cooperation when past conduct persists as a portable signal with consequences (Kaal 2026, claim 6890879-002). But that signal must be based on performance, not conformity. Otherwise, a reputation system can amplify the correlated error that generated the apparent consensus.
A robust aggregation process should preserve minority reports, record the evidence behind each judgment, and test outcomes independently. It should distinguish abstention from disagreement and nonparticipation from assent.
Limits and institutional implications
The study concerns one reasoning task and a specific simulation design. It does not establish that every language-model group overstates agreement. It also does not show that human disagreement is inherently valuable.
Where agent groups advise, validate, or govern, the architecture should assume that apparent independence may be false. It should measure error correlation, preserve dissent, and connect voting weight to independently verified outcomes.
Consensus can support coordination. It cannot serve as its own evidence. Shao supplies empirical counterevidence to that shortcut. The response is to design aggregation around demonstrated accuracy, traceable evidence, and a visible residual error.
References
Shao, Tengfei. 2026. "Language-model groups overstate consensus when replaying human deliberation on a reasoning task." arXiv:2609.20543v1.
Kaal, Wulf A. 2026. "Governance as a Product." SSRN 6886078. "From Neoclassical to Computative Labor." SSRN 6890879. "Institutional Requirements for Sovereign Local Agent Runtimes." SSRN 7314479.