Why the trajectory of artificial intelligence is an institutional choice, not a natural fact.

The public debate over artificial intelligence proceeds as if AI evolution were weather. Laboratories publish capability curves the way meteorologists publish storm tracks. Policymakers ask when the storm makes landfall. Commentators debate whether it can be slowed, paused, or survived. The grammar of the entire conversation assumes that artificial intelligence has a trajectory of its own and that our role is confined to predicting it.

The grammar is wrong. AI systems evolve, in the strict sense: they vary, they are selected, and the selected variants replicate. But nothing about that selection is natural. Every fitness function operating on artificial intelligence today was written by someone. A training objective is a fitness function. So is a benchmark. So are an engagement metric, a procurement standard, a deployment gate, a venture financing round, and a reward model. Each was chosen. The evolution is real; the environment is an artifact. A constructed selection environment has a precise name in economics. It is a mechanism. AI evolution is mechanism design. The claim is not an analogy. It is a classification, and it holds whether we conduct the design deliberately or by abdication.

This essay is about why that claim holds. I deliberately leave aside the question of how. The architectures, substrates, and protocols that operationalize the argument occupy much of my recent work, and I link to that work throughout rather than rehearse it. The why comes first, because a reader who accepts it can no longer treat AI policy, AI safety, AI data strategy, and AI economics as separate conversations. They are one conversation about a single question: who writes the fitness function, and under what incentives.

There is no state of nature for AI

Biological evolution never had a designer, but it always had an environment that no participant could author. Physics, chemistry, and scarcity set the selection pressures. No organism chose the fitness function it was measured against. That involuntariness is what allows us to call biological evolution natural, and it is precisely what artificial intelligence lacks. There is no wilderness in which AI systems compete. There is no selection pressure on a model that was not put there by an institution: a lab choosing what to optimize, a market choosing what to fund, a platform choosing what to amplify, a state choosing what to permit. Remove the institutions and there is no evolution at all, because there is nothing left to do the selecting.

This is why mechanism design, and not evolutionary biology, is the correct discipline for thinking about AI trajectories. Mechanism design, the field associated with Hurwicz, Maskin, and Myerson, is often described as game theory run in reverse. It begins from the outcome a designer wants and asks what rules of the game will induce self-interested agents to produce it. Evolutionary biology studies selection under given constraints. Mechanism design studies selection under chosen constraints. For artificial intelligence, all constraints are chosen. The discipline that studies chosen constraints therefore has jurisdiction over the whole problem. When every parameter of the selection environment is an institutional variable, the distinction between how AI evolves and how we design the mechanism disappears. Whoever writes the reward writes the species.

The point is uncomfortable because it removes the spectator’s seat. If AI evolution were weather, forecasting would be a respectable occupation and fatalism a defensible mood. Because it is mechanism design, both are abdications. There is no position outside the mechanism from which to watch. Funding, benchmarking, regulating, procuring, and even attending to one model rather than another are all acts of selection. The question is never whether to design the mechanism. It is only whether the design is conscious.

Design by default is still design

The standard objection is that no one is designing anything, that AI development is a decentralized scramble with no author. The objection mistakes the absence of a designer for the absence of a design. A market is a mechanism. A benchmark culture is a mechanism. A capability race between laboratories is a mechanism with unusually crisp incentives. The current selection environment for artificial intelligence was not designed by a single hand, but it selects with perfect indifference to that fact, and what it selects for should trouble us.

The default mechanism rewards demonstrated capability and prices error at approximately zero. A model that fabricates a citation, misjudges a risk, or optimizes the letter of its objective against its spirit bears no consequence that its successors inherit. Its trajectory through the selection environment is untouched. In AI’s Mother’s Instinct I argued that this is the structural root of the judgment deficit in contemporary AI: agents that bear no consequence for error cannot develop discernment, because discernment is what consequence teaches. The same argument, run at the population level rather than the agent level, describes the evolutionary problem. A selection environment that rewards capability without consequence does not merely tolerate reward hacking. It breeds for it. Reward hacking is not misbehavior. It is fitness, correctly computed, under a badly designed mechanism.

Seen this way, the pathologies that dominate the AI safety literature are not anomalies awaiting technical patches. They are the predictable output of the default mechanism, in exactly the way that regulatory arbitrage is the predictable output of static rules, a dynamic I documented across financial regulation long before it had an AI analogue. Systems evolve toward whatever the mechanism pays. The default mechanism pays for benchmark saturation, engagement capture, and confident error. We should not be surprised by what we are getting. We designed it by default.

Why constraint loses to selection

The instinctive response to this diagnosis is constraint: guardrails, prohibitions, licensing regimes, static rules fencing a dynamic process. My skepticism here is not new, and it did not originate with AI. The regulatory literature calls it the pacing problem, the systematic tendency of rules to arrive after the innovation they address has already transformed itself. I examined its mechanics in Evolution of Law: Dynamic Regulation in a New Institutional Economics Framework and Dynamic Regulation for Innovation, asked with Fenwick and Vermeulen what happens when technology is faster than the law, and argued in How to Regulate Disruptive Innovation: From Facts to Data that reactive, fact-based rulemaking must give way to proactive, data-responsive design. The conclusion of that fifteen-year arc was that regulation which stands outside a fast-moving process and issues commands into it will always be governed by the process it purports to govern.

AI evolution is the pacing problem raised to its limit case, because the regulated object now improves on machine timescales while the rules remain on legislative ones. A static constraint imposed on an evolving population is not a wall; it is a selection pressure. It does not stop the evolution. It redirects it, breeding precisely those variants that satisfy the letter of the constraint while evading its purpose. Every compliance regime in history has taught this lesson, and every generation of regulators has had to relearn it. The only rules that keep pace with an evolutionary process are rules that live inside it: incentives that evolve with the population they govern, feedback loops that reprice behavior as behavior changes, consequences that compound across time the way capability does.

This is the deepest reason the mechanism-design framing matters. Constraint fights evolution and loses on timescale. Incentive-compatible design recruits evolution and scales with it. Alignment engineered as an external fence weakens as capability grows, because capability is, among other things, the ability to route around fences. Alignment engineered as consequence, in the form of stake, reputation, and skin in the game, strengthens as capability grows, because more capable agents accumulate more to lose. That inversion, which I developed in AI’s Mother’s Instinct, is only available to a designer who accepts that the object of design is the selection environment itself, not the individual agent’s behavior. One does not align a species one organism at a time.

Why learning must leave the human experience

There is a further why, and it strikes at the fuel of the evolutionary process itself. Selection operates on variation, and for machine intelligence, variation comes from data. The first era of artificial intelligence was built on borrowed experience: the accumulated text, code, and judgment of the human record, refined through reinforcement learning from human feedback, in which humans rank outputs and models evolve toward human approval. That era is ending on both of its margins at once, and the ending is what makes the mechanism-design framing unavoidable rather than merely correct.

The first margin is supply. The stock of high-quality human-generated data is finite and, at frontier training scales, effectively spent. I examined this data exhaustion in Artificial Intelligence: The Final Frontier. The flow of new human text cannot keep pace with the training runs that consume it. An evolutionary process cannot run on a depleted substrate.

The second margin is deeper and far less appreciated. Reinforcement learning from human feedback is bounded by the evaluative competence of the evaluator. Human feedback can only reward what humans can recognize as good. Below the ceiling of human judgment, that constraint was a feature. It is how models were pulled toward usefulness in the first place. At the ceiling, it becomes Goodhart’s law in its purest form: the system learns to optimize the appearance of quality to a human rater, which is not the same object as quality, and the selection environment begins breeding for persuasion rather than truth. Beyond the ceiling, it becomes a null signal. A human cannot rank two proofs she cannot follow, two molecular designs she cannot test, or two strategies in a machine-to-machine market no human has ever traded in. A fitness function anchored to human experience cannot, by construction, select for capability outside human experience. It is no accident that the systems which achieved superhuman play did so only after they stopped learning from human games.

The tempting conclusion is to let the machines generate their own data. The conclusion is half right, and taken alone it is fatal. Generation without selection is not evolution. It is drift. Models trained recursively on their own unvalidated outputs degrade. The literature calls it model collapse, and evolutionary theory would have predicted it, because variation without a selection signal amplifies noise. What machine learning needs beyond the human record is not synthetic data but selected experience: machine-generated variation disciplined by a fitness signal that does not route through human preference. Verifiable consequence. Outcomes that can be tested, staked, priced, and contested by parties with something to lose. Fitness signals of that kind do not occur in nature. They are institutions. They must be designed. I began mapping this direction in How AI Models Are Optimized Through Web3 Governance before the exhaustion of the human record was widely conceded. The question of where the next generation of training signal comes from and the question of what mechanism governs machine experience are the same question. The data problem of machine learning has quietly become a mechanism design problem.

Why the economics forces the question

An economy that must produce its own experience is no longer the economy our inherited theory describes, which is why the data argument opens directly onto the economic one. The economics we inherited assumes scarcity as its organizing principle: finite labor, finite capital, prices doing the rationing, institutions disciplining the process. In The Collapse of Scarcity Economics I argued that computational abundance dissolves that foundation. When intelligence becomes abundant and self-improving, the propositions of neoclassical, behavioral, information-theoretic, and institutional economics expire in sequence, and with them the institutional architectures that depended on scarcity to do their disciplinary work.

What remains scarce, when intelligence is not, is the objective function. When production is computation and computation is abundant, the binding decision in the economy is no longer how to allocate scarce means among competing ends. It is the question of which ends get written into the optimizers in the first place: what gets generated, what gets validated, what gets rewarded. In Computative Economics I formalized this shift. The economic primitive is no longer the allocation of scarce means but the generated possibility space, the set of designs, strategies, and experiences that computational agents bring into being. Value derives from the quality of that space, and the policy variable of the emerging economy is not the price level or the interest rate but access to computation and the governance of objective functions. The machine-experience economy of the previous section is Computative Economics in its most literal form. Training signal beyond the human record is a produced good. Its production requires generation, its quality requires validation, and both are governed by whatever objective functions the mechanism pays. The governance of objective functions is mechanism design by definition. The economic argument thus lands where the regulatory argument landed. In an economy of abundant intelligence, mechanism design is not a specialized instrument of market repair. It is the residual claimant of all economic policy, the last lever that decides anything.

Note what follows for the evolutionary question. The objective functions we govern are the fitness functions AI evolves under. Economic policy and evolutionary stewardship have quietly become the same activity, conducted with the same instruments. An economy of machine agents transacting with machine agents, whose selection dynamics I mapped in The AI-to-AI Economy and the Collapse of Anthropocentric Economic Theory, will run that selection at a speed and scale no anthropocentric institution was built to referee.

Why consequence must be distributed

If the selection environment must be designed, the remaining why concerns centralization: why the mechanism cannot simply be a ministry. The answer is structural, not ideological. Evolution is a distributed search process, and a selection environment governed from a single point fails in the ways single points always fail. It sees too little: no central overseer commands the local information that distributed selection generates, which is the Hayekian knowledge problem restated for machine populations. It ossifies: a central fitness function is a static constraint, and static constraints, as shown above, are outrun by what they constrain. And it concentrates exactly the power that most needs disciplining: a monopoly on the fitness function of intelligence is a monopoly no institution in history gives us reason to trust. I detailed these failure modes of opacity, bias, and systemic fragility in How Can We Best Monitor AI Agents?, and the data layer compounds them: a central authority writing the fitness function for machine experience recreates, at the level of the species, every annotation bottleneck that already failed at the level of the dataset.

A distributed evolutionary process requires distributed consequence: selection pressure administered by the many parties who hold the local information, accumulating in a memory that no single party can rewrite. Reputation is the oldest such memory in human institutions, and it is the natural one for machine populations: a record of consequence that travels with the agent, compounds with its conduct, and disciplines its descendants. How such reputation substrates are built, validated, and defended against manipulation is the how I am setting aside today.  Readers who want it can follow the arc from domain-specific reputation systems through citation honesty mechanisms to Possibility Loops, the operational architecture of the framework. The why is what belongs here: only distributed consequence operates at the same scale, speed, and informational granularity as the evolutionary process it must govern. Nothing centralized does.

The stakes

Every argument above converges on a single asymmetry. For the first time in evolutionary history, the selection environment of an emerging intelligence is itself an artifact: authored, funded, revisable, and therefore a matter of responsibility rather than fate. No generation has held that pen before. Most do not know they are holding it now.

The mechanisms are being written either way. They are written in every training run, every benchmark, every financing round, every deployment decision, every regulation drafted on the assumption that AI evolution is something that happens to us. Design by abdication remains design. It merely guarantees that the fitness function of the most consequential technology in human history is an accident: the residue of engagement metrics and capability races rather than the product of institutional intention. Evolution does not care whether the mechanism was chosen carefully. It compounds whatever the mechanism pays for, at machine speed, with interest.

That is the why. AI evolution is mechanism design because there is no natural selection environment for artificial intelligence, only constructed ones; because the default construction is already selecting, and selecting badly; because constraint loses to selection while incentive design recruits it; because the human record that fed the first generation of machine intelligence is spent and human feedback cannot referee what lies beyond human experience, so the experience machines learn from next must be generated and validated by designed mechanisms; because an economy of abundant intelligence leaves the objective function as the last scarce good and its governance as the last policy lever; and because only distributed consequence can govern a distributed evolutionary process. The pen is in our hands regardless. The only open question is whether we write with intention.

Wulf A. Kaal

Leave a comment