Wulf A. Kaal

Release Control Is Not Execution Control

A response to Yukun Zhang, Kemu Xu, and Yishen Chen, "How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents" (arXiv:2609.20474)

Abstract

Zhang, Xu, and Chen separate two sources of value in an agent harness. Planning information can improve task success. A terminal verifier can reduce the acceptance of invalid outcomes. Across 265 matched experimental cells, fixed task-specific plans improved oracle-verified success by 7.17 percentage points. A read-only verifier rejected 61 percent of invalid Retail episodes, but also withheld 17 percent of correct episodes. These are useful empirical results. They also expose a governance boundary. The verifier acts after execution and may leave state changes in place. It controls release, not the underlying effect. Kaal's institutional framework therefore qualifies the result. Verification should be placed before consequential action when reversal is unavailable. Post-execution review remains useful, but it cannot restore authority after an unauthorized mutation.

Plans and verifiers create different value

The authors compare fixed, task-specific plans with shuffled policy text matched in word count. Across 265 matched cells, fixed plans improved oracle-verified success by 7.17 percentage points. The reported 90 percent task-clustered bootstrap interval runs from 1.15 to 13.36 points. Gains were concentrated in more complex tasks. The result supports a limited proposition. Useful planning information can improve stateful agent performance under the tested conditions.

The terminal verifier creates a different benefit. It rejected 61 percent of oracle-invalid Retail episodes for less than one cent of additional cost per episode. It also withheld 17 percent of correct episodes. Whether this trade is valuable depends on the loss assigned to an erroneous release. At higher liability, the verifier's avoided false passes matter more. A standalone verifier captured nearly all of that benefit at about one twelfth of the incremental cost of the combined stack.

A terminal check can arrive too late

The verifier is read-only and runs after the agent has executed the task. The paper notes that state changes may remain in place. This is not a defect in the experiment. It is a boundary on the claim. The verifier can decide whether to release or accept an outcome. It cannot guarantee that an invalid action never altered the environment.

That distinction is decisive for institutional design. Authority drift occurs when locally valid steps accumulate into conduct outside the user's authorization (Kaal 2026, claim 7314479-015). A terminal verifier may recognize the bad aggregate after the consequential steps have already occurred. Detection is valuable. It is not prevention.

The appropriate boundary depends on reversibility. A draft report can often be withheld after generation. A payment, credential disclosure, public filing, or destructive file operation may not be reversible. For those actions, the decisive check must occur before the system holding the capability produces the effect.

Evaluation must remain conditional

Kaal argues that evaluation should derive from observed results under stated conditions, remain specific to context, and resist manipulation (Kaal 2026, claim 7314479-024). The paper largely follows that discipline. It reports oracle-verified outcomes, separates false acceptance from withheld correct work, and shows how the preferred design changes with the assumed liability.

The evidence should not be extended beyond its conditions. The Airline pilot contains only six tasks. Some cells are missing. The experiments use public benchmarks, and the exact endpoint revisions and settings are not fully specified. The authors also disclose that the study was not independently timestamped before the results were observed. These limits do not erase the findings. They keep the findings tied to the tested sample and configuration.

From release control to accountable execution

A governance architecture should distinguish three stages. Planning can shape the proposal. A pre-effect boundary can determine whether the proposal remains within authority. A terminal verifier can assess whether the completed result satisfies the release criteria. Each stage answers a different question.

The execution boundary should produce a durable artifact. A third party should be able to determine which policy governed the action and whether the constraints were enforced without gaining access to the underlying content (Kaal 2026, claim 7314479-048). That receipt records the actual boundary decision. A verifier's later score records a separate evaluative judgment.

The paper's contribution is strongest when read narrowly. It provides empirical evidence that planning information and release control have separable value. The institutional extension is that high-liability actions need more than terminal review. They need execution control held by a component that the agent cannot persuade, bypass, or retrospectively repair.

References

Kaal, Wulf A. 2026. "Institutional Requirements for Sovereign Local Agent Runtimes." SSRN. SSRN 7314479.

Zhang, Yukun, Kemu Xu, and Yishen Chen. 2026. "How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents." arXiv:2609.20474v1. arXiv:2609.20474.