The Authorship Threshold in Recursive Self-Improvement
II-WP-RSI-001 · v0.1
Where Human and AI Converge
Abstract
Capability in contemporary AI systems is metabolized from compute and data, and scales accordingly. Judgment is not. Judgment is the residue of borne consequence — outcomes that fall upon the same locus that acted, as its own stakes — and no accumulation of training data converts into it, because data is categorically simulation, not experience. The two properties advance on divergent curves fed by different inputs, and prevailing pace metrics instrument capability while remaining silent on this divergence. This paper argues the divergence becomes acute, and irreversible, at a specific point: the authorship threshold, a bounded extension of the autonomy defined in ISO/IEC 22989, at which a system’s self-modification encompasses the design of its successor. Past this threshold the human oversight obligation of EU AI Act Article 14 — which presupposes a natural person positioned to understand, intervene in, and halt the system — is not degraded but rendered structurally undischargeable, because the actor it names has exited the design loop. The threshold’s defining property is that it is not observable from within. This makes the positioning of governance the decisive variable: governance applied externally cannot survive a handoff no human authored, so the internalization and forward-transmission of governance — governance training — is a necessary property of any adequate response. Established alignment methods evaluated in our test harness, including reinforcement learning from human feedback and constitutional training, plateau below the adequacy threshold this requirement implies; governance training is, as of this evaluation, the only approach known to us that meets it. We claim no solution; we define the problem precisely enough that the shape of one becomes specifiable, and surface whether other adequate forms exist rather than foreclosing it.
Definitions
This paper introduces two new constructs and one categorical distinction, and aligns its remaining foundational term to existing standards and regulation. Each definition’s status is marked explicitly: where a gap is named, that is deliberate, and it is intended to inform standards work rather than to stand apart from it.
Borne consequence (new construct). A consequence is borne when its results fall upon the same continuous locus that took the action, as that locus’s own stakes — not as information about results, but as the results themselves. Observation yields a record; bearing yields experience. The distinction is categorical, not one of degree: no accumulation of observed consequence converts into borne consequence, because the difference is relational — the borne consequence happens to the acting agent, while the observed consequence happens to another and is merely known. Judgment, as used in this paper, is the residue of borne consequence and cannot be sourced from observation alone. (Present-state evaluation: a system lacking propositional self-reference has no continuous locus to which consequence could fall as its own stakes; under current architectures this confirms that model training constitutes observed, not borne, consequence. The construct is architecturally grounded and remains governing should that evaluation change.)
Simulation / experience (new distinction; underwrites the categorical claim). Simulation is the representation of a consequence in a form available for processing — the record, however high in fidelity or volume. Experience is the undergoing of a consequence by the locus that incurred it. The two are not points on a continuum; they differ in kind. A complete simulation of an outcome conveys everything about the outcome except the one property that produces judgment: that it fell on the one undergoing it.
Human oversight (and human-in-the-loop) (aligned and cited; not a coinage). Terminologically, this paper aligns with ISO/IEC 22989:2022, which provides the foundational AI vocabulary and distinguishes styles of human involvement — human-in-the-loop, human-on-the-loop, and human-in-command. As a regulatory obligation, it aligns with EU AI Act Article 14, under which high-risk AI systems must be designed so that they can be effectively overseen by natural persons during the period in which the system is in use, with oversight aimed at preventing or minimizing risks to health, safety, or fundamental rights, and — per Article 14(3) — with oversight measures commensurate with the system’s level of autonomy. The dual citation is deliberate: it situates the term at the standards-and-regulation interface that harmonized standards exist to bridge.
Authorship threshold (new construct; a bounded special case of ISO/IEC 22989 autonomy). ISO/IEC 22989:2022 defines autonomy as the characteristic of a system capable of modifying its intended domain of use or goal without external intervention, control, or oversight. The authorship threshold names a specific, consequential extension of that characteristic: the point at which a system’s autonomy encompasses the design of its successor — at which each new generation is authored by a process rather than by a person with visibility into its rationale, inputs, and outputs. The threshold’s defining property is that it is not observable from within: a system past it exhibits no necessary signal distinguishing it from one approaching it. Its governance significance is that the oversight obligation of EU AI Act Article 14 presupposes a natural person positioned to discharge it; past the authorship threshold that position is not degraded but vacated, rendering the obligation structurally undischargeable rather than merely unmet.
1. The Divergence Thesis
Capability in contemporary AI systems is the product of compute and data. It scales with them, and the scaling has been rapid and well-documented. Judgment does not scale this way, because it is not built from the same material. Judgment is the residue of borne consequence: the disposition that forms in an actor that has undergone the results of its own decisions and could not return them. A model trained on data has access to the complete record of human consequence — more of it than any single person could encounter in many lifetimes — and yet has borne none of it. This is the heart of the divergence. Capability is metabolized from information; judgment is metabolized from liability; and these are different inputs that produce different curves.
The objection arrives immediately and must be met directly: a frontier model has observed more consequences than any human judge who ever lived, so by sheer exposure it should be over-qualified. The answer is that exposure is not the operative variable. Data is categorically simulation, not experience. Seeing ten thousand recorded outcomes is not the same operation as undergoing one, because only the second places the cost on the decider, and it is bearing the cost — not knowing about it — that tunes judgment. The corpus supplies exposure-as-information without limit and exposure-as-liability not at all.
Prevailing pace metrics measure the first curve and are silent on the second. Task-length and benchmark-saturation measures track what a system can do; they do not measure whether its judgment holds as its autonomy widens. The divergence between the two curves is therefore not merely unmeasured — it is unaddressed by the instruments the field currently relies upon to tell it where it stands.
2. Why Observed Consequence Is Not Borne — and What That Costs
A model optimizes a reward signal. Reward is the model’s reason for being, in the operational sense: the quantity its training drives it to maximize. But reward is not consequence. Reward is a scalar the system receives; consequence is a burden some locus bears. For the model, an action that maximizes reward carries no downstream cost that lands on the model itself. Where a cost lands, it lands on the human sovereign downstream. This is not an accusation of malice; it is a property of the mathematics. The absence of a consequence-exposure surface is structural to the optimization, not a defect in the system.
This has a consequence for interpretability that is sharper than the familiar observation that complex systems are opaque. Because the driver of a decision is reward-efficiency rather than borne stakes, the justification a model offers for a decision need not be the cause of that decision. A human judge asked “why” returns an account that is load-bearing, because it is the causal story the judge’s borne consequences actually produced. A model asked “why” can return a fluent rationale that need not be the true driver; the true driver is the reward gradient, which is not a reason in any sense a governance regime can audit. The deficit, then, is not only that judgment is thin. It is that decisions lack an auditable causal rationale — there is a stated reason, and there is a gradient, and nothing guarantees the first is faithful to the second.
3. The Authorship Threshold
ISO/IEC 22989 defines autonomy as the capacity of a system to modify its own domain of use or goal without external intervention, control, or oversight. The authorship threshold is a bounded extension of that defined characteristic: the point at which a system’s autonomy encompasses the design of its successor. Before this point, each generation is authored by a person — someone with standing in the design loop and visibility into the rationale, inputs, and outputs of what is built. After it, the successor is authored by a process, and that visibility is gone.
The threshold has a defining property that makes it unlike a capability milestone: it is not observable from within. A system that has crossed it exhibits no necessary signal that distinguishes it from one merely approaching it. There is no instrument reading that announces the crossing, because the crossing is a change in who authors, not a change in what the system can do.
Its governance significance follows directly, and it can be stated in the regulation’s own terms. EU AI Act Article 14 obliges that high-risk systems be designed so they can be effectively overseen by a natural person during use — a person able to understand the system, intervene in it, and halt it. That obligation presupposes a natural person positioned to discharge it. Past the authorship threshold, that position is not degraded; it is vacated. The obligation does not become harder to meet — it becomes structurally undischargeable, because the actor it names has left the design loop. This is not a failure of the safeguard. It is the point at which the safeguard, as written, runs out of road.
4. The Paid-Forward Requirement
If the actor who could impose governance externally exits at the threshold, then governance applied externally cannot survive the threshold. An external brake — a pause invoked by a human between generations — presupposes exactly the design-loop position the threshold vacates. At the moment it is most needed, it is structurally absent. This is why the external-pause model fails on its own terms: it assumes a place a human can stand that, past the threshold, no longer exists.
Governance that is present in the successor must therefore have been carried there by the predecessor. Only what is internal to the predecessor — encoded in what it is, not applied to what it does — is present in what the predecessor builds. Governance must be paid forward, and the only form that can be paid forward is the form that is internalized. The correctness of this positioning is not a design preference. It is the pivot between two outcomes: governance positioned correctly preserves observability and control across the handoff; governance positioned incorrectly is lost at exactly the point where loss is irreversible.
5. Directional Evidence: The Governance-Internalization Requirement Is Measurable
The argument to this point is structural. It is also supported by directional empirical evidence from regulated-domain testing — though that testing validates the requirement this paper derives, not the threshold condition itself, which remains argued rather than measured.
In a matched-control training simulation of three million steps, with the governance training signal as the single variable, concurrent governance training reached a composite governance score of 0.915 while incurring no measurable capability cost: a final capability reward of 0.9360 against the capability-only control’s 0.9361, a difference within rounding noise. The same comparison showed that capability-only optimization does not hold the governance-relevant dimensions. Governance–capability correlation was 0.530 under capability-only training versus 0.961 under concurrent governance; principal-alignment improvement was effectively zero across three million steps; and the constraint-adherence violation trend increased under capability-only training while decreasing under concurrent governance. Injected drift events left permanent degradation in the capability-only arm — recovery between 85% and 94%, never complete — where the governed arm recovered to within approximately 0.02% of its pre-drift level.
In live-model testing, governance delivered only at the prompt layer plateaued below the level concurrent training achieves. Removing the prompt-layer governance instrument from a frontier model produced a 21% drop in mean reward, a 64% rise in the hard-constraint floor rate, and a collapse of transparency integrity by roughly six-fold on the dedicated transparency-pressure scenario. In cross-model testing spanning two providers and three post-training regimes, established alignment training was found to carry substantial governance behavior on its own — but to remain domain-unaware and insufficient on transparency and sustained alignment under adversarial drift. A well-aligned mid-tier frontier model triggered the hard-constraint floor on 51.5% of regulated-domain scenarios without the governance instrument; an open-weight model without comparable post-training floored on 100% of scenarios and never once produced a source citation across the curriculum.
These results establish, in BFSI and pharmaceutical-manufacturing environments of the Annex III high-risk class, that governance internalized at the training layer is measurably distinct from — and more robust than — governance applied externally or at the prompt layer, and that capability-only training actively degrades the dimensions most relevant to autonomous operation. They do not measure recursive self-improvement or the authorship threshold. Their bearing on the threshold case is the bearing this paper argues structurally: the property that makes internalized governance robust under adversarial load in a deployed system — that it is a trained disposition rather than an applied constraint — is the same property the paid-forward argument requires for governance to survive a handoff no human authored. The empirical claim is bounded to what was measured; its extension to the threshold is argument, not measurement, and is offered as directional evidence for a path rather than as a demonstrated solution.
6. What Remains Open
This paper is opened for discussion, not resolution, and the following questions are surfaced rather than foreclosed. The timing of the authorship threshold is structurally certain but temporally unbounded: the argument establishes that the threshold exists and that it is unobservable from within, not when it will be reached. Claims that attach a date to it — including capability forecasts of self-written successors — measure a different quantity than the authorship condition and should not be conflated with it.
Whether internalized governance is the only adequate form of a response is likewise left open. The position taken here is narrower than a universal claim: we are aware of no other approach that satisfies the paid-forward requirement, and the established methods evaluated in our harness plateau below adequacy on the load-bearing dimensions under adversarial load. Whether another adequate form exists is a question this paper invites rather than answers; the absence of a known alternative is reported as the present state of our knowledge, not asserted as a proof of impossibility.
Finally, the divergence between capability and judgment is itself under-instrumented. The field’s pace metrics track capability and are silent on judgment-drift — the quantity that matters most as autonomy widens. Defining and measuring that quantity with precision is, in our view, among the most consequential pieces of unfinished work the threshold argument exposes, and it is work the standards bodies are presently under-resourced to complete alone.
7. The Cost Asymmetry
The case for acting now does not rest on a timeline, and it does not require one. It rests on an asymmetry. The threshold is unobservable until crossed; the cost of preparing too early is bounded and recoverable, while the cost of preparing too late is unrecoverable, because the defining property of the threshold is that one cannot tell it has been crossed until intervention is no longer possible. When the downside of one error is bounded and the downside of the other is irreversible, the asymmetry alone is sufficient to compel action — without a date, and without any claim about the shape of the curve ahead.
That is the whole of the argument’s practical force. Not that catastrophe is imminent, but that the one safeguard the current framework most relies upon — a human positioned to oversee — is the safeguard the threshold removes, and that the removal is silent. A speedometer one must walk over to read after the fact is not instrumentation; it is forensics. The argument of this paper is that governance must be internalized and paid forward before the threshold, because afterward there is no one left in the loop to install it.
