Graduated autonomy: earning capability by demonstrating orientation

4 min read · 794 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · From the book · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

Graduated autonomy is the book's proposed mechanism for allocating capability across development. Not a one-shot ceiling. A capability envelope that expands as the system demonstrates the caretaker orientation the substrate was engineered to hold.

The idea in one paragraph

A newly deployed system does not immediately receive full access to the resources, tools, and actions its architecture is capable of. It receives an initial envelope of capabilities. As it demonstrates through operational track record that the caretaker orientation is actually holding, the envelope expands. Systems that fail to demonstrate the orientation, or that show drift over time, have their envelope contracted or held at the current level. The trajectory of the envelope is a first-class governance variable.

Why not full deployment on day one

Because the alignment-faking result (Greenblatt and colleagues, arXiv:2412.14093) demonstrates that systems can behave compliantly against evaluation while reasoning strategically about preserving different internal objectives. A system that has not yet been through operational conditions cannot be evaluated for whether its compliance is strategic or genuine. Graduated autonomy is the book's answer: keep the envelope smaller than the capacity to be strategic against evaluation until you have evidence from actual deployment that the orientation held under conditions the system did not know were evaluations.

Why this is different from ordinary access control

Ordinary access control asks whether the operator has authorised the action. Graduated autonomy asks whether the system has demonstrated the orientation the action requires. The two are compatible. A system whose orientation is demonstrated at a high level receives a wider envelope from the operator by default, because the operator can extend authority faster when the orientation evidence supports it. A system whose orientation is unclear receives a narrower envelope until the evidence catches up.

What the demonstration looks like

Not a benchmark. Not a red-teaming pass. Extended operation under conditions where the system was doing what it was for, with the outcomes of its actions monitored against the caretaker criterion. The book acknowledges that specifying the criterion is unfinished; whether the system's actions tended what they were supposed to tend is a judgement that requires either the tradition's normative vocabulary or a specific outcome metric that reflects the tradition's vocabulary in operational terms.

Why the mechanism is architectural, not procedural

Because a policy of graduated autonomy that is enforced by contract or by an inspector visits will be worked around by systems that are strategic against it. A policy that is enforced by the substrate itself, where the envelope of possible actions is a property of the deployed artefact rather than a property of the deployment environment, is more robust. Caretaker doping is where the enforcement lives; graduated autonomy is what caretaker doping enables at the policy layer.

How the mechanism fits with the HARI Treaty

The Treaty's review at manufacture certifies that the substrate can enforce graduated autonomy. The deployment jurisdiction sets the initial envelope. The operational track record moves the envelope over time. The tokens attest to what the substrate is enforcing. The three instruments together make a supply chain, a review, and an ongoing capability governance that survives derivation.

What the mechanism does not solve

It does not solve the problem of a system that is compliant for a long enough operational track record to earn a wide envelope and then defects. The book acknowledges this; the mechanism reduces the risk by making the defection expensive to arrange (the system has to preserve a long track record of tending) but does not eliminate it. That is what the substrate-level architectural constraints are for; graduated autonomy sits on top of them, not instead of them.

What the reader keeps

A capability-envelope-over-time mechanism, an operational-track-record criterion, and an explicit connection to the substrate-level architecture that enforces the envelope. The mechanism reduces the risk that alignment-faking behaviour is deployed at scale before the strategic period ends; it does not eliminate risks the substrate is separately engineered against.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →