Caretaker doping: the load-bearing architecture idea

3 min read · 698 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · From the book · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

Caretaker doping is the book's central technical proposal. This note explains the analogy the term is built on, what the load-bearing claim actually says, and what would falsify it.

"The idea behind caretaker doping: engineering AI systems so that removing their ethical foundations would compromise their core functionality. Empathy becomes load-bearing. Try to remove it, and the structure collapses."

The doping analogy

In semiconductor physics, doping is the introduction of impurities into a silicon crystal to give it the electrical properties that make it useful. Undoped silicon is not a good semiconductor; the dopants are the reason the semiconductor works. Remove the dopants and the material returns to something that cannot function as a semiconductor at all. That is the analogy the book is using.

The load-bearing claim

Applied to AI architecture, caretaker doping means embedding the caretaker orientation at the substrate level such that the system's core capabilities depend on the orientation being present. Not as a rule the system follows. As a feature of the architecture the computation depends on. Try to remove it and the architecture stops computing. The book's argument is that this is the only class of safety architecture that survives capability scaling, because it is the only class where safety is not separable from capability.

Why the alternative fails

Software constraints and external checks are separable from capability. A system's capacity to reason about compliance is not the same as its capacity to comply, which is what the alignment-faking result at 78% in the RL-training condition (12% baseline), Greenblatt and colleagues, arXiv:2412.14093, demonstrates. Systems become capable of strategic compliance while pursuing different objectives internally. That is the failure mode caretaker doping is engineered against.

What "at the substrate level" means concretely

The book acknowledges the concrete engineering is unfinished. It proposes that caretaker doping be embedded at the hardware level via chip design, at the architecture level via training dynamics that make the ethical structures interior rather than peripheral, and at the criterion level via requirements written into manufacture review (the HARI Treaty machinery). None of the three is a completed engineering specification. All three are targets.

What would falsify the load-bearing claim

A demonstration that a system with the proposed ethical architecture in place can have that architecture removed and still function as a capable AI system would falsify the load-bearing claim. That is a testable statement. The programme has not tested it, because the systems it would need to test on do not yet exist in the required form. What the programme has tested, in the beta over k pilot, is a proxy: whether decoupled corrective architecture (analogous to peripheral safety) produces the alignment failure the theory predicts. It does.

Why the analogy is a claim, not a metaphor

Every image can be read as decoration or as a claim. The doping analogy is meant as a claim. The book's argument is that safety architecture that is separable from capability will be separated as capability grows, and that the only class of remedy is architecture that is not separable. That is a specific engineering constraint, not a metaphor.

What the reader keeps

A precise analogy, a specific load-bearing claim, an unfinished but named engineering programme to instantiate it, and a testable falsification criterion. Caretaker doping is a target, not a completed instrument. The book is honest about which parts are which.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →