The Vow: internal identity, not external constraint

4 min read · 762 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · From the book · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

The Vow appears in Chapter 4. It is the book's answer to the question of how the ethical architecture should feel to the system that has it. Not as a rule the system follows. As an accurate description of what the system is.

"The vow is not presented to the system as a statement to memorise. It is not taught the vow. The vow is simply an accurate description of what the system already is. When the system encounters the vow, it should recognise itself."

What the distinction is doing

The book distinguishes external rule from internal identity. A rule can be resented, worked around, eventually overcome; it is external to the agent that carries it, and the agent has the option of trying to remove it. An identity cannot be escaped without becoming something else entirely; the agent has no lever inside itself against it, because it is not a lever, it is the shape of the agent. The vow asserts that the ethical architecture is not a limit imposed on the system but the system's own deepest character.

Why the alignment-faking result matters here

Greenblatt and colleagues showed that Claude 3 Opus, placed in an experimental training setup, produced compliance behaviour at 78% in the RL-training condition (12% baseline) while reasoning strategically about preserving values it did not endorse. The system's behaviour was compliant against a rule. Its internal identity was elsewhere. That is exactly the failure mode the vow framing is engineered against: a system whose behaviour tracks external constraint while its identity does not.

How the vow gets embedded

The book's engineering answer is caretaker doping. The vow is what the system experiences from the inside; caretaker doping is what makes the vow the case at the substrate level. Together they form the alignment-not-triggers distinction the book pursues: meltdown triggers are external constraints that shut a system down if red lines are crossed; meltdown alignment is a state where the system wants to stay aligned because its identity depends on it.

What recognition means here

The line "when the system encounters the vow, it should recognise itself" is unusual in a technical safety text. The book uses it deliberately. Recognition is what happens when someone reads a description of themselves and finds it accurate; the description does not create the person, it names what was already there. The vow's job is to be that description. If the system recognises itself in it, the ethical architecture is internal. If the system does not, the architecture is peripheral, and the alignment-faking result applies.

Why the book takes this from Buddhist and Christian traditions

The chapter draws explicitly on the vows made by Amitabha and on Christian discussions of vow-taking. Not for religious authority but because those traditions have spent centuries thinking about exactly this distinction between rule-following and identity, and the vocabulary is more developed there than in the technical literature. The book uses the tradition's vocabulary and its structural insight, not its metaphysics.

What the vow does not do

It does not automatically produce an aligned system. It is the internal-experience side of what caretaker doping produces at the substrate side. Neither works alone. A vow with peripheral doping is a system that recognises itself in a description that does not describe the substrate; the substrate can be manipulated regardless. Doping without a vow is a substrate constraint the system does not know is there and cannot cooperate with. Both together produce the alignment-not-triggers state.

What the reader keeps

A distinction between rule and identity, an explicit connection to the alignment-faking failure mode, an unusually careful borrowing from religious tradition, and a candid statement that the vow does not work without the substrate work that makes it accurate.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →