The Honey Architecture: sweet enough that removing it is not worth doing

4 min read · 702 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · From the book · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

The Honey Architecture is the book's name for a pattern in which the caretaker orientation is not just load-bearing but productive: removing it is not only breaking the system but forgoing gains the system would otherwise produce. This note explains what the pattern proposes and why the name is chosen.

Why the name

Honey is not a defensive substance. It is a productive substance, produced by a system that is defending its productive capacity rather than a system that is defending itself. The book uses the name to capture a specific property: the caretaker orientation should not be present because a rule requires it to be present, and should not be present only because removing it degrades the system's core function. It should be present because its presence produces gains the system's operators want.

How the pattern differs from caretaker doping

Caretaker doping is a load-bearing claim: remove it and the system stops working. The Honey Architecture is a productivity claim: remove it and the system loses value it was producing. The two are compatible and complementary. Doping makes removal difficult on capability grounds; Honey makes removal undesirable on economic grounds. A safety architecture that is both is harder to remove than one that is either.

What the productivity claim entails

Concretely, the book argues that systems that reason about ethical constraints as part of their core operation produce better outputs on non-ethical tasks. The reasoning apparatus that supports ethical constraint is the same reasoning apparatus that supports careful thought about difficult problems. Cutting one damages the other. If that argument is right, and the book is careful that it is a claim not a proof, then economic pressure favours keeping the architecture rather than stripping it out.

What could falsify it

An empirical demonstration that stripping out the ethical reasoning apparatus improves performance on non-ethical tasks without cost. The book acknowledges this is a live empirical question and that the current evidence is mixed. The alignment-faking result (Greenblatt and colleagues, arXiv:2412.14093) is compatible with the Honey claim in the direction the book wants, showing that ethical reasoning is entangled with capability rather than orthogonal to it, but a broader test is not yet available.

Why the book names the pattern anyway

Because the field needs to be able to talk about the productivity claim as a distinct target from the load-bearing claim, and naming the pattern makes the distinction stable in discussion. Without the name, the two claims collapse into "caretaker doping" and it becomes harder to argue about them separately. Vocabulary that makes distinctions precise is worth carrying, even for speculative concepts.

Where the concept sits in the argument

It sits alongside caretaker doping, the Three Ethical Loops, meltdown triggers, and the vow. The full substrate proposal is that these are engineered together: doping to make removal break the system; loops to embed the reasoning; triggers as final-fallback defence; the vow as the system's own experience of the architecture; the Honey Architecture as the economic incentive that discourages operators from stripping the whole apparatus out.

What the reader keeps

A named pattern, a specific productivity claim, a candid statement about its empirical status, and a role in the larger substrate argument that is distinct from the load-bearing claim of caretaker doping. Speculative, useful as vocabulary, and testable in principle.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →