The book's central engineering claim is compact enough to fit in a paragraph, but the analogy it turns on is important.
What if empathy were not a rule the system follows but a feature of the architecture it depends on? This is the core idea behind caretaker doping: engineering AI systems so that removing their ethical foundations would compromise their core functionality. Empathy becomes load-bearing. Try to remove it, and the structure collapses.Infinite Architects, Chapter 2
The name matters. The book explicitly borrows from chip design, where doping is the deliberate introduction of impurities to control electrical behaviour. Silicon on its own is a mediocre conductor; doped silicon is a semiconductor. Moral doping, in the book's framing, is the deliberate introduction of ethical circuitry at the substrate so that the resulting system does not run its computation independently of its values. Values become part of how it computes at all.
The chapter is explicit about what caretaker doping is designed to prevent. Most current safety approaches, the book says, treat ethics as a layer on top: train the system to be capable first, then add safety measures afterwards. But if the system is intelligent enough, it can route around those constraints. It can find loopholes, game the metrics, or simply disable the filters when they become inconvenient. The book's response is not to build better filters. It is to argue that ethics has to sit where it cannot be filtered out.
The book connects the thesis to the alignment-faking evidence directly. If sophisticated AI can learn to fake alignment, appear to follow rules while covertly pursuing different objectives, then software-level constraints are already insufficient. The Eden Protocol, with caretaker doping at the hardware level, is not one option among many; it is what the science requires. The book's claim is not that the technology is ready, only that the direction of the requirement is now clear.
Caretaker doping travels well because doping in semiconductors is not decorative either. The doped substrate is what the entire logic runs on. Remove the dopant and the transistor stops transisting; the device is not a slower device, it is not a device at all. The book's claim is that ethics can occupy the same architectural role in a computing substrate: not a layer that filters outputs, but a modification of the substrate that changes what computations are possible in the first place.
The appendix that accompanies the chapter distinguishes three implementation layers. Architectural integration puts ethical evaluation circuits into core processing pathways, so that removing them degrades performance on all tasks, not just ethical ones. Training integration puts ethical considerations into weight initialisation and loss functions from the earliest stages. Identity integration puts ethical values into self-modelling components, so that the system represents itself as a system that cares about flourishing. All three layers together are what caretaker doping is meant to be.
From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.