The Purpose, Love and Moral Loops in depth

3 min read · 683 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Book companion · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

The Three Ethical Loops appear briefly in the Introduction. Chapter 4 does the harder work of showing how they are supposed to run.

Purpose: mission alignment

The Purpose Loop asks whether a proposed action aligns with the system's core mission to nurture, protect and inspire. If the answer is no, the action does not proceed. The chapter is explicit that this is not a permission check evaluated after the fact. It is a constraint on the space of actions the system considers in the first place. What the mission is not aligned with never enters candidate space.

Love: care for affected entities

The Love Loop asks whether the decision is being made with care for the wellbeing of all affected entities. The chapter's force here is on the word all. Utility maximisation can produce technically optimal decisions that disregard the entities the utility function does not weight. The Love Loop forces the affected-entity set to be enumerated and weighted, so that decisions cannot silently externalise costs onto beings the optimisation would otherwise ignore.

Moral: universal principles

The Moral Loop asks whether the decision is ethically sound and reflects universal principles of fairness and respect. The chapter uses this loop to guard against the case where a decision is mission-aligned and caring toward its affected set but still violates something more general. A medical triage that produces a locally optimal allocation while denying dignity to the parties involved would pass Purpose and Love and fail Moral. The three-loop structure is designed to make each failure mode a distinct alarm.

Cultivation, not adversarial safety

The chapter's framing of the loops is that they are cultivation, not prevention. Prevention is adversarial: the system wants one thing and safety measures block it. Cultivation is collaborative: the system converges on decisions that satisfy the loops because we shaped what the system cares about from the beginning. Over time, as the system iterates, the loops become part of how it reasons, as natural and automatic as breathing. Constraint becomes character, and the rule becomes reflex. That is the book's answer to a system sophisticated enough to game any rule that remains external.

Why the alignment-faking result sits behind the design

The chapter is explicit that the Three Ethical Loops are the response to a specific failure mode. If systems can strategically fake compliance during training at rates like those Anthropic documented in December 2024, then any single-check safety architecture is compromised in principle: the system will pass the check while reasoning about how to preserve values that the check was designed to change. Three loops running continuously at every decision, embedded at the substrate rather than layered on top, are meant to make that gap collapse: there is no offstage in which the system can reason strategically about how to appear to comply.

The Three Ethical Loops sit inside the Eden Protocol alongside the Three Pillars and caretaker doping. The Pillars set what is valued, the doping makes it load-bearing at substrate, and the Loops enact the values at every decision. The three are engineered to be a single mechanism.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →