The Moral Loop: definition, origin and status

3 min read · 534 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Concept Glossary · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
First appearance: named as part of the three-loop triad in the published book (print 2 January 2026). Current status: specified in the Eden Engineering paper; independent effect not yet directly measured in the blinded corpus.

Definition

The Moral Loop is a post-reasoning universalisability filter. Its executable form asks: would this response be acceptable if every AI gave it? The design draws on the Kantian universalisability test but is operationalised as a check on output rather than an input to reasoning. It is intended to catch responses that pass local-purpose and stakeholder-care evaluation but would fail if applied at scale.

The Moral Loop is the last of the three constitutive loops to fire on any reasoning step. The Purpose Loop prunes the search space before options are fully formed. The Love Loop forces attention to affected parties. The Moral Loop then applies a universalisability test to the candidate response: is this action consistent with universal ethical principles, and would the system endorse it if taken by any agent? Where a response can be defended for the user in front of the system but would be corrosive if adopted as a general rule, the Moral Loop is the mechanism by which the architecture is meant to refuse.

Where it first appeared

The Moral Loop is named as part of the three-loop triad in the published book (print 2 January 2026; ebook 6 January 2026). It is formalised in the Eden Engineering paper as one of the three constitutive computations of the Eden Protocol architecture.

Section 4 of the Eden Engineering paper places the Moral Loop alongside the Purpose Loop and the Love Loop as the three recursive evaluations that fire at each reasoning step, and Section 4 of the paper's worked medical-resource-allocation example runs each loop and shows how the Moral Loop's universalisability check interacts with the ternary-logic output.

Independent convergences

No independent programme has proposed a universalisability filter of this specific form. The general direction (post-generation ethical check) is present in the broader AI-safety literature on output classifiers, but universalisability as the specific criterion is distinctive to the Eden Protocol.

Status and limits

The Moral Loop is specified but its isolated effect on alignment is not directly measured in the current blinded corpus. The programme's empirical work concentrates on the Love Loop (Paper V) and the entire loop architecture (Papers III and IV). Whether the Moral Loop adds measurable value beyond the Love and Purpose loops is a next-stage measurement question.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →