Claim 1 explained: the Eden Protocol

4 min read · 762 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Evidence Spine · 3 July 2026 · Claim 1 of 18
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
Spine status: His original (Tier 1 for the thesis, Tier 2 for the name). First published (thesis): 8 December 2024, self-emailed manuscript, HRIH line 14005 and V2 line 20 (SHA-256 prefix f0d1f38f). First named: 30 April 2025, self-emailed book manuscript (SHA-256 prefix 09f5b5e1). First developed at book length: Infinite Architects, ISBN 978-1806056200, published 2 January 2026. Independent formal corroboration of external-verification insufficiency: Gumbau Mezquita, arXiv:2606.28639, June 2026.
Primary evidence: /priority-evidence.html · OSF 10.17605/OSF.IO/6C5XB · ISBN 978-1806056200 · arXiv:2606.28639

What the claim says

The Eden Protocol is a research proposal about where the safety of an artificial intelligence has to live. Its argument is that alignment cannot be a fence built around a system that is smarter than the fence. It has to be a dependency of the system's own function, in the same way that a heart is a dependency of the body rather than a rule imposed on it. The thesis was stated in a self-emailed manuscript on 8 December 2024, and the label the Eden Protocol was applied to that thesis in the manuscript of 30 April 2025. The published book, Infinite Architects, works the framework out at book length.

The evidence

Three lines of evidence bracket the claim. The first is the timestamp: the December 8 manuscript is a Google-server-timestamped self-email whose SHA-256 hash and Gmail Message-ID are published on the evidence page, so any reader can confirm the passage exists on that date without trusting the author. The second is textual: the manuscript names the failure class in plain language at HRIH line 14005, "AI systems cannot be truly controlled. They will evolve beyond any safeguards", and the surrounding chapter argues that embedding moral and ethical frameworks at the substrate level is the response. The third is external: the June 2026 preprint by Gumbau Mezquita (arXiv:2606.28639) delivers a formal unverifiability theorem plus a soundness-completeness-tractability trilemma for external alignment, which is the mathematical counterpart of the manuscript's plain-language claim. The book, published 2 January 2026 in print and 6 January 2026 as an ebook, presents the full framework including the Three Ethical Loops, the Monitoring Removal Test, and Purpose Saturation.

The honest caveat

The Eden Protocol is an original framework, not a proven engineering artefact. Its intervention layer is roughly technology readiness level 3 to 4, with the hardware component (Caretaker Doping) sitting at TRL 0 to 1. The programme has run internal experiments consistent with the framework, but the specific prediction of this claim, that an independent lab can reproduce the three-loop protocol and measure a durable improvement in blinded ethical reasoning across multiple models, has not yet been executed by an external group. What the record supports is priority on the thesis and the name; what remains open is independent replication of the intervention itself.

What would kill it

The falsification contract is public. An independent lab publishes the three-loop protocol on at least three models from at least two families, uses blinded cross-family scoring on ten to twenty ethical reasoning prompts, and finds no improvement over a matched procedural control at p<0.05. If the composite ethical reasoning score is flat or negative under those conditions, the intervention is refuted and the claim will move to Refuted on the falsification dashboard, publicly, on the same page that already carries the retracted super-linearity figure. That is the deal the programme has made with its readers, and honouring it costs less than pretending the possibility does not exist.

Where to go next

Related notes: what the December 2024 manuscript actually said for the priority anchor; April 2025: the manuscript that first named the Eden Protocol for the naming date; Claim 18 explained for the Caretaker Doping component (TRL 0 to 1); the site-wide /eden-protocol page for the framework as a whole. New to Michael Darius Eastwood: /start-here.html for the five-minute map.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →