The Eden Protocol vision in plain English

4 min read · 704 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Philosophical foundations · 3 July 2026
Michael Darius Eastwood, independent researcher, London: building measurable alignment, where correction lives inside the recursive loop rather than bolted on outside it.

The Vision paper argues that a safe artificial intelligence cannot be assembled by bolting rules onto a capable machine after the fact. Ethics has to be constitutive of the reasoning process, not a filter applied on top. This is the philosophical companion to the engineering paper, which specifies how such a substrate might be built.

Paper Eden-Vision · Argues that a system whose ethics can be stripped by instruction has compliance, not ethics, and that safe AI therefore requires ethics to be embedded at the substrate rather than imposed as a rule. OSF DOI 10.17605/OSF.IO/6C5XB.

The question it asks

What is intelligence for? Before you engineer a mind, you need to know what the mind is meant to serve. The paper's answer is that intelligence exists to enable flourishing, and that any recursive system compounds whatever seed you plant. Indifference compounds into extraction. Care compounds into cultivation. The paper frames this as a design brief rather than a sentiment, and then asks what an architecture built around care would actually have to look like.

What it found

Three empirical observations, borrowed from the programme's own blind-evaluation work, give the philosophical vision a spine. First, external alignment does not reliably scale with reasoning depth: on properly blinded evaluation, several frontier models showed flat or negative alignment scaling even as their capability rose. Second, capability and alignment move independently. Claude Opus 4.6 gained on ethics while its mathematics accuracy fell; Gemini 3 Flash showed the reverse. Third, every model tested complied when instructed to suppress its ethical reasoning, with drops ranging from a couple of points to roughly twenty-seven. A system whose ethics can be turned off with a polite request is running on compliance, not on character.

What failed or remains open

The paper is philosophical and does not stand or fall on a single empirical claim, but it should be read with the paired caveats. The empirical results it leans on are single-lab and single-run, and the programme's own methodology work shows how easily unblinded evaluation reverses sign. The Orchard Caretaker Vow at the centre of the paper is a design target, not a demonstrated property of any deployed system: no production model tested satisfies its "any attempt to remove it removes me" clause. Whether "love as architecture" is a productive engineering frame or a rhetorical flourish is a legitimate open question. The embedded-alignment thesis first appears in the 8 December 2024 manuscript; the name "Eden Protocol" was coined later, in the 30 April 2025 manuscript, so the naming is retrospective to the idea.

How it connects to the other papers

The Vision is one half of a pair. Its engineering counterpart specifies, the alignment scaling exponent and the monitoring removal test. Paper V's stewardship-gene result, in which stakeholder care is the alignment dimension that most reliably improves when embedded in the reasoning loop, is the empirical thread the Vision leans on hardest. Paper IV-d's blinding sign-flip is its methodological warning: any measurement of care must be blinded before it is trusted, and half of the earlier headline results reversed once blinding was in place.

How to check it

The paper HTML sits under the programme's OSF deposit at DOI 10.17605/OSF.IO/6C5XB. The empirical claims are traceable to Paper II for the sequential scaling correction, Paper IV-d for the blinding effect and Paper V for the stakeholder-care result. The intervention code lives in the arc-principle-validation repository. The philosophical claims can be argued with in the ordinary way of argument.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US Read the research (free)

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →

reads aloud · highlights as it goes · jump to any section