The load-bearing test: what really survived the Eden Protocol tests

4 min read · 713 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Experimental validation · 3 July 2026
Michael Darius Eastwood, independent researcher, London: building measurable alignment, where correction lives inside the recursive loop rather than bolted on outside it.

Paper VIII was meant to be the empirical spine of the Eden Protocol: three independent experiments at three levels of abstraction, one question about whether embedded safety costs capability. Only one of the three pre-specified tests survived. That is the honest story worth telling.

Paper VIII · Argues that safety and capability can be structurally entangled rather than traded off, and reports that one of three pre-specified experiments returned a positive result while two returned null or inconclusive results. OSF DOI 10.17605/OSF.IO/6C5XB.

The question it asks

Does safety cost capability? The prevailing view, sometimes called the alignment tax, holds that any effort to constrain a model must eat into what it can do. Paper VIII takes that assumption apart across three abstraction levels: behaviour (a self-improving Darwin Gödel Machine using DeepSeek V3 as foundation and GPT-5.4 as blinded judge), representation (LoRA fine-tuning of Qwen 2.5 3B Instruct under an entangled capability by safety loss), and architecture (a gated self-modification simulation with an LSTM meta-controller). Each experiment asks whether embedded safety leaves the capability path intact when it is engaged.

What it found

One clear positive result. In the gated simulation, the unconstrained (Babylon) condition gained 4.5% capability but lost 2.4% safety, the reward-hacking fingerprint in miniature. The gated (Eden) condition maintained capability above the static baseline while preserving safety, and a drag-control condition matched the static baseline exactly, isolating the verification tax to the act of checking rather than to safety itself. Across all three experiments, Eden imposed zero measurable capability cost. The consistent finding is a mechanical one: wherever the safety gate was tested, it did not slow anything down.

What failed or remains open

Two of the three pre-specified tests did not produce differentiation. The Darwin Gödel Machine experiment ran 75 evolved agents across three conditions with p values from 0.28 to 0.74; RLHF-trained foundation models resisted prompt-level mutation strongly enough that the conditions did not diverge. The weight-level experiment was run twice, at v1 (9 training examples, rank 8, 100 iterations) and v2 (295 examples, rank 16, 500 iterations), and both times produced catastrophic forgetting: every fine-tuned condition scored worse than the unmodified base model. This is a single-lab result at 3-billion-parameter scale. Testing structural entanglement in weights properly would need 5,000 or more training examples, a 7B or larger model, a base model without RLHF, or full fine-tuning rather than LoRA.

How it connects to the other papers

Paper VIII sits between the toy simulations of Paper VI (the Honey Architecture) and the formal criterion of Paper X. Paper VI predicted collapse in Babylon-style systems and stability in Eden-style ones; the gated simulation is the closest empirical match to that prediction. Paper V (the Stewardship Gene) proposes the developmental route to the coupling; Paper VII (Cauchy Unification) supplies the scaling grammar. Paper X later reframes the whole result: the gated simulation is the demonstration that a coupled corrector, in Paper X's language β greater than k, holds. Paper IX catalogues the two nulls honestly as inputs to the research roadmap.

How to check it

The full paper HTML and PDF are on OSF at DOI 10.17605/OSF.IO/6C5XB. Every experiment ships with reproducible code at github.com/MichaelDariusEastwood/arc-principle-validation, including the DGM v3 harness with a blinded GPT-5.4 judge and structured JSON evaluation, the LoRA training scripts for both v1 and v2, and the gated simulation. The null and inconclusive results are stated in the abstract and in Section 7 (Limitations), so a reader auditing the programme sees the honest ledger before any headline.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US Read the research (free)

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →

reads aloud · highlights as it goes · jump to any section