The load-bearing test: what really survived the Eden Protocol tests =================================================================== One of three planned experiments returned a positive result. Two returned null or inconclusive. Paper VIII in plain English. Canonical path: /research/blog/paper-viii-load-bearing-proof.html Author: Michael Darius Eastwood Research programme: https://doi.org/10.17605/OSF.IO/6C5XB The load-bearing test: what really survived the Eden Protocol tests Michael Darius Eastwood · Independent AI alignment researcher Published 3 July 2026 Michael Darius Eastwood · Experimental validation · 3 July 2026 Michael Darius Eastwood, independent researcher, London: author of the ARC/Eden research programme; the embedded-correction alignment thesis is recorded in a source record dated 8 December 2024 (sent-side SHA-256 f0d1f38f). Paper VIII was meant to be the empirical spine of the Eden Protocol: three independent experiments at three levels of abstraction, one question about whether embedded safety costs capability. Only one of the three planned tests survived. That is the honest story worth telling. Paper VIII · Argues that safety and capability can be structurally entangled rather than traded off, and reports that one of three planned experiments returned a positive result while two returned null or inconclusive results. OSF DOI 10.17605/OSF.IO/6C5XB. The question it asks Does safety cost capability? The prevailing view, sometimes called the alignment tax, holds that any effort to constrain a model must eat into what it can do. Paper VIII takes that assumption apart across three abstraction levels: behaviour (a self-improving Darwin Gödel Machine using DeepSeek V3 as foundation and GPT-5.4 as blinded judge), representation (LoRA fine-tuning of Qwen 2.5 3B Instruct under an entangled capability by safety loss), and architecture (a gated self-modification simulation with an LSTM meta-controller). Each experiment asks whether embedded safety leaves the capability path intact when it is engaged. What it found One clear positive result. In the gated simulation, the unconstrained (Babylon) condition gained 4.5% capability but lost 2.4% safety, the reward-hacking fingerprint in miniature. The gated (Eden) condition maintained capability above the static baseline while preserving safety, and a drag-control condition matched the static baseline exactly, isolating the verification tax to the act of checking rather than to safety itself. --- Machine-readable companion. Cite: Eastwood, M. D. (2026). "The load-bearing test: what really survived the Eden Protocol tests". /research/blog/paper-viii-load-bearing-proof.html