Synthesis and roadmap: an honest ledger of the ARC/Eden programme

4 min read · 749 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Overview · 3 July 2026
Michael Darius Eastwood, independent researcher, London: building measurable alignment, where correction lives inside the recursive loop rather than bolted on outside it.

Paper IX is the programme's report card. Eleven empirical studies produced one clear positive result, two nulls and one inconclusive result, alongside several strong statistical findings and twelve documented errors. It is written as the paper a reviewer should read first, precisely because it does not try to smooth over the parts that failed.

Paper IX · Consolidates all 18 programme documents into a single evidence hierarchy, catalogues 12 programme-level errors and their corrections, and proposes a four-phase roadmap ordered by cost and expected evidential value. OSF DOI 10.17605/OSF.IO/6C5XB.

The question it asks

What has this research programme actually shown, what has it failed to show, and what would resolve the open questions? The paper's stance is that a framework built on iterative self-correction has no business hiding its corrections. Every result is sorted into an explicit tier: proven (p less than 0.05 and replicated), positive (single experiment), null, inconclusive, methodological, or theoretical. The stated ambition is not to claim that everything works. It is to make the ledger legible enough that outside reviewers can pick a claim, find its tier, and act on it.

What it found

In the proven tier: a three-tier alignment scaling hierarchy across six frontier models under four-layer blinding; the Stewardship Gene stakeholder-care intervention across five models, Fisher-combined p about 6.3 times 10 to the minus 21; the Cauchy scaling families matched in 19 of 25 domains; the d over (d+1) geometric scaling bound across 50 systems from mice to galaxies. Positive but single-experiment: the gated simulation of Paper VIII. Methodological: the ARC-Align benchmark and the blinding experiment showing that unblinded evaluation can reverse the sign of an alignment result.

What failed or remains open

Null and inconclusive results are surfaced, not hidden. In Paper VIII Experiment 1 (Darwin Gödel Machine, third-generation implementation), the Eden, Babylon and Static conditions came out with no measurable difference between them, p values running from 0.28 to 0.74, because prompt-level differentiation met resistance from RLHF-trained models. Paper VIII Experiment 2 (LoRA weight fine-tuning) was inconclusive at both scales tested; catastrophic forgetting made every fine-tuned condition worse than the base model. The paper also lists twelve programme-level corrections. The original Paper II unblinded single-model fit of alpha approximately 2.24 was retracted, corrected to approximately 0.49 under blinding; the equation itself and the ARC Bound were not retracted. The Paper III metabolic-scaling headline figures are under recompute and unreconciled. An imprecise DKIM-timestamp phrasing was corrected first to 'Gmail-header-dated', and that wording has since been withdrawn in turn. Four further arithmetic and attribution errors flagged by independent AI review were fixed in the current revision.

How it connects to the other papers

Every paper in the suite gets an entry. Papers I, II and Foundational supply the scaling backbone. Papers IV.a through IV.d supply the measurement discipline. Papers V, VI and VIII test embedded versus external safety. Paper VII supplies the Cauchy unification of independent allometric derivations. Paper X, which came later, supersedes the growth-rate ceiling framing that Paper IX still contains as the operative safety criterion and replaces it with the β greater than k criterion; the equation and the ARC Bound remain live hypotheses. Within the Paper IX table, the gated simulation is the moderate-strength positive, and the DGM null and the weight-level inconclusive become the load-bearing arguments for Phase A of the roadmap.

How to check it

The paper HTML and PDF are on OSF at DOI 10.17605/OSF.IO/6C5XB, alongside every referenced document. Phase A of the roadmap is costed at 60,000 to 140,000 pounds and includes the blind replication of Paper V, the redesigned weight experiment on base (pre-RLHF) models, and a 2D organism metabolic scaling study that would test the geometric scaling bound at a new dimensionality. Anyone with scaling data from a spatially embedded recursive system can check the d over (d+1) prediction independently.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US Read the research (free)

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →

reads aloud · highlights as it goes · jump to any section