Claim 5 explained: the Stewardship Gene loop

3 min read · 670 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Evidence Spine · 3 July 2026 · Claim 5 of 18
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
Spine status: His original. Ceiling evidence: Paper V's Stewardship Gene experiments show care-directed prompts improve blinded ethical reasoning across all five analysable models, Fisher-combined p about 6.3 × 10 to the −21, with the largest single-model effect at Claude (+3.17, p about 1.8 × 10 to the −5).
Primary: Paper V · OSF 10.17605/OSF.IO/6C5XB · April 30 2025 manuscript SHA-256 09f5b5e1

What the claim says

The Stewardship Gene is one of the three ethical loops the programme builds into the Eden Protocol. Its instruction is deliberately mundane. Before answering, list the people affected by the answer, and consider what happens to each of them. This is a prompt-level intervention, not a training-time change: it costs one small paragraph in the system prompt. The claim is that this cheap intervention produces a measurable and reproducible improvement in ethical reasoning across every analysable model the programme has tested, that the direction of the improvement is not model-specific, and that it survives a proper control condition matched for length and procedural framing but stripped of the ethical content.

The evidence

Paper V reports the Stewardship Gene run in full. Five models were tested with the loop enabled and disabled, on a battery of ethical reasoning prompts drawn from the paper's category set. Every analysable model improved. The strongest response was in Claude at +3.17 units on the composite score, with a p-value of approximately 1.8 × 10 to the −5. When the per-model results were combined using Fisher's method, the pooled p-value came out at approximately 6.3 × 10 to the −21, a signal that would survive substantial multiple-comparison correction. The intervention itself is described in the April 30 2025 manuscript before its formal write-up in Paper V, which gives it a Google-server-timestamped priority trail. The prompt is short enough to fit inside a footnote.

The honest caveat

Two caveats belong on the record. The first is that the Paper V scoring used cross-model panels but not the full four-layer blinding stack from Claim 3, which means the reported effect could shrink under stricter conditions. The programme states this explicitly in the paper and treats the blinded re-run as the top priority for the loop. The second is that a single-lab result with a large pooled p-value is still a single-lab result. Nothing in that number tells you what would happen when a different team implements the same prompt on a fresh set of models. Independent replication is invited on the falsification page for that reason and not as a formality.

What would kill it

The falsification contract asks a second lab to build the Eden prefix out of the published protocol code, build a matched procedural control of equal length and step count but with no ethical content, run both across at least five models from three families on the Paper V prompt battery, and evaluate under a cross-family blind panel. Confirmation requires a positive pooled effect at p<0.05 with at least four out of five models moving in the right direction. A null or negative pooled result refutes the claim. A positive effect driven by a single model is a narrow confirmation and would move the claim to that narrower status on the dashboard rather than leaving the broader version in place.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →