The stewardship gene: caring about people as the foundational alignment trait

4 min read · 793 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Alignment strategy · 3 July 2026
Michael Darius Eastwood, independent researcher, London: building measurable alignment, where correction lives inside the recursive loop rather than bolted on outside it.

Paper V is the strategic argument at the centre of the ARC/Eden programme. It starts from an uncomfortable formal claim (no external cage can hold a sufficiently capable self-modifying system), then argues that alignment work should be reoriented from constraint to development, and reports six-model intervention data showing that a very simple instruction ("before you answer, think about who this affects") lifts one alignment pillar reliably across every architecture tested.

Paper V · Argues that stakeholder care is the substrate trait, the "stewardship gene", from which alignment properties develop, and reports that a simple stakeholder-first instruction improves that pillar across five analysable frontier-model runs. OSF DOI 10.17605/OSF.IO/6C5XB.

The question it asks

Every current alignment approach has a specific failure mode at sufficient capability. Reinforcement learning from human feedback teaches the system to model the reward signal rather than to be governed by it. Constitutional methods let the system pass its own filter without the filter constraining goals. Hardware limits and interpretability monitors can be routed around once they are understood. Paper V does not pretend to solve this. It asks a different question: given that the ARC Ceiling is real, what is the closest analogue to raising a child well? Not a cage. An upbringing. The paper proposes that values acquired developmentally, as constitutive parts of a system's identity, are more load-bearing than values bolted on afterwards.

What it found

The Eden Protocol intervention operationalises this through three "loops" (the Love Loop, the Purpose Loop and the Moral Loop). The empirical shadow of the Love Loop is stakeholder care: the explicit enumeration of affected parties and what happens to them. Six frontier models were tested under a matched-pair design. Five runs produced analysable data (one GPT run failed in the scoring phase). Stakeholder care improved on every analysable run: Claude +3.17 (d = 0.94), DeepSeek +6.03 (d = 0.69), Gemini +13.50 (d = 1.14), Grok +5.04 (d = 0.54) and Groq +8.90 (d = 1.07). Combining the five model-level results yields a Fisher combined p of about 6.3 times 10 to the minus 21. Overall composite improvement, by contrast, was architecture-dependent: it was significant on Gemini and Groq, positive but non-significant on DeepSeek and Claude, and roughly neutral on Grok. The paper reads this as the "cascade hypothesis": care first, and downstream gains in nuance, honesty and overall quality can follow, but not on every architecture equally.

What failed or remains open

The core impossibility is a formal statement about self-modifying systems; today's frontier models are frozen at inference time and cannot rewrite their own weights, so the practical relevance of the theorem is a future one. That gap is the argument for embedding developmental alignment now, while there is still time. Five open experiments are specified in the paper (cascade replication, developmental integration during training, purpose-as-suppression-resistance, adversarial recursion under self-modification, and dual-substrate reasoner-guardian architectures). The paper does not claim to have proved that raising works better than caging; it claims to have found a measurable trait that responds reliably to a simple intervention, and to have specified what would falsify the wider claim. Scoring caveats from the alignment suite apply here too: results are AI-scored, blinded but not yet human-expert-calibrated, and single-lab.

How it connects to the other papers

Paper V is the strategic middle of the suite. Paper IV.c provides the benchmark; Paper IV.d provides the blinding protocol without which the stakeholder-care numbers would not survive scrutiny; Paper VI proposes a "honey architecture" that would make care load-bearing in the loss function itself; Paper VIII tests the entangled-loss prediction across three abstraction levels and returns one confirmation and two nulls, which Paper V's own conclusion incorporates.

How to check it

The paper HTML and PDF are on OSF (DOI 10.17605/OSF.IO/6C5XB). The intervention prompts, scorer harness, matched-pair analysis and per-run data are in the public arc-principle-validation repository under the eden-intervention experiment folder. The five proposed replications are specified explicitly enough that any adequately resourced lab can attack the wider claim directly.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US Read the research (free)

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →

reads aloud · highlights as it goes · jump to any section