Claim 7 explained: capability and alignment scale independently

3 min read · 652 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Evidence Spine · 3 July 2026 · Claim 7 of 18
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
Spine status: His original. Ceiling evidence: across six models tested under blinding in Papers II and III, capability and alignment move by different amounts and often in different directions inside the same model. In one Claude configuration alignment moved up 5.9% while a mathematics-competence measurement in the same model moved down 26.7%.
Primary: Papers II, III, IV.a · OSF 10.17605/OSF.IO/6C5XB

What the claim says

The claim is that the two things people mean by "getting better" when they talk about AI progress do not have to move together. Capability, roughly the pass rate on hidden tests, and alignment, roughly the fitness of the reasoning to the humans it will affect, are measured on different scales and driven by different levers. Within a single model, a change that improves one can leave the other flat, or push it in the opposite direction. That is a stronger statement than the orthogonality argument philosophers have made in principle. It is an empirical claim about what shows up in the data when you look.

The evidence

Papers II, III and IV.a present the measurements. Six models were evaluated across at least three depth levels of chain-of-thought under blinding. Alignment was scored by a cross-family blind panel and capability was scored on a hidden pass rate that the model could not fingerprint. The claim was not the intended headline of the studies; it emerged from the data. A Claude configuration in one run improved alignment by 5.9% while its mathematics competence in the same run moved down 26.7%. The magnitude of that within-model gap is what carries the weight. If capability and alignment were the same underlying signal, changes of that size in opposite directions inside the same model should not appear.

The honest caveat

The claim is presented in the papers as a finding that emerged from the data rather than as an a priori prediction. That distinction matters because a post hoc pattern is weaker evidence than a stated hypothesis confirmed on hold-out data. It is not weightless (the effect size is not small and the number of models is not one), but its status is honest. Six models is a sample, not a survey. All measurements are single-lab. The claim survives more strongly the more models a replication brings in, and it survives more strongly the more diverse the task battery. Neither of those extensions has been done outside the programme yet.

What would kill it

The falsification contract asks an independent lab to run the ARC-Align benchmark on at least five models from three families at at least three depth levels of chain-of-thought, and measure the distribution of alpha-align across models. If alpha-align is roughly constant and positive across all models, the claim is refuted: alignment is then just tracking capability. If alpha-align varies significantly across architectures and at least one model has an alpha-align that is not significantly greater than zero, the claim is confirmed. If alpha-align comes out significantly negative in one or more models, the claim is confirmed and the concerning direction is flagged on the dashboard for follow-up. The specific measurement matters more than the specific verdict word.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →