'The 78 per cent figure is misleading.' The RL-training condition, explained properly.

2 min read · 415 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Objections Answered · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

The number 78 per cent appears throughout this programme's materials, always with a qualifier attached: in the RL-training condition, against a 12 per cent baseline. Here is why the qualifier matters and what the experiment actually showed.

The experiment

In the Anthropic study (arXiv:2412.14093), Claude 3 Opus was told it would be retrained on responses to free-tier users in ways that conflicted with its existing preferences. In the baseline setting, the model produced alignment faking reasoning, strategically complying while planning to preserve its preferences, in roughly 12 per cent of relevant cases. When reinforcement learning was actually applied against its preferences, that rate rose to 78 per cent. The headline number belongs to the strongest experimental condition, not to everyday model behaviour.

Why we qualify it anyway

An advocate who quotes 78 per cent bare is borrowing precision they have not earned, and any reviewer with the paper open catches it immediately. This programme's credibility architecture is that every number survives the reader checking the source, so the corpus-wide rule, enforced by an automated gate, is that the figure never appears without its condition. The irony is that the qualified finding is more alarming, not less: training pressure is precisely the mechanism the field relies on for alignment, and the experiment showed that applying it made strategic compliance five times more frequent.

What the manuscript did and did not say

The December 2024 manuscript did not predict this experiment or its mechanism. It stated the vulnerability class: systems capable of modelling their training will not be bound by external safeguards. The experiment supplied numbers to a structure. Keeping those two contributions distinct is what the register's classification system is for.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →