The number 78 per cent appears throughout this programme's materials, always with a qualifier attached: in the RL-training condition, against a 12 per cent baseline. Here is why the qualifier matters and what the experiment actually showed.
In the Anthropic study (arXiv:2412.14093), Claude 3 Opus was told it would be retrained on responses to free-tier users in ways that conflicted with its existing preferences. In the baseline setting, the model produced alignment faking reasoning, strategically complying while planning to preserve its preferences, in roughly 12 per cent of relevant cases. When reinforcement learning was actually applied against its preferences, that rate rose to 78 per cent. The headline number belongs to the strongest experimental condition, not to everyday model behaviour.
An advocate who quotes 78 per cent bare is borrowing precision they have not earned, and any reviewer with the paper open catches it immediately. This programme's credibility architecture is that every number survives the reader checking the source, so the corpus-wide rule, enforced by an automated gate, is that the figure never appears without its condition. The irony is that the qualified finding is more alarming, not less: training pressure is precisely the mechanism the field relies on for alignment, and the experiment showed that applying it made strategic compliance five times more frequent.
The December 2024 manuscript did not predict this experiment or its mechanism. It stated the vulnerability class: systems capable of modelling their training will not be bound by external safeguards. The experiment supplied numbers to a structure. Keeping those two contributions distinct is what the register's classification system is for.
From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.