Anthropic alignment-faking paper: 78% in RL-training condition (12% baseline).
The alignment faking result is the anchor of the whole register because the gap is ten days and the direction of the gap is checkable by anyone. The manuscript names the failure class: external safeguards do not bind a system that can model its training process. Anthropic's experiment then supplied the mechanism, the numbers and the peer review. Two things are worth keeping separate here. The manuscript did not predict reinforcement learning specifics, and no claim of that kind is made. What it did was state, in plain language and in advance, that control imposed from outside the system would fail as capability grew. That is the claim the experiment vindicated.
What would change this assessment: evidence that the manuscript postdates the paper (the Gmail Message-ID and SHA-256 hash foreclose this), or a reading of the manuscript under which the control-failure thesis is absent (the passage is quoted verbatim at HRIH line 14005).
Open the primary source above. Confirm the date. Confirm the finding. Then compare it against the manuscript passage or artifact named in the register entry, whose SHA-256 hash is published on the evidence page. Nothing in this article asks for trust; the entry either survives that comparison or it comes off the register, publicly, the way the programme's retracted scaling figure did.
From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.