Convergence 1: Anthropic alignment-faking paper: 78% in RL-training condition (12% baseline)

2 min read · 423 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Research Notes · 3 July 2026 · Part of the Research evidence spine
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
Register entry 1 of 19. Domain: AI Safety. Institution: Anthropic. External date: 2024-12-18. Measured against: A1 (10 days). Classification: PREDICTION.
Primary source: arXiv:2412.14093

What happened

Anthropic alignment-faking paper: 78% in RL-training condition (12% baseline).

Why it is in the register

The alignment faking result is the anchor of the whole register because the gap is ten days and the direction of the gap is checkable by anyone. The manuscript names the failure class: external safeguards do not bind a system that can model its training process. Anthropic's experiment then supplied the mechanism, the numbers and the peer review. Two things are worth keeping separate here. The manuscript did not predict reinforcement learning specifics, and no claim of that kind is made. What it did was state, in plain language and in advance, that control imposed from outside the system would fail as capability grew. That is the claim the experiment vindicated.

The honest caveat

What would change this assessment: evidence that the manuscript postdates the paper (the Gmail Message-ID and SHA-256 hash foreclose this), or a reading of the manuscript under which the control-failure thesis is absent (the passage is quoted verbatim at HRIH line 14005).

How to check this entry

Open the primary source above. Confirm the date. Confirm the finding. Then compare it against the manuscript passage or artifact named in the register entry, whose SHA-256 hash is published on the evidence page. Nothing in this article asks for trust; the entry either survives that comparison or it comes off the register, publicly, the way the programme's retracted scaling figure did.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →