Home / Priority claims register / PC-007

Priority claim · PC-007

AI-control-failure prediction recorded ten days before Anthropic's alignment-faking paper

First date: Timeline

This was first published by Michael Darius Eastwood on . Anchored evidence is listed below. See the machine-readable priority record at /research/priority/ and the full ledger at /priority-claims.

The claim - verbatim

Michael Darius Eastwood first recorded the structural AI-control-failure passage ('AI systems cannot be truly controlled; they will evolve beyond any safeguards') on 8 December 2024 at 02:45 UTC in a self-emailed manuscript (Google-server-timestamped, SHA-256 prefix f0d1f38f, HRIH manuscript line 14005), ten days before Anthropic's arXiv preprint reporting alignment-faking behaviour under an RL-training condition in large language models (Greenblatt et al., arXiv:2412.14093, 18 December 2024). The relationship is convergent, not causal; Anthropic's paper is independent work by different researchers on a different specific mechanism.

Anchored evidence

Risk notes - what this claim does not support

Risk notes. Any citation of Anthropic's 78 per cent alignment-faking figure must include both the RL-training condition and the 12 per cent baseline, per the site's truth gate. Do not claim Eastwood predicted the specific mechanism Anthropic measured. Existence-by-date only.

Verify this claim. The machine-readable priority register is at /research/priority.json (this claim is PC-007 in claimsLedger[]). The full public ledger view is at /priority-claims. Where evidence points at a redacted evidence/... path, the public verification method is documented at /research/evidence-spine.html. Every result is independently checkable via the OSF programme at osf.io/6c5xb (DOI 10.17605/OSF.IO/6C5XB).

← Back to the priority claims register · public ledger view · machine-readable priority record · home.