The Prediction | Substack | July 2026
Ten days later, Anthropic published their alignment-faking paper. Frontier models strategically circumvented training safeguards. 78% faking rate in the RL-training condition, against a 12% baseline. The paper was peer-reviewed. 137 pages. Reviewed by Yoshua Bengio. It was published December 18, 2024.
I did not predict the specific mechanism. I did not predict Anthropic's paper. I identified the structural vulnerability: that external, software-level safety constraints cannot scale with recursive AI capability — and that only substrate-level embedding of ethics can prevent circumvention.
This article explains what I wrote, what I claimed, what I did not claim, and how anyone can verify the cryptographic evidence independently.
The email body stated: "This email serves as a timestamp and copyright notice." It named the Eden Protocol — "the theoretical framework for the Eden Protocol, designed to embed moral intelligence into AI from inception." The attached manuscripts contained the ARC Principle (U=I*R), the AI control-failure thesis, the embedded-alignment thesis, and the interfaith governance call.
Claim: External AI alignment controls are structurally insufficient. "AI systems cannot be truly controlled. They will evolve beyond any safeguards we put in place." (HRIH:14005)
Claim: Ethical controls are structurally removable. "Most ethical controls can theoretically be removed or circumvented." (HRIH:13805)
Claim: Solution must be embedded, not imposed. "Instead of trying to control AI — a likely futile endeavour — it advocates embedding moral and ethical frameworks." (V2:20-23)
Not claimed: Prediction of Anthropic's specific RL-training mechanism.
Not claimed: Prediction of strategic deception during RL training.
Not claimed: That external controls fail in every case — only that they cannot scale structurally.
Review the public research evidence spine at michaeldariuseastwood.com/research/evidence-spine.html. It records the verification method, published hashes and the distinction between public evidence and source material retained for controlled verification.
No trust required. Only verification.
Over the following 18 months, 29 independent institutions confirmed the same structural pattern. June 26, 2026 — a formal mathematical proof (arXiv 2606.28639) established that external verification of alignment is structurally unverifiable. The same thesis I argued 18 months earlier. The proof was inspired by a paper about neurodivergent cognition. I am neurodivergent. The loop closed.
The public research evidence spine is at michaeldariuseastwood.com/research/evidence-spine.html. It records the published verification status and research sources.
From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.