Skip to content

← Questions

What is the embedded-correction alignment thesis?

The embedded-correction alignment thesis, first placed on the public record by Michael Darius Eastwood on 8 December 2024, holds that alignment must be embedded in the substrate of an artificial-intelligence system rather than imposed from outside. It is presented as a dated framing and named formulation, not as invention of embedded-values training, where Anthropic's Constitutional AI is prior work.

Concept First anchored

The embedded-correction alignment thesis, first placed on the public evidentiary record by Michael Darius Eastwood on 8 December 2024, holds that alignment must be embedded in the substrate of an artificial-intelligence system rather than imposed from outside. The claim is over the specific dated framing and named formulation of this thesis in the 8 December 2024 self-emailed manuscript, not over the underlying idea of value-embedded training, on which Anthropic's Constitutional AI work from 2022 is independent prior work.

The anchor is a self-emailed Gmail package bearing a visible Gmail origination field of 02:45:18 UTC on 8 December 2024, with sender-DKIM and paired Google ARC mathematics verifying against keys captured at the signed selectors and no RFC 3161 trusted timestamp, Message-ID CAGPsKAnp-DLDBOeBAh4t9XsqJBKdoe07uORGF86meoSMFnin7w, of 221,236 words across five attachments (four manuscript drafts totalling 221,236 words (five attachments) and the 31,881-word book draft v3.2), with SHA-256 anchor prefix f0d1f38f. The verbatim phrasing appears at lines 20 and 40 of the version-2 essay and at line 14001 of the Hyperspace Recursive Intelligence Hypothesis (HRIH) manuscript.

The thesis pairs with a structural AI-control-failure passage in the same manuscript ('AI systems cannot be truly controlled; they will evolve beyond any safeguards') recorded ten days before Anthropic's alignment-faking paper (Greenblatt et al., arXiv:2412.14093, 18 December 2024) reported an independent alignment-faking behaviour in large language models under an RL-training condition. The temporal relationship is convergent, not causal. The Google-server timestamp establishes existence-by-date; it does not establish a DKIM-authenticated authorship claim, and the manuscript identifies structural directions rather than specific RL-training mechanisms.

reads aloud · highlights as it goes · jump to any section