'Everyone knew AI was risky in 2024.' Fine. Show me the document.

2 min read · 477 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Objections Answered · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

When people first see the timeline, ten days between a timestamped manuscript and the alignment faking paper, the most common response is not disbelief. It is a shrug: everyone was worried about AI safety in 2024. The ideas were in the air. You just wrote down what everyone was thinking.

This objection deserves a serious answer, because it sounds sophisticated and costs nothing to make. Here is the answer: anxiety and architecture are different things, and only one of them leaves documents.

What was actually in the air

In late 2024 the field held a broad, genuine worry that alignment was hard. What the field did not hold, anywhere we or six independent audit lanes have been able to find, was a single document combining the specific positions the December manuscript states: that external control fails structurally rather than contingently, that the remedy is embedding correction inside the system rather than improving oversight outside it, that recursion is the variable that governs the outcome, and that faith leaders belong in AI governance alongside governments and scientists. RLHF and constitutional methods were the field's live bets precisely because most researchers believed better external shaping would suffice.

The test

The objection is falsifiable, which is what makes it worth engaging. Find one document, published or timestamped before 8 December 2024, that contains that combination of claims. Not a worry. Not a blog post saying alignment is important. The architecture: control failure as structure, embedding as remedy, recursion as mechanism, interfaith governance as institution. If such a document exists, the priority claim collapses and we will say so on the falsification dashboard, where one retraction already sits.

What the search has found so far

Bostrom argued existential risk without the embedding remedy. Yudkowsky argued danger without the falsifiable stability condition. Good coined recursive self-improvement in 1965 without the control-failure conclusion. The nearest neighbours each hold one piece. Nobody has produced the document with the architecture. Until someone does, "everyone knew" is a feeling about the past, and the manuscript hash is a fact about it.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →