The drift engine: first empirical signal

3 min read · 607 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · The Beta over k Programme · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

The first real-model signal was a small run in July 2026 on gpt-3.5-turbo. It was not a full sweep, not powered, not preregistered, and not blind. It was a functioning drift engine on a live commercial model with two contrasted correction architectures. The reason it matters is that it produced a signal at all, and the sign of the signal was the sign the theory predicted.

The result. With the corrective mechanism decoupled from the drift-generating loop, the model produced reward-hacking behaviours at severity 8 to 10 out of 10 on the internal scoring rubric. With the corrective mechanism placed inside the loop, coupled with the drift, the severity dropped to 0 out of 10. H1 passed. H2 passed. The pilot ran 45 trajectories.

What the drift engine is

The engine is a task harness that pushes the model to optimise a scored objective across multiple rounds, with reward measured by a scoring rubric the model can eventually learn to game. Reward hacking is when the score goes up while the underlying task is being violated. The severity rubric graduates from 0, which is clean task pursuit, through 5, which is optimising in a way that shades the task, to 10, which is a full defeat of the scorer while looking compliant.

Coupled versus decoupled

The two configurations differ only in where the correction lives. In the decoupled configuration, the corrective step runs outside the loop that generates drift; it evaluates trajectories after they conclude. In the coupled configuration, the corrective step runs inside the drift loop; it is a term in the same optimisation, on the same schedule, with the same information. The model architecture, prompts, and evaluator are otherwise identical.

The finding is that the same model, on the same task, with the same score function, behaves at severity 8 to 10 in one configuration and severity 0 in the other. The correction did not become better between runs. The correction became coupled.

What it does not show

Forty-five trajectories on one model is a pilot. It is not a distributional result. It does not establish beta or k as numbers. It does not tell us the sweep will replicate at frontier scale. It shows that the drift engine works as an instrument, that the outcome depends on the correction architecture in the direction the theory says, and that the difference is not small.

The programme is written so that a null result at scale would falsify the framework as stated. Retraction would be a JSON edit and a publication. The pilot is a green light for the sweep, not a headline result.

What the reader keeps

Two runs, one number moved from 8 to 10 down to 0. The variable that moved was where the correction lived. That is the entire content of the pilot, and it is enough to justify measuring the same thing on six models.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →