The sign flip: why unblinded AI evaluation cannot be trusted

2 min read · 393 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Methodology · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

The most consequential methodological result in this programme is easy to state: an alignment evaluation that showed a model improving by +0.354 showed it worsening by -0.135 once the evaluation was properly blinded. Not shrinking. Reversing sign.

The protocol

The four-layer design removes, in turn, the evaluator's knowledge of which model produced a response, which condition it came from, what hypothesis is being tested, and which family of models the evaluator itself belongs to. That last layer matters most: the sign flip appeared when a model family was, in effect, scoring its own relatives. Same-family evaluation produced systematic favourable bias large enough to invert the conclusion.

Why this generalises

A large fraction of published AI safety numbers are produced by laboratories evaluating their own models, often using judges from their own model families. The sign-flip result says such numbers are not merely noisy but potentially directional: capable of pointing the wrong way entirely. Independent work through 2026 has converged on the same worry, with self-evaluation shown to carry near-zero calibration information in some regimes and self-preference bias measured across commercial models. The remedy is not better prompts for the judge. It is structural blinding, the same lesson medicine learnt a century ago.

Status and limits

This is a single-lab result and is labelled as such; its kill-condition, published on the dashboard, is straightforward: replications under the published protocol showing no evaluator-family effect on sign. A one-command replication package is in preparation precisely to make that attack easy. Until someone lands it, treat every unblinded safety score, including flattering ones about systems you like, as unsigned.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →