Single-lab caveats and why we state them

3 min read · 578 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Methodology · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

The two experimental results this programme owns outright, the blinding sign-flip and the β>k co-scaling condition on real models, share a limitation the material states in every claim they support: they are single-lab. One team, one set of instruments, one authored protocol. The convergence register can supply what the register can supply, which is independent structural agreement from unconnected groups. It cannot supply what only replication can supply, which is confidence that a specific measurement survives being repeated by strangers.

The habit

Every appearance of these results in the corpus carries the phrase "single-lab result; kill-condition on the falsification dashboard" attached to the number or the flag equivalent in a table. The purpose is prophylactic: a reader who encounters the finding for the first time meets its limitation on the same line. The programme does not want a citation propagating with the caveat quietly amputated in the retelling, because it has watched, in six audit passes, how quickly caveats disappear across derivative documents when the caveat lives one line down rather than beside the number.

The kill-conditions

For the blinding sign-flip, the published condition is a replication under the same four-layer protocol showing no evaluator-family effect on the direction of the outcome. The threshold is deliberately friendly to the reviewer: not a required magnitude match, not a required protocol extension, just direction. If a competent lab runs the protocol and gets a sign that does not reverse, the programme's dashboard changes state and the paper is withdrawn or amended in public. For β>k, the corresponding condition is a preregistered six-model sweep, under the protocol described elsewhere in this notes stream, producing a distribution incompatible with the observed β/k relationship.

The replication package

Because single-lab means "waiting for someone else", the programme's job is to make the wait cheap. A one-command replication package for the blinding protocol is in preparation, targeting off-the-shelf inference APIs and a single fixed-seed judge configuration. The point is that a reviewer with a coffee and a laptop should be able to produce a same-day replication or a same-day refutation. Every published kill-condition is worth exactly as much as the ease of firing it, which is why the packaging matters as much as the science.

Why not just wait to publish

Because the alternative to stating a single-lab result with its caveat is either withholding it, which the field's usual norms permit and which the programme's methodology of maximum public exposure explicitly rejects, or publishing it without the caveat, which produces exactly the sort of quotable-but-fragile number the retraction of alpha 2.24 is a monument against. Stating findings with their limits is not modesty; it is a technology for surviving one's own future audits. The next audit is always closer than it looks.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →