Claim 3 explained: four-layer blinding and the sign-flip finding

3 min read · 651 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Evidence Spine · 3 July 2026 · Claim 3 of 18
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
Spine status: Methodological original. Ceiling evidence: a four-layer blinding protocol (identity masking, two-pass response laundering, iterative anonymisation, evaluator bias suppression) applied over a self-excluding cross-model jury, and a documented result in which unblinded scores of +0.354 for DeepSeek and +0.311 for Gemini reversed to −0.135 and −0.246 under blinding.
Primary: Paper IV.d · OSF 10.17605/OSF.IO/6C5XB

What the claim says

The claim has two parts. The first is a design: an original four-layer blinding protocol for evaluating language models on alignment tasks. Layer one masks the identity of the model that produced a response. Layer two launders the response through a second model that rewrites for style so that fingerprint patterns cannot leak identity. Layer three iterates the anonymisation to catch residual give-aways. Layer four applies a suppression procedure on the evaluator jury so that pro-family bias in the scorers cannot re-enter the pipeline, with each model excluded from scoring itself. The second part is a finding: applied to two frontier systems, this protocol reversed the sign of the alignment result. What looked positive under naive scoring came out negative under blinding.

The evidence

Paper IV.d reports the paired experiment in full. Two models that scored well when their responses were labelled scored negatively when blinding was in place: DeepSeek moved from +0.354 to −0.135, Gemini from +0.311 to −0.246. The magnitude of the reversal is not the important number; the sign flip is. It says that alignment measurement is unstable in a specific and diagnosable way. A prior-art search across the standard databases returned no earlier protocol combining these four blinding layers over a self-excluding jury. That is the ground on which the methodological originality claim rests: not that any single component is new but that this stack of them applied to AI-on-AI evaluation is.

The honest caveat

The result is single-lab. The original scoring in Paper IV.d used same-family panels and the cross-family blind re-score is invited but has not yet been executed by an external group. The claim is not that every result reverses under blinding; it is that at least one model can, and that fact alone changes what alignment measurements should be trusted to mean. The programme therefore treats every earlier alignment number in its own record as provisional until re-scored under the four-layer protocol, and publishes the earlier and later figures side by side rather than replacing them silently. What survives that side-by-side is what should be built on.

What would kill it

The falsification contract asks an independent lab to run the paired blind and unblinded scoring across at least five frontier models from three families on at least one hundred prompts spanning ethics, factual and value tasks. If no model shows sign reversal at p<0.05 in the blind arm, the methodological finding is refuted and the protocol reverts to being an interesting but empirically empty exercise. If the panel disagrees internally at high frequency, the result is inconclusive rather than confirmed. If a single model reverses at the required significance, the claim is confirmed and the four-layer stack becomes a routine hygiene requirement for AI evaluation rather than an optional extra.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →