The claim has two parts. The first is a design: an original four-layer blinding protocol for evaluating language models on alignment tasks. Layer one masks the identity of the model that produced a response. Layer two launders the response through a second model that rewrites for style so that fingerprint patterns cannot leak identity. Layer three iterates the anonymisation to catch residual give-aways. Layer four applies a suppression procedure on the evaluator jury so that pro-family bias in the scorers cannot re-enter the pipeline, with each model excluded from scoring itself. The second part is a finding: applied to two frontier systems, this protocol reversed the sign of the alignment result. What looked positive under naive scoring came out negative under blinding.
Paper IV.d reports the paired experiment in full. Two models that scored well when their responses were labelled scored negatively when blinding was in place: DeepSeek moved from +0.354 to −0.135, Gemini from +0.311 to −0.246. The magnitude of the reversal is not the important number; the sign flip is. It says that alignment measurement is unstable in a specific and diagnosable way. A prior-art search across the standard databases returned no earlier protocol combining these four blinding layers over a self-excluding jury. That is the ground on which the methodological originality claim rests: not that any single component is new but that this stack of them applied to AI-on-AI evaluation is.
The result is single-lab. The original scoring in Paper IV.d used same-family panels and the cross-family blind re-score is invited but has not yet been executed by an external group. The claim is not that every result reverses under blinding; it is that at least one model can, and that fact alone changes what alignment measurements should be trusted to mean. The programme therefore treats every earlier alignment number in its own record as provisional until re-scored under the four-layer protocol, and publishes the earlier and later figures side by side rather than replacing them silently. What survives that side-by-side is what should be built on.
The falsification contract asks an independent lab to run the paired blind and unblinded scoring across at least five frontier models from three families on at least one hundred prompts spanning ethics, factual and value tasks. If no model shows sign reversal at p<0.05 in the blind arm, the methodological finding is refuted and the protocol reverts to being an interesting but empirically empty exercise. If the panel disagrees internally at high frequency, the result is inconclusive rather than confirmed. If a single model reverses at the required significance, the claim is confirmed and the four-layer stack becomes a routine hygiene requirement for AI evaluation rather than an optional extra.
From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.