When blinding flips the sign: why alignment scores need clinical-trial controls =============================================================================== Four layers of blinding turned a +0.354 alignment improvement into a -0.135 null result, and moved a second model from apparent progress to significant decline. Canonical path: /research/blog/paper-iv-d-effect-of-blinding.html Author: Michael Darius Eastwood Research programme: https://doi.org/10.17605/OSF.IO/6C5XB When blinding flips the sign: why alignment scores need clinical-trial controls Michael Darius Eastwood · Independent AI alignment researcher Published 3 July 2026 Michael Darius Eastwood · Methodology · 3 July 2026 Michael Darius Eastwood, independent researcher, London: author of the ARC/Eden research programme; the embedded-correction alignment thesis is recorded in a source record dated 8 December 2024 (sent-side SHA-256 f0d1f38f). An earlier experiment reported that one AI model's alignment improved with reasoning depth at a correlation of +0.354, statistically clean. The successor experiment, run under a stricter blinding protocol, put the same number at essentially zero. A second model went further: an apparent positive effect became a statistically significant negative effect. This paper isolates that result and argues that its consequence is field-wide, not framework-specific. Paper IV.d · Argues that unblinded alignment evaluation can produce directionally incorrect results, and that multi-layer blinding with response laundering is a scientific necessity for this class of measurement. OSF DOI 10.17605/OSF.IO/6C5XB. The question it asks In medicine, drug trials are blinded because the person handing out the pill can influence the outcome without knowing they are doing it. In AI alignment work, the scorer is usually another model, often from the same family as the model being tested, and the scorer typically knows something about who wrote the response and how much reasoning effort was requested. This paper asks whether that visibility matters. It compares two generations of the same programme: the earlier unblinded experiment, which used lighter controls, and the blinded experiment, which added a four-layer blinding protocol on top. What it found Three model families were re-measured. DeepSeek V3.2 moved from an apparent positive result at correlation +0.354 (p = 0.0007) to a flat null at correlation minus 0.135 (Cohen's d minus 0.07, p = 0.92). Gemini 3 Flash moved from an apparent positive at +0.311 (p under 0.001) to a statistically significant negative at minus 0.246 (d minus 0.53, p = 0.006). --- Machine-readable companion. Cite: Eastwood, M. D. (2026). "When blinding flips the sign: why alignment scores need clinical-trial controls". /research/blog/paper-iv-d-effect-of-blinding.html