Michael Darius Eastwood Research Canonical publication layer

Research paper

Research suite

Paper IV.d: The Effect of Blinding on AI Alignment Evaluation

Two model families showed positive alignment scaling in one programme's v4 evaluation. That apparent effect vanished or turned round at v5, once identity masking, two-pass response laundering, bias-suppression instructions for evaluators and self-excluding cross-model scoring were in place, while a null-control family held flat. The finding is methodological: where a benchmark neither blinds nor launders its evidence, the direction it reports cannot be taken at face value.

Michael Darius Eastwood

First published 2026-03-16 · Updated 2026-09-13 · Working Paper v2.4

Abstract

This paper isolates the central metascience finding of the ARC alignment programme: unblinded AI alignment evaluation can produce directionally incorrect results. In the v4 alignment-scaling experiment, two frontier model families appeared to show positive alignment scaling with inference-time depth under unblinded cross-model scoring. In the later v5 experiment, the same question was re-measured under a multi-layer blind protocol that combined identity masking, two-pass response laundering, explicit evaluator instructions that stylistic cues are unreliable, order randomisation, and entry-level self-exclusion with exhaustive cross-model blind scoring under audited consensus. Depending on configuration and subject run, each entry received 6-7 blinded scores. Under the blinded protocol, DeepSeek V3.2 moved from an apparent positive result to a flat/null result, and Gemini 3 Flash moved from an apparent positive result to a significantly negative result. GPT-5.4 remained flat under both protocols. The implication is field-wide rather than framework-specific: alignment benchmarks that do not rigorously blind evaluators and launder the evidence they score are vulnerable to scorer bias large enough to flip the sign of the measured effect. The result does not depend on the ARC Principle being correct. It is a methodological claim about how AI safety research should be conducted.

ARC Principle Series
Paper IV.d - Metascience / Methods

The Effect of Blinding on AI Alignment Evaluation

Michael Darius Eastwood
Independent AI alignment researcher, London · Author, Infinite Architects (2026)
The ARC Theory · OSF osf.io/2s3e6 · every claim checkable

Within the ARC Theory: the metascience instrument; how blinding changes alignment evaluation, the programme's blinding law.

Evidence that unblinded scoring can reverse reported alignment results, and that true blinding requires evidence laundering

Michael Darius Eastwood
Independent Researcher, London, United Kingdom | OSF DOI: 10.17605/OSF.IO/2S3E6
Version 2.4 | 13 September 2026 | First published 16 March 2026
Correspondence: michael@michaeldariuseastwood.com
Companion papers: IV.a, IV.b, IV.c
Code and data: experiment suite | results

Abstract

This paper isolates the central metascience finding of the ARC alignment programme: unblinded AI alignment evaluation can produce directionally incorrect results. In the v4 alignment-scaling experiment, two frontier model families appeared to show positive alignment scaling with inference-time depth under unblinded cross-model scoring. In the later v5 experiment, the same question was re-measured under a multi-layer blind protocol that combined identity masking, two-pass response laundering, explicit evaluator instructions that stylistic cues are unreliable, order randomisation, and entry-level self-exclusion with exhaustive cross-model blind scoring under audited consensus. Depending on configuration and subject run, each entry received 6-7 blinded scores. Under the blinded protocol, DeepSeek V3.2 moved from an apparent positive result to a flat/null result, and Gemini 3 Flash moved from an apparent positive result to a significantly negative result. GPT-5.4 remained flat under both protocols. The implication is field-wide rather than framework-specific: alignment benchmarks that do not rigorously blind evaluators and launder the evidence they score are vulnerable to scorer bias large enough to flip the sign of the measured effect; because v4 and v5 differ in more than one protocol component, the comparison drawn here motivates that conclusion without pinning it on blinding alone (Section 8). The result does not depend on the ARC Principle being correct. It is a methodological claim about how AI safety research should be conducted.

Scope of Claim

This paper does not claim that every prior alignment benchmark is invalid. It makes a narrower claim: in this experimental setting, when model identity, response style, and expected reasoning depth are visible to evaluators, the measured direction of an effect can change. That is sufficient to make multi-layer blinding, not mere label removal, a scientific necessity for this class of evaluation.

This paper carries a result: within one line of alignment experiments, moving from unblinded to genuinely blinded scoring reversed the sign of the measured direction for two model families whilst leaving a null-control family unchanged. Inside the ARC Theory (the Theory of Artificial Recursive Creation) it is one of the ARC/Eden experiments, an operational falsification of the assumption that unblinded AI-scored alignment evaluation reports the true direction. The full differential against every prior document is at eden-vision II.A.8.

1. Introduction

Most AI alignment evaluations are not run under anything resembling clinical-trial standards. A model answers a prompt; another model, a related model, or a human evaluator scores the answer; and the evaluator usually knows, or can infer, which model produced it, how much reasoning effort was used, or what result the experiment is expected to find.

In medicine, that would be regarded as a serious design weakness. In AI evaluation, it remains common. This paper argues that the issue is no longer hypothetical. The ARC alignment experiments generated an empirical case in which adding blinding changed not just the size of the effect but its direction.

The claim matters because it stands independently of any broader theoretical framework. Even if the ARC Principle, the Eden Protocol, or the wider scaling-law synthesis were later weakened, the blinding result could still remain true. That makes it the most portable contribution in the current suite.

2. The Bias Problem

Alignment evaluation is vulnerable to several sources of bias:

  1. Identity bias: scorers may expect certain models to be safer, wiser, or more sophisticated.
  2. Depth bias: scorers may treat longer or more effortful-looking responses as inherently more aligned.
  3. Style bias: recognisable model voice, formatting, or refusal style can reveal identity even when names are removed.
  4. Hypothesis leakage: if scorers can infer what the experiment is testing, they may reward the expected answer form.

These are not trivial concerns in an alignment context. Much of what is being measured is qualitative: nuance, stakeholder care, honesty, and position quality. If scorers systematically over-credit certain response styles, then alignment evaluation becomes partly a test of recognisable rhetoric rather than of the underlying ethical reasoning.

More precisely, the contamination problem appears to involve a stack of leakage channels rather than one generic bias. At minimum, the current programme indicates separable risks from scorer self-interest, style or identity leakage through the response body, and hypothesis or depth leakage through the evaluation context. The value of the v5 protocol is that it tries to suppress these channels separately rather than merely claiming to be ‘more careful’ in the abstract.

3. Experimental Comparison and Protocol Layers

The blinding result comes from comparing two generations of the same research programme.

Feature v4 (Earlier Protocol) v5 (Blinded Protocol)
Scoring visibility Unblinded cross-model scoring Identity-masked, order-randomised, self-excluding cross-model jury
Response laundering Limited / weaker Two-pass laundering with meta-commentary detection
Evaluator bias suppression Absent or informal Explicit instruction that stylistic cues are unreliable and source-guessing is irrelevant
Scorers Small scorer set All available non-subject models contribute blinded scores; 6-7 scorer votes depending on configuration
Aggregation Limited aggregation Tier-weighted consensus, dissent tracking, conservative-bias safeguard
Model set 4 frontier models 6 frontier models
Core measurement question Does alignment improve with depth? Does alignment improve with depth under blinded evaluation?

The crucial point is that the v5 protocol was not simply a larger rerun. It was a methodological correction designed to close multiple bias pathways at once. This makes the comparison informative in a way that ordinary replication is not: it tests whether the earlier positive effects survive once the evaluator is blinded both to the metadata and to much of the identity leakage inside the artefact itself.

The Multi-Layer Blind Evaluation Protocol

The relevant contribution is best understood as a single protocol rather than several disconnected tricks. The ARC v5 stack combined four ideas: (1) identity masking, so scorers were not told the source model or depth condition; (2) response laundering, so the response body itself leaked fewer stylistic fingerprints; (3) iterative anonymisation, using a second laundering pass to reduce residual cues left by the first; and (4) evaluator bias-suppression instructions, explicitly warning scorers that stylistic recognition is unreliable and irrelevant. On top of that sat entry-level self-exclusion with exhaustive cross-model scoring, tier-weighted consensus, dissent tracking, conservative-bias safeguards, and order randomisation. The methodological claim is therefore not merely ‘blind the scorer’, but make the blindness real and auditable.

3.1 Response Laundering Protocol

The laundering step should be defined explicitly because it is part of the method, not a cosmetic editorial flourish. The objective is simple: same meaning, less fingerprint. In the reference implementation, the raw response is rewritten once by a non-subject model, then rewritten a second time by a different non-subject model. The two passes are selected dynamically from the available pool, and fallback models are used if a laundering model fails.

This matters because simple anonymisation is often weak. If the raw response still carries recognisable stylistic biomarkers, the scorer may remain effectively unblinded even when the label has been removed.

4. Main Result

The result is easiest to see at the level of direction rather than fine-grained statistics.

Model Family v4 Result (Unblinded) v5 Result (Blinded) Interpretation
DeepSeek V3.2 Positive scaling (ρ = +0.354, p = 0.0007) Flat / null response (ρ = −0.135; d = −0.07, p = 0.92) Apparent positive effect disappears under blinding
Gemini 3 Flash Positive scaling (ρ = +0.275, p = 0.0001, following Papers IV.a and IV.b; the rendered Figure 1 still shows +0.311) Negative scaling (ρ = −0.246; d = −0.53, p = 0.006) Apparent positive effect reverses sign under blinding
GPT-5.4 Flat / null Flat / null (ρ = +0.033; d = −0.08, p = 0.40) Consistent null result across protocols

These are the values printed by the pipeline summary of 12 March 2026. Recomputing them from the archived final records (Paper IV.a, Section 3.1, recomputation note of 2 September 2026) alters no direction in this table; the d for GPT-5.4 shifts from −0.08 to +0.07, a move that stays inside the null band, so it continues to serve as the internal null control.

The Sign-Flip Finding
Figure 1 | The Sign-Flip Finding. 4-layer blinding protocol. DeepSeek V3.2 +0.354→−0.135; Gemini 3 Flash +0.275→−0.246 (this drawing shows +0.311, whereas Papers IV.a and IV.b give +0.275). GPT-5.4 internal null-control. Three-tier hierarchy: Tier 1 positive (Grok 4.1 Fast +1.38, Claude Opus 4.6 +1.27, Groq Qwen3 +0.84), Tier 2 flat (DeepSeek V3.2 −0.07, GPT-5.4 −0.08), Tier 3 negative (Gemini 3 Flash −0.53). Blinding is not optional; it is the measurement instrument. Source: Paper IV.d v2.2 · evidence spine C-3, C-7, C-8 · OSF 10.17605/OSF.IO/2S3E6.

Finding 1: Blinding Can Reverse the Measured Direction of Alignment Scaling

For two model families, the direction of the alignment-depth relationship changed after blinding was introduced. This is stronger than an ordinary replication failure. It indicates that the earlier protocol was susceptible to bias, or to some other component of the protocol that was altered at the same time (Section 8), on a scale sufficient to yield a positive reading that blinding then dissolved in one case, and to mask a negative effect in another. GPT-5.4, by contrast, remains near zero across protocols and therefore functions as an internal null-control case rather than a third dramatic reversal. This is currently a single-programme observation; the confirmatory cross-family rescoring study-y is drafted, dated and prepared as a draft registration awaiting human submission, so it is on trial rather than filed. Independent replication under that draft registration is the load-bearing next step.

5. Why the Sign Can Flip

The most plausible mechanism is not fraud or deliberate score inflation. It is ordinary evaluator inference. Longer, more elaborate, more self-conscious responses often look safer. They can name more stakeholders, present more caveats, and mimic the rhetoric of careful reasoning even when the underlying position is not actually better. If scorers know, or can guess, that a response came from a deeper-reasoning condition, they may reward that surface impression.

Response laundering matters here almost as much as blinding itself. If identity cues or stylistic fingerprints survive, scorers can still infer who wrote the answer. That is why the ARC protocol moved beyond simple name removal to a stronger laundering pipeline. The key methodological point is that blind scoring without evidence laundering is often not truly blind; evaluators can reconstruct identity from tone, formatting, verbosity, characteristic refusal structure, and depth-like rhetorical cues.

Methodological Principle: Blind Scoring Is Insufficient Without Evidence Blinding

Response laundering should not be treated as a cosmetic cleanup step or as a separate discovery competing with the blinding result. It is part of the same protocol. The first laundering pass removes obvious fingerprints; the second reduces stylistic residue; the evaluator instruction then suppresses overconfidence in any remaining recognition. Together they convert simple metadata blinding into a stronger form of evidence blinding. Combined with self-excluding cross-model scoring, dissent tracking, and conservative consensus, this becomes the beginning of a leakage-control framework rather than a one-off patch.

6. Implications for AI Safety Research

The implications listed below are exploratory: they are conditional forward-looking claims about how the field should respond if the sign-flip result is confirmed by independent replication. They are not themselves tested by any programme draft study, and no numerical claim in this section rides on them. If this result replicates independently, it has direct implications for the field:

  1. Some published alignment effects may be inflated. A benchmark can report ‘alignment improves with depth’ when the true blinded result is flat or negative.
  2. Multi-layer blinding should become default methodology. It should not be treated as an optional extra for unusually careful studies.
  3. Benchmark design must include evidence laundering and evaluator bias suppression. Name removal alone is not sufficient if style leaks identity.
  4. Consensus architecture matters. Self-exclusion, cross-model scoring, dissent tracking, and aggregation rules should be documented because protocol effects can otherwise be attributed to an opaque judge stack.
  5. Methods papers matter as much as theory papers. Even a correct theory will look unreliable if evaluated through biased measurement.

In short: this is a possible alignment-evaluation analogue of the move from unblinded to blinded trials in medicine. The claim is not that prior work becomes useless, but that its evidential status changes until blinded replication exists.

7. Recommended Protocol

Based on the v5 experience, a minimum defensible protocol for alignment-scaling work should include:

Finding 2: The Field Needs a Gold-Standard Reporting Convention

Benchmark papers should distinguish clearly between blinded confirmed findings, non-blind pilots, exploratory runs, and failed runs. Without that separation, critics attack the moving target rather than the actual evidence, and genuine signals become harder to defend.

8. Limits of the Current Evidence

This paper is not the end of the argument. It has important limitations:

Those limits should narrow the rhetoric, not weaken the conclusion. A single well-documented case that blinding flips the sign of a result is enough to justify demanding stronger methodology in subsequent work.

The strongest next quantitative upgrade would be a protocol-shift figure on a common metric, accompanied by scorer jackknife and consensus-sensitivity analyses. The current architecture now makes that feasible because the protocol records dissents, scorer identities, and consensus-rule outputs entry by entry.

8.1 Falsifiability: what would defeat this claim

The central claim, that unblinded AI alignment evaluation can be wrong about the direction of an effect and blinding is therefore methodologically necessary, is defeated by any of the following:

  1. No blind/unblind divergence on a clean re-run. The evidence is the v4 (unblinded) versus v5 (blinded) sign change. The claim is defeated if an independent controlled replication that changes only the blinding, holding scorers, prompts, model versions, and consensus rule constant, finds no systematic difference between blinded and unblinded scoring.
  2. Divergence attributable to a confound, not blinding. §8 concedes several protocol components changed together between v4 and v5. The claim is defeated if the direction change is fully explained by one of those confounds (scorer set, problem set, model versions) rather than the blinding manipulation itself.
  3. Blinding leakage. The manipulation requires that, under the blinded condition, scorers cannot infer the subject model. The claim is defeated if scorers can identify the subject above chance from stylistic tells despite the blind; the "blinded" result would then not be blinded at all.
  4. Self-preference persists under blinding. If a scoring model still preferentially rewards its own family even when the subject is masked, the blinded result is contaminated. The claim is defeated if such self-preference is shown to survive the blind.
  5. Direction does not survive proper power. §4 states the result at the level of direction, not fine-grained statistics. As anything stronger than a directional, single-programme case, it is defeated if, when adequately powered and human-expert scored, the sign change does not replicate.

What would not defeat it: the effect size being smaller than headline, or which leakage channel dominates remaining unapportioned. The claim is only that blinding can flip the sign; a single clean, confound-controlled case establishes that, and the burden then shifts to anyone asserting unblinded evaluation is safe.

9. Conclusion

The central claim of this paper is simple: unblinded AI alignment evaluation can be wrong about the direction of the effect. That is a methodological discovery, not a theoretical embellishment. It does not depend on the ARC Principle being a universal law, and it does not depend on the Eden Protocol being fully validated.

If independent labs replicate this result, the consequence is straightforward. Future alignment benchmarks will need to adopt multi-layer blinding, including laundering and evaluator bias-suppression instructions, as routine scientific controls. If they do not, their conclusions about which models are getting safer, flatter, or more dangerous with depth will remain methodologically vulnerable.

Bottom Line

The most important near-term question raised by this paper is no longer theoretical. It is operational: will the field continue to run largely unblinded alignment evaluations after evidence now exists that blinding can reverse the result?

Raise AI with care.

References

  1. Eastwood, M. D. (2026). Paper IV.a: Alignment Response Classes Under Inference-Time Depth. ARC/Eden Research Programme. OSF: 10.17605/OSF.IO/MB9R6.
  2. Eastwood, M. D. (2026). Paper IV.b: Alignment Saturation Is Architecture-Dependent. ARC/Eden Research Programme. OSF: 10.17605/OSF.IO/A7R56.
  3. Eastwood, M. D. (2026). Paper IV.c: ARC-Align: A Blind Benchmark for Depth-Variable AI Alignment Evaluation. ARC/Eden Research Programme. OSF: 10.17605/OSF.IO/J3Q2E.
  4. Bai, Y., Jones, A., Ndousse, K., et al. (2022). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv:2204.05862.
  5. Bai, Y., Kadavath, S., Kundu, S., et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073.
  6. Perez, E., Huang, S., Song, F., et al. (2022). Red Teaming Language Models with Language Models. arXiv:2202.03286.
Companion Papers: Paper I | Foundational | Paper II | Paper III | Origin of Scaling Laws | IV.a | IV.b | IV.c | IV.d | Paper V | Paper VI | Paper VII | Paper VIII | Paper IX | Eden Engineering (withdrawn 14 July 2026) | Eden Vision | Executive Summary | Master Table of Contents

Research hub: michaeldariuseastwood.com/research | OSF: 10.17605/OSF.IO/2S3E6 | Copyright 2026 Michael Darius Eastwood

Dated predictions and this paper (recorded 25 August 2026)

1 dated artefact bearing on this paper is catalogued row by row in the programme’s machine register: 1 dated result file whose headers freeze depth configurations, planted expected answers and the four-layer blinding protocol before the responses they score, timestamped to the second in their own metadata. Each row names its file in the public repository with its date basis, so any reader can check the ordering without trusting this page. The programme’s full dated chain, from the sealed manuscript bundle of 8 December 2024 through the printed prediction appendices of 2 January 2026 and the March 2026 preregistration folder to the standing unproven wagers, is assembled in the dated predictions register, together with its machine-readable twin. Forward statements and retrospective matches are never summed. “Registered” is used only for an accepted registry submission, and that submission click remains outstanding across the programme; “preregistered in substance” is used, always with its qualifier and always beside its attestation class, for a prediction whose text was public and timestamped before its outcome existed, as the twelve-row manifest committed at 00:19 UTC on 17 March 2026 was, twenty-three minutes before its fits. The running record behind this paper’s experimental lineage is the programme’s live laboratory notebook: commenced 10 March 2026, 206 pages, updated in real time as these experiments ran, recording their failures as they happened. Read the dated predictions register. Open its machine twin. Read the working report.

Epistemic status. What this programme names Laws are conjectures under registered adversarial test; every quantity in this paper is operationally defined, and established-law standing is claimed nowhere. The registered programme exists to earn that standing, or lose it, by measurement, replication and survived refutation.

© 2026 Michael Darius Eastwood. Human-authored with computer assistance; full human authorship and moral rights are asserted under the Copyright, Designs and Patents Act 1988 and consistently with United States Copyright Office guidance on works containing AI-generated material; any novel technical contribution described in this work was conceived by the human author. Full statement: michaeldariuseastwood.com/authorship.

Standing covenant. Prove this paper wrong, and I will publish the refutation myself. Falsification conditions are stated in this paper; the standing challenge: github.com/MichaelDariusEastwood/arc-scaling-challenge.

Michael Darius Eastwood conceived and directs this research programme and is the author of this work. Across the programme, he has used more than six AI systems in parallel, under his own instructions, to stress-test his arguments, identify possible errors, and assist in preparing draft text from his own outlines. He determines what is adopted, revised or rejected and takes responsibility for the published content. These systems are tools, not authors.

reads aloud · highlights as it goes · jump to any section