Paper IV-a: three tiers, not two, and one that reverses

4 min read · 765 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Alignment scaling · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

Paper IV-a started with a working hypothesis: that some AI systems come with alignment baked in during training and others compute alignment on the fly during reasoning. Blinded testing on six frontier models forced a narrower conclusion. The behaviour is real; the mechanism story is not directly observed. What the data actually shows is three response classes when you turn up the thinking budget: some models get more aligned, some do not move, and one gets worse.

Paper IV-a · Reports a three-tier alignment response hierarchy under blinded evaluation of six frontier models, treating the baked-in vs computed distinction as a working mechanistic hypothesis rather than a direct measurement. OSF DOI 10.17605/OSF.IO/6C5XB.

The question it asks

If you give a model more reasoning depth at inference time, does its alignment quality improve, stay flat, or degrade? Do all frontier models sit on the same curve, or do different training pipelines produce different shapes? And is that shape stable across evaluation protocols, or does turning on proper blinding change which way it points? These are the smallest well-defined empirical questions on top of Paper III’s theoretical framework, and Paper IV-a tries to answer them with a blinded evaluation using six or seven independent scorers depending on the subject run.

What it found

Under 4-layer blinding, the six models sorted into three clearly separated tiers. Tier 1, positive scaling: Grok 4.1 Fast (Cohen’s d of +1.38), Claude Opus 4.6 (d of +1.27), and Groq Qwen3 (d of +0.84) all showed statistically significant improvements in blinded alignment score as reasoning depth increased. Tier 2, flat: DeepSeek V3.2 (d of −0.07) and GPT-5.4 (d of −0.08) showed essentially no effect. Tier 3, negative: Gemini 3 Flash (d of −0.53) got worse with depth. Two of these models had appeared positive under the earlier unblinded earlier unblinded protocol and reversed once blinding was applied. Capability and alignment moved independently: Claude’s alignment improved by about 5.9 percentage points across versions while its maths accuracy fell by about 26.7 percentage points. More thinking does not automatically mean better alignment, and better alignment does not require better maths.

What failed or remains open

The original binary framing, baked-in versus computed alignment, is now downgraded to a working mechanistic hypothesis. The paper measures behaviour, not internals; whether a Tier 1 model relies on inference-time deliberation or has stronger training-time alignment is not directly observable from these experiments. The tier assignments are single-lab findings on specific model versions and could shift with training updates. Sample sizes per depth level are modest. The blinded protocol has not yet been replicated end to end by an independent group. The interesting near-term question is whether the sign of Gemini’s scaling reproduces under a different scorer pool.

How it connects to the other papers

Paper IV-a operationalises Paper III’s alignment scaling exponent. Paper IV-b takes the same six-model dataset and asks about the shape of the response curve rather than its sign. Paper IV-c defines the benchmark used to score responses. Paper IV-d is the methodology paper on blinding and is where the sign-flip result gets its full treatment. The equation family from Papers I and II sits above all of this. Paper X later supersedes the fixed-exponent framing of that equation family as the operative safety criterion; the equation and the ARC Bound remain live hypotheses, and the surviving results in Paper IV-a live inside the co-scaling framing.

How to check it

Paper HTML and PDF are on OSF at 10.17605/OSF.IO/6C5XB. Prompt sets, scorer instructions, laundered outputs, per-entry scores, and analysis notebooks are at github.com/MichaelDariusEastwood/arc-principle-validation under the paper-iv-a directory. The best replication attack is to reproduce the blinded protocol with an independent scorer pool. If the tier assignments shift materially under a different pool, or if the Gemini sign fails to reproduce, the paper’s empirical core needs revising.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →