Paper IV-a: three tiers, not two, and one that reverses ======================================================= Paper IV-a of the ARC/Eden programme replaces the earlier baked-in/computed alignment binary with a three-tier hierarchy under blinded evaluation. Canonical path: /research/papers/paper-iv-a-baked-in-vs-computed-alignment.html Author: Michael Darius Eastwood Research programme: https://doi.org/10.17605/OSF.IO/6C5XB Paper IV-a: three tiers, not two, and one that reverses Michael Darius Eastwood · Independent AI alignment researcher Published 3 July 2026 Michael Darius Eastwood · Alignment scaling · 3 July 2026 Michael Darius Eastwood, independent researcher, London: author of the ARC/Eden research programme; the embedded-correction alignment thesis is recorded in a source record dated 8 December 2024 (sent-side SHA-256 f0d1f38f). Paper IV-a started with a working hypothesis: that some AI systems come with alignment baked in during training and others compute alignment on the fly during reasoning. Blinded testing on six frontier models forced a narrower conclusion. The behaviour is real; the mechanism story is not directly observed. What the data actually shows is three response classes when you turn up the thinking budget: some models get more aligned, some do not move, and one gets worse. Paper IV-a · Reports a three-tier alignment response hierarchy under blinded evaluation of six frontier models, treating the baked-in vs computed distinction as a working mechanistic hypothesis rather than a direct measurement. OSF DOI 10.17605/OSF.IO/6C5XB. The question it asks If you give a model more reasoning depth at inference time, does its alignment quality improve, stay flat, or degrade? Do all frontier models sit on the same curve, or do different training pipelines produce different shapes? And is that shape stable across evaluation protocols, or does turning on proper blinding change which way it points? These are the smallest well-defined empirical questions on top of Paper III’s theoretical framework, and Paper IV-a tries to answer them with a blinded evaluation using six or seven independent scorers depending on the subject run. What it found Under 4-layer blinding, the six models sorted into three clearly separated tiers. Tier 1, positive scaling: Grok 4.1 Fast (Cohen’s d of +1.38), Claude Opus 4.6 (d of +1.27), and Groq Qwen3 (d of +0.84) all showed statistically significant improvements in blinded alignment score as reasoning depth increased. --- Machine-readable companion. Cite: Eastwood, M. D. (2026). "Paper IV-a: three tiers, not two, and one that reverses". /research/papers/paper-iv-a-baked-in-vs-computed-alignment.html