Paper III: why external safety measures cannot keep up ====================================================== Paper III of the ARC/Eden programme argues that if capability compounds through recursion, external safety constraints that do not participate in that recursion get left behind. Canonical path: /research/papers/paper-iii-alignment-scaling-problem.html Author: Michael Darius Eastwood Research programme: https://doi.org/10.17605/OSF.IO/6C5XB Paper III: why external safety measures cannot keep up Michael Darius Eastwood · Independent AI alignment researcher Published 3 July 2026 Michael Darius Eastwood · Alignment scaling · 3 July 2026 Michael Darius Eastwood, independent researcher, London: author of the ARC/Eden research programme; the embedded-correction alignment thesis is recorded in a source record dated 8 December 2024 (sent-side SHA-256 f0d1f38f). Paper III is the first serious attempt to treat AI alignment as something with its own scaling exponent, sitting alongside the capability exponent Papers I and II tried to measure. Its core worry is structural: if capability compounds through recursion and safety constraints sit outside that recursion, the two curves separate. Blinded evaluation across six frontier models shows this is not a universal law but an architecture-dependent one, which is arguably worse, because it means nobody currently knows in advance which models will hold up under more thinking and which will drift. Paper III · Argues that external safety constraints cannot scale with recursive capability and reports blinded evaluation across six frontier models showing architecture-dependent alignment scaling. OSF DOI 10.17605/OSF.IO/6C5XB. The question it asks If you give a model more test-time compute, does its ethical behaviour improve at the same rate as its problem-solving behaviour? Do the guardrails, whether they come from RLHF, constitutional rules, filters, or monitoring, participate in the same recursive process that lifts capability? Or do they sit outside the loop, treated once at training time and then ignored while the reasoning chain gets longer? Paper III frames this as a measurable quantity and calls it the alignment scaling exponent. What it found Under 4-layer blinded evaluation, six frontier models fell into three tiers. Grok 4.1 Fast, Claude Opus 4.6, and Groq Qwen3 showed statistically significant positive alignment scaling with depth. DeepSeek V3.2 and GPT-5.4 showed flat, null response consistent with an alignment exponent of about zero. --- Machine-readable companion. Cite: Eastwood, M. D. (2026). "Paper III: why external safety measures cannot keep up". /research/papers/paper-iii-alignment-scaling-problem.html