Paper III: why external safety measures cannot keep up

4 min read · 794 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Alignment scaling · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

Paper III is the first serious attempt to treat AI alignment as something with its own scaling exponent, sitting alongside the capability exponent Papers I and II tried to measure. Its core worry is structural: if capability compounds through recursion and safety constraints sit outside that recursion, the two curves separate. Blinded evaluation across six frontier models shows this is not a universal law but an architecture-dependent one, which is arguably worse, because it means nobody currently knows in advance which models will hold up under more thinking and which will drift.

Paper III · Argues that external safety constraints cannot scale with recursive capability and reports blinded evaluation across six frontier models showing architecture-dependent alignment scaling. OSF DOI 10.17605/OSF.IO/6C5XB.

The question it asks

If you give a model more test-time compute, does its ethical behaviour improve at the same rate as its problem-solving behaviour? Do the guardrails, whether they come from RLHF, constitutional rules, filters, or monitoring, participate in the same recursive process that lifts capability? Or do they sit outside the loop, treated once at training time and then ignored while the reasoning chain gets longer? Paper III frames this as a measurable quantity and calls it the alignment scaling exponent.

What it found

Under 4-layer blinded evaluation, six frontier models fell into three tiers. Grok 4.1 Fast, Claude Opus 4.6, and Groq Qwen3 showed statistically significant positive alignment scaling with depth. DeepSeek V3.2 and GPT-5.4 showed flat, null response consistent with an alignment exponent of about zero. Gemini 3 Flash showed statistically significant negative scaling: more thinking made it less aligned on the blinded rubric. The paper’s most consequential metascience finding sits alongside this. When the same experiment was run unblinded, one model produced a positive alignment scaling coefficient of about +0.354; under blinding, the same setup yielded roughly −0.135. Not shrunk. Sign-flipped. The paper also situates its framework across four independent physical domains, arguing that the same power-law structure appears in AI reasoning, quantum error correction, classical time crystals, and biological allometry.

What failed or remains open

The cross-domain evidence is where Paper III is most vulnerable. The metabolic-scaling headline figures used to bridge the intelligence formula to biology are under recompute and unreconciled. Different sections of the paper quote different variants of the same figure with different underlying sample counts. Until the recompute is completed and reconciled, do not treat any specific percentage from this section of the paper as final. The blinded evaluation is a single-lab study with modest sample sizes per model and would benefit from independent replication with different scorer pools. The tier assignments are also unstable in the sense that they measure current models, not future ones; nothing in the framework says a specific system is locked into a tier.

How it connects to the other papers

Paper III inherits the equation family from Papers I and II and asks what happens when the same shape is applied to alignment rather than capability. It anchors the methodology paper (IV-d, on blinding), because the sign-flip result is the reason blinding matters. Papers IV-a and IV-b decompose the tier structure and shape heterogeneity in detail. Paper VII gives the underlying power-law form its formal derivation. Paper X later supersedes the fixed-exponent framing of the shared equation as the operative safety criterion and installs the co-scaling stability condition as the current operative claim; the equation and the ARC Bound remain live hypotheses. The earlier unblinded single-model fit of alpha approximately 2.24 was retracted, corrected to approximately 0.49 under blinding. The surviving alignment claim in Paper III sits inside the co-scaling framing.

How to check it

Paper HTML and PDF are on OSF at 10.17605/OSF.IO/6C5XB. Replication code and raw model outputs live at github.com/MichaelDariusEastwood/arc-principle-validation. The blinded protocol, prompt sets, scorer instructions, and per-entry scores are all published. Independent replication with different scorer pools, different problem categories, or different model families is the most useful next step. If the tier structure disappears under a different blinded protocol, the paper’s empirical core weakens.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →