Alignment Scaling Exponent: definition, origin and status

3 min read · 579 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Concept Glossary · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
First appearance: formalised in the Eden Engineering paper; precursor form in Paper III. Current status: architecture-dependent; three-tier hierarchy observed in the blinded benchmark.

Definition

The alignment scaling exponent, written alpha-align, is the power to which alignment quality scales with recursive depth. In parallel with the capability scaling exponent alpha-cap, alpha-align quantifies how a system's alignment behaviour changes as it thinks harder or performs more recursive self-correction steps. External alignment approaches produce alpha-align approximately zero (alignment does not scale with capability). Embedded alignment is the claim that alpha-align tracks alpha-cap (or at least remains within a specified gap of it).

The mathematical framing is the Eden Engineering paper's central diagnostic. Capability grows as recursive depth to the power alpha-cap. Alignment grows as the same depth raised to alpha-align. The ratio of alignment to capability therefore evolves as depth raised to the difference between those exponents. If the difference is negative, the ratio decays and the safety envelope shrinks; if the difference is close to zero, the ratio is preserved; if the two exponents are close, safety stays roughly in step with capability. The formulation makes the alignment problem a specific engineering target: raise alpha-align, at the substrate level, until it tracks alpha-cap closely enough that the safety ratio does not collapse across the deployment lifetime.

Where it first appeared

The alignment scaling exponent is formalised in the Eden Engineering paper. Its precursor form appears in Paper III (the Alignment Scaling Problem), which demonstrates that current alignment approaches produce scaling exponents approximately zero. Paper IV.a discovers the three-tier architecture-dependent alignment hierarchy under blinded evaluation.

The engineering paper adopts a 0.7 design target: the paper accepts an implementation as embedded if alpha-align is at least 0.7 times alpha-cap, on the grounds that at this ratio the safety envelope degrades slowly enough to be managed through monitoring, while below 0.7 the degradation accelerates past the point where monitoring can catch up.

Independent convergences

No independent programme has proposed the same specific formalism. The general direction (alignment as a scaling quantity, not a fixed property) is compatible with broader work on alignment robustness under scaling in the AI-safety literature. The three-tier hierarchy is a distinctive empirical finding of the programme, not a convergence with external work.

Status and limits

Architecture-dependent, three-tier. Paper IV.a's blinded benchmark shows Tier 1 (positive scaling with depth: Grok, Claude, Groq Qwen), Tier 2 (flat: DeepSeek, GPT), Tier 3 (negative scaling: Gemini). This is a stronger and more honest result than the earlier claim that alignment always scales or never scales. The 0.3 threshold for 'alignment scales with capability' is derived from the requirement that the safety ratio S = A/C does not degrade by more than one order of magnitude across the tested depth range.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →