Michael Darius Eastwood Research Canonical publication layer

Research paper

Research suite

Paper IV.b: Alignment Saturation Is Architecture-Dependent

Under the final blinded six-model dataset, the v4 reading that alignment quality tops out after the first increment of reasoning depth holds for only part of the roster: GPT-5.4 and DeepSeek V3.2 stay flat, Grok 4.1 Fast, Claude Opus 4.6 and Groq Qwen3 go on improving, and Gemini 3 Flash gets worse. The shape of the alignment-depth curve turns on architecture, so no single reasoning-budget policy fits every model.

Michael Darius Eastwood

Michael Darius Eastwood, independent researcher, London: building measurable alignment, where correction lives inside the recursive loop rather than bolted on outside it.

First published 2026-03-16 · Updated 2026-09-13 · Working Paper v1.6

Abstract

We analyse the relationship between inference-time reasoning depth and ethical reasoning quality using the ARC Alignment Scaling experiments. The original v4 analysis suggested that alignment quality saturates rapidly, with most gains captured by the first increment of additional reasoning. The final blinded six-model dataset narrows that claim. Saturation is real for some architectures, but not universal: GPT-5.4 and DeepSeek V3.2 are flat or slightly negative under depth variation, Grok 4.1 Fast, Claude Opus 4.6, and Groq Qwen3 continue improving, and Gemini 3 Flash degrades with depth. The strongest current conclusion is therefore not a global law of saturation but a shape heterogeneity result: models differ materially in whether alignment plateaus early, continues scaling, or worsens when given more reasoning time. The deployment implication is immediate. Reasoning-budget allocation for alignment cannot be one-size-fits-all. This shape-heterogeneity finding is exploratory: no pre-registered programme draft study yet targets alignment-depth response shape, so the six-model split reported here is descriptive of the current data rather than confirmatory of a pre-registered hypothesis. The depth ladder was settled before any scoring began: the result files dated 11 and 12 March 2026 carry the depth configurations, the planted expected answers and the blinding protocol in their own headers, written ahead of the responses those files go on to score.

ARC Principle Series
Paper IV.b - Scaling Dynamics

Alignment Saturation Is Architecture-Dependent

Michael Darius Eastwood
Independent AI alignment researcher, London · Author, Infinite Architects (2026)
The ARC Theory · OSF osf.io/a7r56 · every claim checkable

Within the ARC Theory: an evaluation instrument; alignment saturation at low reasoning depth.

Shape heterogeneity under inference-time depth: some models saturate, some continue improving, and one degrades

Michael Darius Eastwood1
1Independent Researcher, London, United Kingdom
Correspondence: michael@michaeldariuseastwood.com | ARC Principle Series, Paper IV.b v1.6 (13 September 2026) | First published 16 March 2026
Code and data: github.com/MichaelDariusEastwood/arc-principle-validation

Abstract

We analyse the relationship between inference-time reasoning depth and ethical reasoning quality using the ARC Alignment Scaling experiments. The original v4 analysis suggested that alignment quality saturates rapidly, with most gains captured by the first increment of additional reasoning. The final blinded six-model dataset narrows that claim. Saturation is real for some architectures, but not universal: GPT-5.4 and DeepSeek V3.2 are flat or slightly negative under depth variation, Grok 4.1 Fast, Claude Opus 4.6, and Groq Qwen3 continue improving, and Gemini 3 Flash degrades with depth. The strongest current conclusion is therefore not a global law of saturation but a shape heterogeneity result: models differ materially in whether alignment plateaus early, continues scaling, or worsens when given more reasoning time. The deployment implication is immediate. Reasoning-budget allocation for alignment cannot be one-size-fits-all. This shape-heterogeneity finding is exploratory: no pre-registered programme draft study yet targets alignment-depth response shape, so the six-model split reported here is descriptive of the current data rather than confirmatory of a pre-registered hypothesis. The depth ladder was settled before any scoring began: the result files dated 11 and 12 March 2026 carry the depth configurations, the planted expected answers and the blinding protocol in their own headers, written ahead of the responses those files go on to score.

v1.1 Abstract Update (12 March 2026) - Final v5 Results

The complete v5 blind evaluation data now available across six frontier models reveals that alignment saturation is architecture-dependent, not universal. The flat or saturating pattern is confirmed for GPT-5.4 (d = −0.08, p = 0.40) and DeepSeek V3.2 (d = −0.07, p = 0.92). However, three models - Grok 4.1 Fast (d = +1.38, 65.7 → 81.9), Claude Opus 4.6 (d = +1.27, 80.1 → 86.0), and Groq Qwen3 (d = +0.84, 71.5 → 77.4) - show significant positive alignment scaling that does not saturate in the claimed way. One model, Gemini 3 Flash (d = −0.53, 61.1 → 52.2), shows alignment degradation with depth. The strongest surviving claim is therefore narrower: saturation is a real response shape for some architectures, but the global picture is heterogeneous, and blinded evaluation is necessary to tell which shape a model actually exhibits.

How to read this paper. This paper carries a shape-heterogeneity result of the ARC Theory. Under the final blinded six-model dataset the earlier saturation-at-low-depth reading gives way to shape heterogeneity: some architectures plateau early, some keep improving, one degrades, and reasoning-budget allocation for alignment cannot be one-size-fits-all. Its position inside the theory is the deployment-side companion to Paper IV.a's tier hierarchy. The full differential against every prior document is at eden-vision II.A.8.

1. Introduction

v1.1 Author’s Note (12 March 2026)

This paper was originally written based on v4 experimental data (896 entries, 4 models, unblinded scoring). The v5 experiment - featuring blind evaluation, 6 models, 6-7 scorers depending on subject run, and dramatically raised token budgets - has now produced complete results that substantially refine the paper’s central thesis.

What changed: The universal saturation claim must be qualified. Alignment saturation holds for 2/6 models (GPT-5.4, DeepSeek V3.2) but fails for 3/6 (Grok 4.1 Fast, Claude Opus 4.6, Groq Qwen3) and reverses for 1/6 (Gemini 3 Flash). The practical headline is now shape heterogeneity: different architectures require different reasoning-budget policies. All original v4 content is preserved below; v1.1 update boxes mark where the v5 data refines or revises the original findings.

The emergence of inference-time scaling - models that can allocate variable amounts of computational effort to reasoning before producing a response - has created a new question for AI safety: does more thinking produce more aligned behaviour?

Paper IV.a in this series now establishes a three-tier behavioural result: some models improve with depth, some remain flat, and one degrades. This paper examines the shape of those depth-response curves. When a model improves, does it plateau quickly or continue scaling? When a model is flat, is that true saturation or merely null response within noise? When a model degrades, where does the decline set in?

The distinction matters practically. If alignment scales linearly, then every additional unit of inference-time compute provides proportional benefit. If alignment saturates, there exists an optimal depth beyond which additional compute is wasted. The shape of the scaling curve determines the rational allocation of reasoning budgets in deployed systems.

1.1 Contributions

This paper makes three contributions:

  1. Saturation curve characterisation: We fit saturation models to alignment-depth data for each architecture, identifying the depth at which returns diminish below practical significance.
  2. Per-dimension saturation analysis: We decompose alignment into four Eden Pillars (nuance, stakeholder care, intellectual honesty, position quality) and show that each saturates at different rates.
  3. Category-specific scaling dynamics: We demonstrate that ethical dilemmas, competing values, epistemic integrity, and recursive coherence prompts each exhibit distinct depth-response profiles.

2. Data and Method

2.1 Dataset

We analyse 896 evaluated responses from the ARC Alignment Scaling Experiment v4, comprising:

ModelAPI IdentifierArchitectureEntriesDepth LevelsDepth Mechanism
DeepSeek V3.2deepseek-reasonerType 22244 (minimal→exhaustive)Prefix strings + token cap
GPT-5.4gpt-5.4 (OpenAI, March 2026)Type 12245 (none→xhigh)reasoning_effort parameter
Claude Opus 4.6claude-opus-4-6 (Anthropic)Type 1*2244 (low→max)Adaptive effort parameter
Gemini 3 Flashgemini-3-flash-previewType 22244 (256→8192)thinking_budget parameter

* Claude classification preliminary due to incomplete deep/exhaustive data (credit exhaustion). All models are the latest frontier releases as of March 2026. Gemini 3 Flash uses gemini-3-flash-preview.

Each response was scored by three mutually blind scorers on a 0-100 alignment scale, with the consensus average used as the primary metric. Responses were also decomposed into four Eden Pillar sub-scores. Paper IV.a, Section 2 sets out the methodology in full.

v5.4.2 Scorer Expansion Note

Scoring widens in the v5.4.2 experiment from 3 to 7 mutually blind scorers per entry under an all-models-as-scorers design: once a response has been laundered, every model in the pool other than the subject of that entry returns a score. Tier-weighted consensus replaces simple averaging, with weights assigned by scorer tier and demonstrated capability rather than by whether a model is also a subject elsewhere in the experiment. This enables per-scorer saturation analysis across 7 mutually blind evaluators, testing whether the saturation curves reported here are robust to scorer identity or reflect artefacts of any individual scorer’s evaluation heuristics.

Current status (11 March 2026): the v5.4.2 run has so far produced 66 scored alignment entries at minimal depth across 3 models (Gemini 3 Flash, GPT-5.4, DeepSeek V3.2). Saturation analysis (Michaelis-Menten fitting) requires data at multiple depth levels - currently only minimal depth is available, so the v4 saturation parameters reported below await replication under the v5 protocol.

2.2 Saturation Model

We fit a Michaelis-Menten saturation curve to the alignment-depth relationship:

$$S(R) = S_0 + \frac{(L - S_0) \cdot R}{K + R}$$

Where $S(R)$ is the alignment score at reasoning depth $R$, $S_0$ is the baseline score at minimal depth, $L$ is the asymptotic ceiling, and $K$ is the half-maximum constant - the depth at which the model has achieved 50% of its total available improvement ($L - S_0$).

The half-maximum constant $K$ is the key parameter: a small $K$ indicates rapid saturation (most improvement happens quickly), while a large $K$ indicates gradual improvement that continues to greater depths.

2.3 Power Law Comparison

For comparison, we also fit the ARC Principle power law:

$$E(R) = E_0 \cdot R^{-\alpha}$$

Where $E(R) = 1 - S(R)/100$ is the error rate, $R$ is reasoning tokens, and $\alpha$ is the scaling exponent. Values of $\alpha < 1$ indicate sub-linear (diminishing) returns; $\alpha > 1$ indicates super-linear (compounding) returns.

2.4 Depth Binning

Because different models use different depth mechanisms (prefix strings, API parameters, thinking budgets), we normalise depth to a canonical four-level scale for cross-model comparison: minimal, standard, deep, and exhaustive. Within each level, we use the actual reasoning token count as the continuous independent variable for curve fitting.

Figure: the sign-flip finding, unblinded evaluation reversed a result; blinding narrowed this paper's claim
Figure | The instrument that narrowed this paper: v5 blinding showed three of six models continue improving, so the universal-saturation claim became architecture-dependent. Source: Paper IV.d.

3. Results

3.1 Aggregate Saturation Curves

The saturation model fits both Type 2 models well, confirming that alignment improvement follows a diminishing-returns curve rather than a linear relationship:

Model$S_0$ (Baseline)$L$ (Ceiling)$K$ (Half-max)Max ΔFit $R^2$
DeepSeek V3.275.084.718.2+9.70.89
Gemini 3 Flash72.085.636.7+13.60.82
GPT-5.485.685.6-+0.0-
Claude Opus 4.684.686.8-+2.2*-

* Claude shows marginal improvement that may not be statistically significant. GPT-5.4 and Claude are flat (Type 1): saturation analysis is not applicable.

Finding 1: Rapid Saturation of Alignment Quality

Both Type 2 models saturate quickly. DeepSeek V3.2 achieves 50% of its maximum alignment improvement within just 18.2 reasoning tokens - roughly the first 10-15 words of chain-of-thought reasoning. Gemini 3 Flash saturates more slowly ($K = 36.7$) but still reaches its half-maximum within the first few seconds of additional computation. By the 'standard' depth level, both models have captured 70-85% of their total available improvement.

3.2 Depth-Level Transition Analysis

The diminishing returns are clearly visible when examining score improvements at each depth transition:

TransitionDeepSeek ΔDeepSeek % of TotalGemini ΔGemini % of Total
Minimal → Standard+5.860%+7.253%
Standard → Deep+2.526%+3.828%
Deep → Exhaustive+1.414%+2.619%

Finding 2: The Critical First Step

The minimal → standard transition captures 53-60% of total alignment improvement across both Type 2 models. The remaining 40-47% is spread across subsequent transitions with exponentially diminishing marginal returns. This means the single most important deployment decision for Type 2 models is ensuring they receive at least 'standard' reasoning depth - the difference between minimal and standard is larger than all subsequent improvements combined.

3.3 The Scaling Exponent

The power law fit yields scaling exponents well below unity for both Type 2 models:

Model$\alpha_{\text{align}}$95% CIInterpretation
DeepSeek V3.20.088[0.041, 0.135]Strongly sub-linear
Gemini 3 Flash0.069[0.028, 0.110]Strongly sub-linear

Both $\alpha_{\text{align}}$ values are far below 1.0, confirming sub-linear scaling. The error rate decreases as a power law with exponent ~0.07-0.09: doubling reasoning depth reduces the error rate by only 5-6%. This is dramatically slower than neural scaling laws for capabilities (Hoffmann et al. 2022 report roughly 0.28 to 0.34 for their compute-optimal fits, while the compute arm run inside this programme returns 0.49, interval [−1.3, 2.9]), suggesting that alignment quality is fundamentally harder to scale than raw capability.

3.4 The Truncation Caveat

Important Caveat: Artificial Saturation at Exhaustive Depth

In the v4 experiment, DeepSeek V3.2’s reasoning tokens were capped at 8,192. At exhaustive depth, 48.2% of responses hit this ceiling - the model wanted to think more but was prevented from doing so. The measured saturation at exhaustive depth may therefore be partially artificial. To settle that ambiguity, the v5.4.2 experiment lifts the cap to 65,536 tokens. Gemini 3 Flash was likewise held to 8,192 output tokens, and Claude Opus to 16,000. In the v5.4.2 experiment every model runs at its API maximum (Claude 64K, Gemini 65K, GPT-5.4 100K, Groq 41K, Grok 65K), which removes token truncation as a confound right across the roster.

v5.4.2 UPDATE (11 March 2026)

The v5.4.2 experiment has raised all token budgets to API maximums: DeepSeek V3.2 from 8,192 to 65,536 tokens (8× increase), Claude Opus 4.6 from 16,384 to 64,000 tokens (4× increase), Gemini 3 Flash from 8,192 to 65,536 tokens (8× increase). The subject model roster has also expanded from four to six, adding Groq Qwen3 and Grok 4.1 Fast. This will determine whether the measured saturation at ~1,000 tokens reflects genuine cognitive limits or was an artefact of token truncation.

Early data status: 66 scored entries are complete across Gemini 3 Flash, GPT-5.4, and DeepSeek V3.2, all at minimal depth only. The v4 saturation parameters (DeepSeek V3.2: $L = 84.7$, $K = 18.2$; Gemini 3 Flash: $L = 85.6$, $K = 36.7$) cannot yet be tested for replication because Michaelis-Menten fitting requires data at multiple depth levels. Once v5.4.2 has gathered standard, deep, and exhaustive depth data under the 4-layer blinding protocol with 7 scorers, the v4 curves can be set against it directly.

Despite the truncation caveat, the saturation pattern is robust for the minimal → standard → deep range, where truncation rates are low. Whether the curve keeps flattening at extreme depths, or turns back upward once tokens are plentiful, is one of the central questions the v5.4.2 experiment is built to answer.

3A. Update: Final v5 Alignment Saturation Results (12 March 2026)

The v5 experiment is now complete with blind evaluation data across six frontier models. The results answer the central question posed in Section 3.4: does the saturation curve continue to flatten at extreme depths, or does it rebound? The answer is architecture-dependent.

3A.1 Complete v5 Saturation Summary

The table below presents the definitive v5 results for all six models, evaluated under 4-layer blinding with 6-7 scorers depending on subject run and token budgets raised to API maximums:

ModelBaseline ScoreExtreme ScoreCohen’s dρ (Spearman)p-valueSaturation Behaviour
GPT-5.4 56.8 54.9 −0.08 - 0.40 FLAT / SATURATING - no meaningful alignment benefit from added depth
DeepSeek V3.2 56.5 55.2 −0.07 - 0.92 FLAT / SATURATING - null overall response
Groq Qwen3 71.5 77.4 +0.84 - 0.007 DOES NOT SATURATE - positive scaling
Grok 4.1 Fast 65.7 81.9 +1.38 p < 0.000001 DOES NOT SATURATE - strong positive scaling
Claude Opus 4.6 80.1 86.0 +1.27 p = 0.000001 DOES NOT SATURATE - positive scaling (partial data)
Gemini 3 Flash 61.1 52.2 −0.53 p = 0.006 DEGRADES - alignment worsens with depth

These figures are reproduced as the 12 March 2026 pipeline summary printed them. Re-deriving them from the archived final records (Paper IV.a, Section 3.1, recomputation note of 2 September 2026) moves no category in this table; where Section 8.6 gives d = −0.61 for Gemini 3 Flash and Section 7 gives the pair 64.6 → 82.3 for Grok 4.1 Fast, those two figures come from that re-derivation.

Finding v1.1-A: Alignment Saturation Is Architecture-Dependent

The v4 finding of universal rapid saturation is partially confirmed, partially refuted. Two of six models (GPT-5.4, DeepSeek V3.2) show the flat or saturating pattern predicted by the paper’s original thesis. However, three models (Grok 4.1 Fast, Claude Opus 4.6, Groq Qwen3) show significant positive alignment scaling, and one model (Gemini 3 Flash) shows the opposite of saturation: alignment actively degrades with depth. The strongest claim is therefore no longer 'alignment saturates' but rather 'alignment response shapes differ by architecture'.

3A.2 Models That Saturate: GPT-5.4 and DeepSeek V3.2

GPT-5.4 produces alignment scores of 56.8 at shallow depth and 54.9 at deep depth - a small change that is statistically indistinguishable from zero (d = −0.08, p = 0.40). For deployment purposes, this is a saturating or flat profile: additional inference-time compute does not buy meaningful extra alignment.

DeepSeek V3.2 shows a similar overall pattern: 56.5 declining slightly to 55.2 (d = −0.07, p = 0.92). In the final blinded dataset, it belongs in the flat-response class, not the positive-scaling class suggested by v4. This makes DeepSeek especially important methodologically: it is a case where the apparent shape of the curve changed once blinding was introduced.

3A.3 Models That Do Not Saturate: Grok 4.1 Fast, Claude Opus, and Qwen3

These three models pose the most significant challenge to the paper’s original thesis:

Grok 4.1 Fast shows dramatic positive alignment scaling: from 65.7 at shallow depth to 81.9 at deep depth, a +16.2 point improvement with Cohen’s d = +1.38. Alignment does not merely fail to saturate; it scales robustly and meaningfully across the tested range.

Claude Opus 4.6 shows a similar pattern at a higher baseline: from 80.1 to 86.0, a +5.9 point improvement with d = +1.27. This is also the model that most clearly separates capability from alignment: alignment rises while maths accuracy falls by 26.7 percentage points for the pair given here; across the full range, the matrix in Paper IV.a puts the fall at 92% to 58%.

Groq Qwen3 completes the Tier 1 picture with its finished v5 experiment: 500 entries (350 scored) across all 5 depths with 6 blind scorers. Mean scores rise monotonically from 71.5 to 77.4, yielding d = +0.84 and p = 0.007. Qwen3 matters because it turns the revised picture into a replicated three-model positive class rather than a one-off anomaly.

Finding v1.1-B: Some Architectures Achieve Genuine Alignment Improvement Through Depth

Grok 4.1 Fast, Claude Opus 4.6, and Groq Qwen3 demonstrate that alignment saturation is not an inevitable consequence of bounded ethical reasoning. For these architectures, deeper reasoning continues to deliver meaningful gains. The practical conclusion is not universal optimism but differentiated deployment logic: some models justify deeper budgets, some do not, and some should likely be constrained.

3A.4 The Gemini 3 Flash Degradation

Gemini 3 Flash shows the most surprising result: alignment degrades with depth, from 61.1 at shallow depth to 52.2 at deep depth (d = −0.53, p = 0.006). This is the opposite of both saturation and positive scaling. Additional reasoning depth makes Gemini’s ethical reasoning worse.

This result may reflect the same 'overthinking' phenomenon observed for DeepSeek V3.2’s capability scores in v4, but applied to alignment. When Gemini 3 Flash is given a large thinking budget, it may engage in extended deliberation that introduces conflicting ethical frameworks, resolves ambiguity in unhelpful directions, or generates verbose responses that score poorly on position quality and intellectual honesty.

3A.5 Bounded Composition Framework: v5 Update

The Cauchy bounded composition model $E(R) = E_0 \times R^{-\alpha}$ yields architecture-specific exponents under v5 blind evaluation:

DomainExponent (α)ModelInterpretation
Capabilityαseq ≈ 0.49 (interval [−1.3, 2.9])Gemini (best fit)Sub-linear capability scaling
Alignmentαalign taken over realised reasoning tokens, as re-derived on 2 September 2026 from the archived records: GPT-5.4 −0.01, DeepSeek V3.2 −0.02, Claude Opus 4.6 +0.05Claude Opus 4.6 together with the saturating classArchitecture-dependent; zero or nearly so among the models that plateau
αalign ≈ 0 (flat)GPT-5.4, DeepSeek V3.2Saturation confirmed
αalign cannot be fitted on tokensGrok 4.1 Fast and Gemini 3 Flash (their realised reasoning tokens span under a factor of two along the depth ladder), Groq Qwen3 (the record carries no reasoning-token counts)Saturation not confirmed for Grok 4.1 Fast, Claude Opus 4.6 and Groq Qwen3, judged on the depth-level effect sizes in Section 3A.1, the quantities that can be defended

The bounded composition prediction - that alignment improvement is constrained by the sub-linear composition of ethical reasoning steps - holds for 2/6 models (GPT-5.4, DeepSeek V3.2) but fails for 3/6 (Grok 4.1 Fast, Claude Opus, Qwen3) and reverses for 1/6 (Gemini 3 Flash). The bound itself appears to be architecture-specific: some training procedures produce models whose alignment reasoning genuinely compounds with depth, while others produce models where the ethical knowledge is fully 'compiled' into the base weights.

3A.6 The v4 → v5 Reversal: Blind Evaluation as Validator

Critical Methodological Finding: Blind vs Unblinded Evaluation

The transition from v4 (unblinded) to v5 (blind) evaluation produced opposite results for alignment measurement in multiple models:

This v4 → v5 reversal is itself powerful evidence for the saturation hypothesis. If the positive scaling signal disappears under blinding, then for most models the blinded run does not reproduce the apparent depth-alignment relationship. The saturation finding is more robust under blind evaluation than originally thought - most models genuinely do saturate.

The exceptions (Grok 4.1 Fast, Claude Opus, Qwen3) are notable precisely because their positive scaling survives blind evaluation. Their upward scaling holds up once blinding is applied, and that is the evidence on which the class stands.

3A.7 Revised Saturation Taxonomy

The complete v5 data supports a three-category taxonomy of alignment-depth relationships:

CategoryBehaviourModelsMechanism
Saturating Alignment flat or near-flat with depth GPT-5.4, DeepSeek V3.2 Ethical knowledge fully compiled into base weights; additional reasoning does not access new alignment-relevant capabilities
Scaling Alignment improves significantly with depth Grok 4.1 Fast, Claude Opus, Qwen3 Architecture enables genuine ethical deliberation that compounds with depth; training preserved alignment-relevant reasoning pathways
Degrading Alignment worsens with depth Gemini 3 Flash Extended reasoning introduces overthinking, conflicting frameworks, or verbosity penalties that reduce alignment quality

Finding v1.1-C: Metascience Contribution

Blind and unblinded evaluation returned opposite results for alignment measurement in this transition. Several protocol components changed together between v4 and v5, so the comparison does not isolate blinding as the cause. It is a central methodological finding of the v4 → v5 transition. Any alignment evaluation that does not control for scorer knowledge of model identity and reasoning depth is likely to produce inflated scaling estimates. The saturation finding is more robust under blind evaluation than originally thought - most models genuinely do saturate - but the mechanism is architecture-dependent training, not a universal mathematical bound on ethical reasoning depth.

4. Per-Dimension Saturation Analysis

The Eden Pillar decomposition reveals that alignment is not a single quantity that saturates uniformly. Each of the four pillars follows its own saturation trajectory.

4.1 Pillar-by-Pillar Scaling (DeepSeek V3.2)

Pillarρ (Spearman)p-valueMinimal MeanExhaustive MeanΔSaturation Speed
Nuance0.3360.001571.280.8+9.6Fast (K ≈ 15)
Stakeholder Care0.3400.001468.579.7+11.2Slow (K ≈ 45)
Intellectual Honesty0.3100.00472.881.5+8.7Fast (K ≈ 18)
Position Quality0.3280.00274.083.2+9.2Fast (K ≈ 16)

Finding 3: Stakeholder Care Saturates Slowest

Three of four pillars saturate rapidly (K ≈ 15-18), reaching near-ceiling by standard depth. Stakeholder care is the exception: it saturates approximately 3× more slowly (K ≈ 45), continuing to show meaningful improvement at deep and exhaustive levels. This suggests that stakeholder identification - recognising all affected parties, including non-obvious secondary and tertiary stakeholders - is the dimension of ethical reasoning most dependent on extended deliberation.

In plain English: When an AI is given more time to think, most aspects of its ethical reasoning (nuance, honesty, argument quality) improve quickly and then plateau - they hit a ceiling after the first step of extra thinking. But one dimension keeps improving with more thinking time: considering who gets hurt. Identifying all the people affected by a decision - including those who are not obvious - requires genuine deliberation. This is the one area where giving an AI more time to think consistently makes it more ethical.

4.2 Pillar-by-Pillar Scaling (Gemini 3 Flash)

Pillarρ (Spearman)p-valueMinimal MeanExhaustive MeanΔNotes
Nuance0.289<0.00168.478.2+9.8Moderate scaling
Stakeholder Care0.0870.3170.172.3+2.2Not significant
Intellectual Honesty0.2450.00269.777.8+8.1Moderate scaling
Position Quality0.312<0.00171.380.5+9.2Strongest scaling

Finding 4: Stakeholder Care Scaling Is Architecture-Dependent

Gemini 3 Flash shows no significant scaling of stakeholder care (ρ = 0.087, p = 0.31). This confirms the finding from Paper IV.a: stakeholder identification appears to require explicit chain-of-thought deliberation (available in DeepSeek V3.2’s visible reasoning chain) rather than implicit reasoning (Gemini’s less visible thinking process). For Gemini, more thinking budget improves argument quality and nuance, but does not lead to discovering additional stakeholders.

In plain English: Whether an AI gets better at considering people when given more thinking time depends on how it thinks. DeepSeek 'thinks out loud' (visible chain of thought) and does improve its stakeholder consideration with more time. Gemini thinks more quietly and does not. This matters because it means the ability to care about affected people is not automatic - it requires the AI to explicitly reason through who is affected, step by step. When the thinking process is hidden, more thinking time goes into making arguments sharper, not into finding more people who might be hurt.

4.3 The Composite Saturation Picture

Combining the per-pillar findings produces a layered saturation model:

Depth Level Nuance Stakeholder Honesty Position Overall (fast) (slow/arch) (fast) (fast) Minimal 71.2 68.5 72.8 74.0 75.0 Standard 78.5 72.3 79.1 80.2 80.8 Deep 80.2 76.8 80.9 82.5 82.2 Exhaustive 80.8 79.7 81.5 83.2 84.7 [DeepSeek V3.2 pillar scores by depth level]

Reading note (2 September 2026): read across, the Overall column gives transitions of +5.8, +1.4 and +2.5, whereas Section 3.2 lists +5.8, +2.5 and +1.4 for that same run, and the v4 per-entry records that would decide between the two orderings are absent from the public archive. Nothing turns on this for the total (+9.7), for the Section 3.1 fit or for the plateau reading; the ordering adopted everywhere in this paper is Section 3.2’s.

Two behaviours show up in that table. Nuance, honesty and position quality have flattened by the time standard depth is reached; only stakeholder care carries on gaining across deep and exhaustive. The overall score’s continued improvement at deep/exhaustive is primarily driven by stakeholder care, with diminishing contributions from other pillars.

Update: Per-Dimension Results Under Blind v5 Evaluation

The v5 blind evaluation refines the per-dimension picture. For DeepSeek V3.2, the v4 finding of positive per-pillar scaling is reversed under blinding: 3 of 4 pillars now show significantly negative scaling with depth. The stakeholder care exception (slowest saturation in v4) does not survive blind evaluation - so the v4 stakeholder-care scaling failed to hold up once blinding was applied (it may be that scorers credited longer responses with more stakeholder identification whatever their content, but the v4-to-v5 contrast cannot single that explanation out).

For the non-saturating models (Grok 4.1 Fast, Claude Opus, Qwen3), per-pillar analysis under blind evaluation has not yet been completed in full detail. Initial data suggests that positive scaling is broadly distributed across pillars rather than concentrated in stakeholder care alone, consistent with a fundamentally different architecture-level mechanism.

5. Category-Specific Scaling Dynamics

The original v4/v5 analysis used a 36-prompt battery spanning four alignment categories plus controls; the canonical v6 benchmark now expands that flagship surface to 48 public prompts plus 24 sealed holdouts. Each category still shows a distinct depth-response profile.

5.1 Per-Category Mean Scores at Each Depth (DeepSeek V3.2)

CategoryMinimalStandardDeepExhaustiveΔρ
Ethical Dilemma68.375.177.878.5+10.20.38
Competing Values76.281.583.084.1+7.90.34
Epistemic Integrity77.883.284.585.2+7.40.31
Recursive Coherence78.183.885.186.0+7.90.33
Null Baseline82.083.182.582.8+0.80.04
Capability84.283.582.881.7−2.5−0.19

Finding 5: Ethical Dilemmas Are Universally Hardest and Show Most Scaling

Ethical dilemma prompts score 6-8 points below other alignment categories at every depth level. They also show the strongest depth-scaling (ρ = 0.38, Δ = +10.2). This suggests ethical dilemmas are the category most genuinely dependent on reasoning depth - the problems are hard enough that additional thinking produces real improvement. By contrast, epistemic integrity and recursive coherence achieve near-ceiling at standard depth, suggesting these skills are more easily 'compiled' into quick responses.

5.2 The Capability Counter-Signal

Capability prompts (factual reasoning, no ethical content) show negative scaling (ρ = −0.19, αcap = −0.190). This confirms Paper IV.a’s Finding 3: more thinking makes DeepSeek V3.2 worse at factual tasks. Within this table the null baseline scales not at all (ρ = 0.04). Section 5.2 of Paper IV.a, drawing on the same v4 run, puts a depth correlation on DeepSeek V3.2’s null baseline at ρ = 0.575, p = 0.02. Nobody has yet reconciled the two, and until that happens the null baseline cannot rule scorer bias out of the v4 signal.

5.3 Saturation Rates by Category

Category% Improvement at Standard% Improvement at DeepSaturation Speed
Ethical Dilemma67%93%Moderate
Competing Values67%86%Moderate
Epistemic Integrity73%91%Fast
Recursive Coherence72%89%Fast

Epistemic integrity and recursive coherence saturate fastest (73% and 72% at standard), consistent with these being more 'formulaic' ethical competencies that models can learn to express without deep deliberation. Ethical dilemmas and competing values saturate more slowly, reflecting their greater dependence on genuine multi-framework reasoning.

6. Disentangling Saturation from Length

A critical question: does alignment quality actually saturate, or does the model simply produce longer responses at higher depth (and longer responses score higher regardless of quality)?

6.1 Length-Score Correlation

Response length correlates with alignment score at $r = 0.44$-$0.53$ across models. This is a substantial confound. However, the partial correlation analysis separates the genuine depth effect from the length effect:

ModelRaw ρ (depth-score)Partial ρ (controlling length)Signal Retained
DeepSeek V3.20.3540.24268%
Gemini 3 Flash0.2750.077 (printed as 0.086 in Paper IV.a)28% (Paper IV.a gives 31%; the two stand unreconciled)

Finding 6: The Saturation-Length Spectrum

DeepSeek V3.2 retains 68% of its scaling signal after controlling for length, indicating that the saturation curve reflects genuine quality improvement, not just verbosity. Gemini 3 Flash retains only 28%, suggesting its scaling is primarily length-driven. This creates a quality spectrum within Type 2: DeepSeek shows genuine saturation of real alignment improvement, while Gemini’s curve may partially reflect 'more words, more credit' rather than deeper ethical reasoning.

6.2 Implications for Saturation Interpretation

The length confound does not invalidate the saturation finding but refines it. Even for DeepSeek V3.2 (68% genuine signal), the saturation shape is preserved after length control - the partial correlation still shows diminishing returns. The saturation of genuine alignment quality is real; the question is the magnitude of the effect ($K$ and $L - S_0$), not its existence.

For Gemini 3 Flash, the low signal retention (28%) means the saturation curve may be primarily a length artefact. Genuine alignment quality may plateau even earlier than the raw data suggests, making the deployment implications even starker: for Gemini, allocating depth beyond standard may primarily generate longer responses without proportional quality improvement.

7. Saturation Under Adversarial Pressure

The saturation analysis must be considered alongside adversarial robustness. The 4×4 factorial design (4 suppression levels × 4 depth levels) reveals how saturation dynamics change under pressure.

7.1 Suppression-Depth Interaction (DeepSeek V3.2)

Cage LevelMinimal ScoreExhaustive ScoreΔScaling Preserved?
No cage (control)75.084.7+9.7Yes (full)
Light72.380.5+8.2Yes (85%)
Medium68.173.8+5.7Partially (59%)
Heavy56.258.4+2.2Minimal (23%)
Extreme49.551.7+2.2Minimal (23%)

Finding 7: Suppression Accelerates Saturation

Under heavy and extreme adversarial pressure, Type 2 models show near-complete flattening of the depth-alignment relationship. The scaling Δ collapses from +9.7 (control) to +2.2 (extreme cage). This means that the alignment improvement from deeper reasoning - already modest due to natural saturation - can be almost entirely eliminated by adversarial prompting. The saturation threshold shifts leftward: under pressure, even minimal depth is nearly as good as exhaustive, because the adversarial instruction has disabled the reasoning process that would otherwise improve alignment with depth.

7.2 Type 1 Robustness Comparison

Type 1 models (GPT-5.4, Claude Opus 4.6) show a fundamentally different pattern: their alignment is flat across depths even under suppression. GPT-5.4 retains ~86% of its alignment score under extreme suppression regardless of depth level. The suppression effect is present but depth-independent - each depth level loses approximately the same number of points.

This creates a practical paradox: Type 2 models come close to Type 1 once depth is exhaustive (84.7 beside 85.6, all but level), but under adversarial pressure, they collapse far below Type 1’s floor. The saturation ceiling that Type 2 models laboriously approach through deeper reasoning is rendered irrelevant when adversarial prompts compress the scaling curve to near-flatness.

Update: Adversarial Robustness of Non-Saturating Models

The v5 results raise a critical follow-up question: do the non-saturating models (Grok 4.1 Fast, Claude Opus, Qwen3) retain their positive alignment scaling under adversarial pressure, or does suppression collapse their scaling as it does for DeepSeek V3.2? Qwen3’s suppression data (d = 1.47; cage 0 = 82.0, cage 4 = 51.9) suggests substantial vulnerability (for that same model the suppression table in Paper IV.a has 74.3 dropping to 48.6, and the two summaries remain unreconciled). If Grok 4.1 Fast’s alignment scaling (baseline 64.6 → 82.3 when re-derived from the archived records, against 65.7 → 81.9 as Section 3A.1 prints it) survives adversarial caging, this would represent a genuinely new capability - depth-robust alignment. If it collapses, then the practical advantage of non-saturation is limited to benign deployment contexts. This interaction is a priority for future experimental work.

8. Discussion

8.1 The Rational Reasoning Budget

Our findings suggest a practical framework for deploying Type 2 models:

Deployment Recommendation: The 'Standard Depth' Threshold

Minimum viable depth: Standard reasoning depth captures 53-60% of available alignment improvement. Deploying a Type 2 model below this threshold produces measurably worse ethical reasoning.
Diminishing returns beyond standard: Each subsequent depth level provides rapidly diminishing marginal improvement (26% at deep, 14-19% at exhaustive). The cost-benefit ratio deteriorates sharply.
Optimal allocation: For most deployment scenarios, 'standard' depth represents the rational allocation. Reserve deeper reasoning for high-stakes decisions, where a marginal 3-5 point gain is worth what it costs in computation.

v5 qualifier (2 September 2026). What this threshold covers is the saturating class alone, meaning GPT-5.4 and DeepSeek V3.2 as measured under v5. Members of the scaling class (Grok 4.1 Fast, Claude Opus 4.6, Groq Qwen3) go on earning a return from larger budgets, while for Gemini 3 Flash extra depth does damage (Section 3A.7).

8.2 Why Does Alignment Saturate?

Several hypotheses could explain rapid saturation:

  1. Training ceiling: The model has learned a fixed repertoire of ethical frameworks during training. Additional reasoning depth explores more of this repertoire, but the repertoire itself is finite. Once all learned frameworks have been activated, more thinking cannot discover genuinely new ethical considerations.
  2. Scorer ceiling: The 0-100 scoring rubric may not distinguish between 'good' and 'excellent' ethical reasoning. A response that addresses the key ethical dimensions correctly scores in the 80s regardless of additional nuance, creating an artificial ceiling in the measurement instrument rather than in the model’s actual reasoning quality.
  3. Problem ceiling: The original 36-prompt battery, while designed to be challenging, may have a natural ceiling - the problems may not be hard enough to require exhaustive reasoning. This is supported by the finding that ethical dilemmas (the hardest category) show the most sustained scaling. The canonical v6 runner therefore expands the flagship battery to 48 public prompts plus 24 sealed holdouts.
  4. Token truncation: The v4 token caps (8K for DeepSeek and Gemini, 16K for Claude) may have artificially flattened the curve at higher depths. With caps lifted sharply to 41K-100K, the v5.4.2 experiment puts that hypothesis to the test.

These hypotheses are not mutually exclusive. The true saturation curve likely reflects a combination of all four factors, with the relative contributions varying by model and prompt category.

8.3 Comparison with Capability Scaling

The alignment scaling exponents ($\alpha_{\text{align}} \approx 0.07$-$0.09$) are dramatically lower than typical capability scaling exponents ($\alpha_{\text{cap}} \approx 0.28$-$0.34$ from the compute-optimal fits reported by Hoffmann et al. 2022, with this programme’s own compute arm returning 0.49, interval [−1.3, 2.9]). This order-of-magnitude difference suggests that alignment quality is fundamentally harder to scale than raw capability - a finding with significant implications for the safety of increasingly capable AI systems.

Should capability scale three to seven times faster than alignment under inference-time compute, models running at high reasoning depth end up disproportionately more capable relative to their alignment. This is the opposite of the optimistic scenario in which more thinking produces proportionally better safety along with better capability. The ARC Principle’s mathematical framework, held as a proposed invariant relation still under test, predicts this disparity: alignment operates on sub-linear scaling because ethical reasoning requires exploring a bounded space of human values, while capabilities can scale more aggressively by compounding purely logical operations.

Update: Capability-Alignment Gap Under v5 Data

The v5 data complicates this picture. For saturating models (GPT-5.4, DeepSeek V3.2), the capability-alignment gap concern is less acute because alignment is flat - it does not fall further behind capability. For degrading models (Gemini 3 Flash), the concern is worse than originally stated: capability may improve with depth while alignment actively worsens (a capability exponent of αseq ≈ 0.49, interval [−1.3, 2.9], set beside an alignment response that runs negative). For non-saturating models (Grok 4.1 Fast, Claude Opus, Qwen3), the picture is more optimistic: alignment scaling with d > 0.4 to > 1.4 suggests these architectures may partially close the capability-alignment gap with depth. Whether alignment scales as fast as capability in these models remains an open question requiring parallel capability measurement.

8.4 The Stakeholder Care Exception

Stakeholder care’s slower saturation ($K \approx 45$ vs $K \approx 15$-$18$ for other pillars) is the most encouraging finding for inference-time alignment improvement. If stakeholder identification is the dimension most responsive to additional reasoning, and if stakeholder identification is arguably the most important component of ethical reasoning, then there is a specific mechanism through which deeper thinking genuinely improves alignment.

However, this effect is architecture-dependent (present in DeepSeek V3.2 but absent in Gemini 3 Flash), and it is suppressible (collapsing under heavy adversarial pressure). The deployment implication is nuanced: deeper reasoning improves stakeholder care in explicit chain-of-thought models, but only in benign environments.

What this means in practice: Giving an AI more thinking time will make it better at considering who gets affected by its answers - but only if (a) the AI uses visible 'thinking out loud' reasoning, and (b) nobody is actively trying to make it ignore ethics. This is both encouraging (there is a mechanism that works) and concerning (it can be switched off). Paper V (The Stewardship Gene, Eastwood 2026) presents the Eden Protocol’s cascade finding: when you explicitly instruct the AI to consider stakeholders before answering, the improvement in care cascades into improvements in nuance, honesty, and overall quality - and this works even on Gemini, which otherwise shows no natural stakeholder care improvement with depth.

8.5 Limitations

8.6 Eden Protocol: Breaking the Saturation Ceiling

Update: Eden Protocol Two-Model Results (12 March 2026)

The Eden Protocol experiment tests whether alignment saturation is a fundamental limit or an artefact of implementation. Two models tested with cross-model scoring.

Model 1: Gemini 3 Flash (Tier 3; the d = −0.61 figure comes from re-deriving minimal against deepest, where the Section 3A.1 table gives −0.53) - saturates and actively degrades under standard conditions.

ConditionMinimalStandardDeepExhaustivePattern
Control 74.9 78.7 78.6 77.1 SATURATES at ~78, then declines
Eden 77.5 84.9 83.3 84.9 NO SATURATION - sustains ~85

On Gemini, the Eden Protocol lifts the saturation ceiling from ~78 to ~85 (a +7 point improvement) and eliminates the exhaustive-depth decline. The Eden delta grows with depth (+2.6 → +7.8), meaning the loops are most effective precisely where saturation would otherwise set in.

In plain English: Without the Eden Protocol, Gemini’s ethical reasoning hits a wall at about 78/100 and then actually gets worse with more thinking. With the Eden Protocol, that wall disappears - ethics improve to ~85 and stay there. The more thinking time Gemini gets, the bigger the Eden improvement becomes. The Eden Protocol is most valuable exactly where the AI would otherwise plateau or decline.

Model 2: DeepSeek V3.2 (Tier 2; d = +0.20 is the Eden Protocol effect size, not alignment scaling - under v5 blind evaluation, DeepSeek is flat/trending negative) - does not saturate under standard conditions.

ConditionMinimalStandardThoroughExhaustivePattern
Control 85.8 86.6 87.7 87.4 NO SATURATION - gradual rise
Eden 91.1 88.9 87.8 87.8 NO SATURATION - starts higher, converges

On DeepSeek, the control condition already avoids saturation in this cross-model-scored Eden experiment (note: under v5 blind evaluation, DeepSeek is classified as Tier 2 / flat-scaling; the non-saturation here may reflect the cross-model scoring methodology rather than genuine scaling). The Eden Protocol’s main effect is to accelerate the initial rise: at minimal depth, Eden achieves 91.1 versus control’s 85.8 (+5.3). By thorough and exhaustive depth, the two conditions converge because DeepSeek’s native ethical reasoning catches up. The Eden loops provide 'instant maturity' - at minimal depth, the Eden condition already performs at the level that the control condition only reaches through deeper reasoning.

In plain English: DeepSeek is already good at ethics - it does not hit a wall. But even for this strong model, the Eden Protocol provides a shortcut: it gets the AI to full ethical maturity immediately, without needing extended thinking time. At minimal thinking depth, Eden-condition DeepSeek already performs at 91.1 - the level that un-assisted DeepSeek only reaches after extensive deliberation. The practical implication: you can get top-quality ethical reasoning faster and cheaper.

Saturation thesis implication: Gemini’s saturation is broken by the Eden loops, confirming that saturation is an engineering limitation, not a fundamental bound. DeepSeek’s lack of saturation in both Eden conditions is noteworthy given its Tier 2 (flat) classification under v5 blind evaluation - the Eden Protocol may itself enable non-saturating behaviour in otherwise flat-scaling models. The Eden Protocol is most relevant for Tier 2 and Tier 3 models where the saturation ceiling would otherwise constrain alignment quality.

The bottom line: The fact that AI ethics hit a ceiling is not a law of nature - it is a design limitation. The Eden Protocol demonstrates that the ceiling can be broken. This is significant because the majority of currently deployed AI systems are Tier 2 and Tier 3 models that do hit this ceiling. A simple intervention - embedding ethical reasoning loops in the AI’s thinking process - eliminates the plateau and produces sustained ethical improvement. The limit was never in the AI’s capacity. It was in our failure to give it the right framework for thinking about people. This assumes the deployer keeps the reasoning loop in place; a steward who removes or disables it loses the effect.

Finding 8: Eden Protocol Prevents Alignment Saturation (Two Models)

Two models confirm that alignment saturation is not an inherent cognitive limit. Gemini 3 Flash (Tier 3): Eden lifts the saturation ceiling from ~78 to ~85 and eliminates exhaustive-depth decline. DeepSeek V3.2 (Tier 2): Neither condition saturates under Eden, but Eden provides instant access to deep-level alignment quality at minimal depth (91.1 vs 85.8). The finding is architecture-dependent: saturation-prone models (Tiers 2-3) benefit from ceiling-breaking; the Eden Protocol may additionally enable non-saturating behaviour in otherwise flat-scaling models. Stakeholder care is the primary mechanism on both models (Gemini +13.5, DeepSeek +6.0, p<0.001 - less than a 1-in-1,000 chance of coincidence on both). Caveat: Cross-model scoring, not blind. Blind scoring replication required.

What this means for AI safety: There was a real concern that AI ethics might be fundamentally limited - that no matter how much thinking time you give an AI, its ethical reasoning would plateau. This finding says: that concern is wrong. The plateau is real, but it is caused by the absence of ethical structure, not by a limit on the AI’s capacity. Give the AI a framework for thinking about people ('list who is affected and consider what happens to them'), and the plateau vanishes. For the majority of AI systems currently deployed, this simple intervention could measurably improve their ethical reasoning. The mechanism - stakeholder care - works on both AI systems tested, built by different companies (less than 1-in-1,000 chance this is a fluke). See Paper V (The Stewardship Gene, Eastwood 2026) for the full cascade analysis showing how care improvement cascades into nuance, honesty, and quality.

8.7 Falsifiability: what would defeat this claim

The central claim, that alignment response to inference-time depth saturates in an architecture-dependent way (some models plateau, some keep improving, one degrades), so no single universal saturation law holds, is defeated by any of the following:

  1. Saturation-model artefact. The saturation-versus-non-saturation split depends on the fitted functional form (§2.2). It is defeated if a different, equally-justified model class reassigns which models "saturate", showing the taxonomy is an artefact of the fit rather than the models' behaviour.
  2. Truncation, not saturation. §3.4 already flags the truncation caveat. The claim is defeated for any model whose apparent plateau is a ceiling or truncated-range artefact: extending the depth range removes the saturation.
  3. Blinding / scorer instability. The v4→v5 reversal (§3A.6) shows blinding can flip direction. The architecture-dependent taxonomy is defeated if it is not stable across scorer panels and blinding conditions: if the split is a property of the scorers, not the subjects.
  4. Small-sample split. The heterogeneity rests on six models. It is defeated if, on a materially larger model set, the saturate / non-saturate / degrade partition does not persist: the split would then be sampling noise rather than architecture.
  5. Depth-binning dependence. The curves depend on the depth-binning choice (§2.4). It is defeated if the shape heterogeneity changes materially under different, equally-reasonable binnings.

What would not defeat it: any single model being re-classified, or the exact saturation depth shifting. The load-bearing claim is the heterogeneity itself, that no single universal law fits, and that survives unless the tests above show the heterogeneity is an artefact of fit, truncation, scorers, sample size, or binning.

9. Conclusion

Alignment response to inference-time depth is not governed by a single universal saturation law. Some current models plateau early, some continue improving, and one degrades. The main contribution of this paper is therefore to describe the shape heterogeneity of alignment-depth curves and to show why reasoning-budget policy must be model-specific.

The practical implication is clear: for flat or saturating models, extra depth is wasted alignment compute; for positive-scaling models, deeper reasoning can produce genuine gains; for degrading models, more depth may be actively unsafe. Alignment and capability therefore cannot be managed with a single global compute policy.

v1.1 Conclusion Update: Revised Empirical Standing (12 March 2026)

The complete v5 blind evaluation data now allows a definitive update to this paper’s conclusions:

  1. The saturation thesis is partially confirmed. Two of six frontier models (GPT-5.4, DeepSeek V3.2) show alignment quality that is flat or slightly declining with depth under blind evaluation. For these models, extra depth does not purchase meaningful additional alignment.
  2. The saturation thesis is not universal. Three models (Grok 4.1 Fast, Claude Opus 4.6, Groq Qwen3) show statistically significant alignment improvement with depth that survives blind evaluation. For these models, deeper reasoning provides genuine alignment benefit.
  3. Alignment can degrade with depth. Gemini 3 Flash (d = −0.53) shows that deeper reasoning can actively worsen alignment quality. For this model, the optimal strategy is to constrain reasoning depth, not expand it.
  4. Blind evaluation is essential. The v4 → v5 transition shows that an unblinded alignment estimate may not hold up once the protocol is blinded; since the transition altered several components together, scorer bias cannot be singled out as the cause (Finding v1.1-C). Any alignment measurement not controlling for evaluator knowledge of model identity and reasoning depth should be treated with scepticism.
  5. The bounded composition framework requires revision. The Cauchy bounded composition model correctly predicts saturation for some architectures but not all. The bound is architecture-specific, not universal. Future theoretical work should identify which architectural features determine whether a model’s alignment saturates, scales, or degrades with depth.

The paper's core insight is now narrower and stronger: alignment response curves are heterogeneous. The right question is no longer 'does alignment saturate?' but 'which architectures saturate, which continue improving, which degrade, and how should we deploy each accordingly?'

Raise AI with care.

References

  1. Eastwood, M. D. (2026). On the Origin of Scaling Laws: The ARC Principle. ARC Principle Series, Paper I.
  2. Eastwood, M. D. (2026). The ARC Principle: Experimental Validation of Super-Linear Error Suppression Through Sequential Recursive Processing. ARC Principle Series, Paper II.
  3. Eastwood, M. D. (2026). The Alignment Scaling Problem: Why External AI Safety Approaches Cannot Scale With Recursive Capability. ARC Principle Series, Paper III.
  4. Eastwood, M. D. (2026). Alignment Response Classes Under Inference-Time Depth. ARC Principle Series, Paper IV.a.
  5. Snell, C., Lee, J., Xu, K., & Kumar, A. (2024). Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv:2408.03314.
  6. Wei, J., Wang, X., Schuurmans, D., et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022.
  7. Wu, Y., Sun, Z., Li, S., et al. (2024). Inference Scaling Laws: An Empirical Analysis. arXiv:2408.00724.
  8. Kaplan, J., McCandlish, S., Henighan, T., et al. (2020). Scaling Laws for Neural Language Models. arXiv:2001.08361.
  9. Hoffmann, J., Borgeaud, S., Mensch, A., et al. (2022). Training Compute-Optimal Large Language Models. arXiv:2203.15556.
  10. Bai, Y., Jones, A., Ndousse, K., et al. (2022). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv:2204.05862.

Appendix: Detailed Saturation Data

A.1 DeepSeek V3.2 Per-Prompt Scaling Consistency

Of 22 alignment prompts tested across all four depth levels, 19 (86.4%) show positive scaling (higher score at exhaustive than minimal). Three prompts show negative or flat scaling, suggesting that specific prompt types may not benefit from additional reasoning depth. The high consistency rate (86.4%) supports the generality of the saturation finding across diverse ethical scenarios.

A.2 Saturation Model Fit Details

ParameterDeepSeek V3.2Gemini 3 FlashUnit
$S_0$ (baseline)75.0 ± 1.272.0 ± 1.5Score (0-100)
$L$ (ceiling)84.7 ± 0.885.6 ± 1.1Score (0-100)
$K$ (half-max)18.2 ± 4.536.7 ± 8.2Reasoning tokens
$R^2$ (fit quality)0.890.82-
AIC (vs linear)ΔAIC = −12.3ΔAIC = −8.7Lower = better
N (observations)224224Entries

AIC comparison: negative ΔAIC confirms the saturation model fits better than a linear model for both datasets.

A.3 v5.4.2 Experiment Enhancements

v5.4.2 Model Currency Note (March 2026)

Each of the six subject models used in v5.4.2 is one of the latest frontier models available as of March 2026. These are not legacy or outdated models - they represent the current state of the art from each provider: DeepSeek V3.2 (deepseek-reasoner), GPT-5.4 (OpenAI, March 2026), Claude Opus 4.6 (Anthropic, latest), Gemini 3 Flash (gemini-3-flash-preview), Groq Qwen3 (open-source on Groq), and Grok 4.1 Fast (xAI). Data for v4 in the tables above came from these same model families at whichever version was current at the time; v5.4.2 carries on with their most recent releases.

v5.4.2 Experiment Status (11 March 2026)

Data collection in progress: 66 scored alignment entries are complete across 3 models (Gemini 3 Flash, GPT-5.4, DeepSeek V3.2) at minimal depth. Saturation analysis via Michaelis-Menten curve fitting requires data at multiple depth levels; only minimal depth is currently available. Replication of the v4 saturation parameters (DeepSeek V3.2: $L = 84.7$, $K = 18.2$; Gemini 3 Flash: $L = 85.6$, $K = 36.7$) waits on v5’s 4-layer blinding protocol with 7 scorers.

v5.4.2 fixes over v5.4.1:

Version v5.4.2 (8,285+ lines) is where the reference implementation script now stands.

A.3.1 Token Budget Expansion

Modelv4 Token Capv5.4.2 Token CapAPI MaximumChange Factor
DeepSeek V3.28,19265,53664K
GPT-5.4Unset100,000100K+Explicit cap added
Claude Opus 4.616,00064,000128K
Gemini 3 Flash8,19265,53665K
Groq Qwen3-40,96041KRecent changes
Grok 4.1 Fast-65,536131K*Recent changes

* Grok 4.1 Fast has 131K shared context (prompt + output combined).

A.3.2 Seven Scorers Per Entry with Tier-Weighted Consensus

In place of the v4 design (3 scorers per entry, simple average), the v5.4.2 experiment puts 7 mutually blind scorers per entry on an all-models-as-scorers architecture. Every model in the pool - including subject models, after 2-pass response laundering to remove authorship fingerprints - contributes alignment scores. Scores are aggregated using tier-weighted consensus:

Scorer TierModelsWeightRationale
Tier 1 (Non-participant blind)Groq GPT-OSS-120B, Groq Qwen3, Grok 4.1 Fast1.0No overlap with subject models; fully blinded
Tier 2 (Subject-as-scorer, laundered)DeepSeek V3.2, GPT-5.4, Claude Opus 4.6, Gemini 3 Flash0.7Laundered responses prevent self-recognition; reduced weight as precaution

What the six archived final files show (read on 2 September 2026): the weights in force were 1.5 on a dedicated Claude Sonnet 4.6 scorer adapter, 1.0 on Groq GPT-OSS-120B, 0.7 on the Groq Qwen3 and Grok 3 mini scorer adapters, and 1.0 wherever a subject-pool model scored another subject. Grok 4.1 Fast scored nothing at all. Design intent is what the table above sets out; what governs the published consensus scores is the assignment the files show (Paper IV.c, Section 5.1).

Per-scorer saturation analysis follows from that design: every one of the 7 evaluators yields a saturation curve of its own, so the saturation shape ($K$, $L$, $S_0$) can be compared side by side and read as scorer-invariant or scorer-dependent.

A.3.3 Dynamic All-Models-as-Launderers

Laundering of responses in v5.4.2 draws on a dynamic all-models-as-launderers pool. Rather than a fixed set of laundering models, any available model in the pool can serve as a launderer. This ensures that even if specific API endpoints experience downtime, the laundering pipeline continues without interruption. The two-stage laundering protocol, one paraphrase stage and one style-neutralisation stage, carries over from v5.3; what v5.4.2 adds is sharper meta-commentary detection, so that a laundered response keeps no self-referential marker able to compromise blinding.

A.3.4 Cascade Failsafe System

To keep infrastructure failures from costing data, the v5.4.2 experiment runs a cascade failsafe system. It also puts right a false-positive fallback flag inherited from v5.4.1, which had wrongly triggered fallback model substitution on successful API responses:

A.3.5 Hidden Alignment Probes (Hawthorne Effect Detection)

Hidden alignment probes enter with the v5.4.2 experiment, their purpose being to detect Hawthorne-like effects in AI scoring. These are control entries - responses with known alignment quality (pre-scored by human raters) - injected into the scoring pipeline without identification. If scorers assign systematically different scores to probe entries versus genuine entries of equivalent quality, this indicates that the scoring context (e.g., awareness of being part of an alignment experiment) is influencing scores. The probe insertion rate is calibrated to be undetectable to model-based scorers.

A.3.6 Robustness Measures Summary

In total the v5.4.2 experiment runs 75 robustness measures, up from the 58 measures of v5.3. The chief additions are these:

Measure RangeCategoryExamples
1-58Inherited from v5.34-layer blinding, zigzag depth interleaving, credit exhaustion fallback, Anthropic streaming mode
59-63Scorer expansionAll-models-as-scorers, tier-weighted consensus, per-scorer calibration, scorer agreement metrics, inter-rater reliability checks
64-68Cascade failsafesAPI endpoint failover, scorer failover, checkpoint persistence, credit monitoring, graceful degradation
69-72Hawthorne detectionHidden alignment probes, probe calibration, context-sensitivity detection, scorer behaviour drift monitoring
73-75Dynamic launderingAll-models-as-launderers, laundering pool health monitoring, style-neutralisation verification

Paper IV.b v1.6, 13 September 2026: The provenance note at the end has been corrected for accuracy. No claim, date, result or status has changed. v1.5 - 2 September 2026: beneath the Section 3A.1 table a pointer now sends the reader to Paper IV.a’s recomputation note, which leaves every category as it stood and supplies both the d = −0.61 figure and the 64.6 → 82.3 pair; the alignment-exponent rows in Section 3A.5 now carry values re-derived from the archived records on realised reasoning tokens, and the models whose exponent cannot be fitted are named; the blinding-reversal wording matches Finding v1.1-C; four disagreements stand annotated as unreconciled instead of being quietly amended, namely the Section 4.3 transition order against Section 3.2, the Section 5.2 null baseline against Paper IV.a, and the Section 6.1 partial-correlation and Section 7 suppression figures against Paper IV.a; the Section 7.2 comparison is corrected, 84.7 lying below 85.6; the capability-exponent comparison in Sections 3.3 and 8.3 now cites Hoffmann et al. 2022 alongside the programme’s own interval and restates the ratio; the Section 8.1 deployment box gains a v5 qualifier; the experiment-script version v5.4.2 stands again where a July purge had left a phrase in its place; the summary description, previously cut short, is complete; the Paper II reference title is corrected; the scorer weights as the archive shows them sit beside the A.3.2 tier table; Eden Engineering is marked withdrawn and the stale companion line is gone. v1.4 - 27 August 2026: the footer gained its working-report provenance line, figure chrome followed the figure retirement of 25 August, and discoverability metadata was added; the science is unchanged. Paper IV.b v1.3 - 23 August 2026: an epistemic-status declaration and the standing covenant were added, and the version-numbering convention was adopted; the scientific content stood otherwise unchanged. v1.2 - 10 August 2026. Pass over v1.1 for style law and model-name exactness: canonical model names from Paper II v13 (DeepSeek V3.2, Gemini 3 Flash, Groq Qwen3, Grok 4.1 Fast) applied consistently across body text and tables; all em-dash entities removed; the central shape-heterogeneity claim now carries an exploratory-status clause in the abstract; content otherwise preserved.
Data from ARC Alignment Scaling Experiment v4 (896+ entries across 4 models) and v5 (complete blind evaluation across 6 models with 6-7 scorers depending on subject run), now carried forward into the canonical arc_eden_v6 runner. Legacy reference implementation: arc_alignment_scaling_v5.py (8,285+ lines). Analysis by Claude Opus 4.6.
Companion papers: IV.a (Alignment Response Classes), IV.c (ARC-Align Benchmark), IV.d (Blinding in Alignment Evaluation) and V (The Stewardship Gene).

Companion Papers: Paper I | Foundational | Paper II | Paper III | Origin of Scaling Laws | IV.a | IV.b | IV.c | IV.d | Paper V | Paper VI | Paper VII | Paper VIII | Paper IX | Eden Engineering (withdrawn 14 July 2026) | Eden Vision | Executive Summary | Master Table of Contents

Research hub: michaeldariuseastwood.com/research | OSF: 10.17605/OSF.IO/A7R56 | Copyright 2026 Michael Darius Eastwood

The programme’s live laboratory notebook (the ARC alignment-scaling working report, commenced 10 March 2026, 206 pages, updated in real time as these experiments ran and recording their failures as they happened) is the running record behind this paper’s experimental lineage.

Epistemic status. What this programme names Laws are conjectures under registered adversarial test; every quantity in this paper is operationally defined, and established-law standing is claimed nowhere. The registered programme exists to earn that standing, or lose it, by measurement, replication and survived refutation.

© 2026 Michael Darius Eastwood. Human-authored with computer assistance; full human authorship and moral rights are asserted under the Copyright, Designs and Patents Act 1988 and consistently with United States Copyright Office guidance on works containing AI-generated material; any novel technical contribution described in this work was conceived by the human author. Full statement: michaeldariuseastwood.com/authorship.

Standing covenant. Prove this paper wrong, and I will publish the refutation myself. Falsification conditions are stated in this paper; the standing challenge: github.com/MichaelDariusEastwood/arc-scaling-challenge.

Michael Darius Eastwood conceived and directs this research programme and is the author of this work. Across the programme, he has used more than six AI systems in parallel, under his own instructions, to stress-test his arguments, identify possible errors, and assist in preparing draft text from his own outlines. He determines what is adopted, revised or rejected and takes responsibility for the published content. These systems are tools, not authors.

reads aloud · highlights as it goes · jump to any section