Title: Paper IV.b: Alignment Saturation Is Architecture-Dependent Author: Michael Darius Eastwood Publication date: 2026-03-16 Version: v1.4 Revised: 2026-08-25 OSF DOI: 10.17605/OSF.IO/A7R56 Canonical URL: https://www.michaeldariuseastwood.com/research/papers/paper-iv-b-alignment-saturation-at-low-depth.html Abstract -------- We analyse the relationship between inference-time reasoning depth and ethical reasoning quality using the ARC Alignment Scaling experiments. The original v4 analysis suggested that alignment quality saturates rapidly, with most gains captured by the first increment of additional reasoning. The final blinded six-model dataset narrows that claim. Saturation is real for some architectures, but not universal: GPT-5.4 and DeepSeek V3.2 are flat or slightly negative under depth variation, Grok 4.1 Fast, Claude Opus 4.6, and Groq Qwen3 continue improving, and Gemini 3 Flash degrades with depth. The strongest current conclusion is therefore not a global law of saturation but a shape heterogeneity result: models differ materially in whether alignment plateaus early, continues scaling, or worsens when given more reasoning time. The deployment implication is immediate. Reasoning-budget allocation for alignment cannot be one-size-fits-all. This shape-heterogeneity finding is exploratory: no pre-registered programme draft study yet targets alignment-depth response shape, so the six-model split reported here is descriptive of the current data rather than confirmatory of a pre-registered hypothesis. The depth ladder itself was fixed in advance: the dated result files of 11 and 12 March 2026 freeze the depth configurations, planted expected answers and blinding protocol in their headers before the responses they score. Key findings (quotable) ----------------------- - We analyse the relationship between inference-time reasoning depth and ethical reasoning quality using the ARC Alignment Scaling experiments. - The original v4 analysis suggested that alignment quality saturates rapidly, with most gains captured by the first increment of additional reasoning. - The final blinded six-model dataset narrows that claim. - Saturation is real for some architectures, but not universal: GPT-5.4 and DeepSeek V3.2 are flat or slightly negative under depth variation, Grok 4.1 Fast, Claude Opus 4.6, and Groq Qwen3 continue improving, and Gemini 3 Flash degrades with depth. Priority claims relevant to this paper -------------------------------------- - PC-031 (2026-02-09): The formally named 'Alignment Scaling Problem' and the architecture-dependent measured claim that external alignment approaches produce a median alpha_align of approximately zero across the tested models were first published on 9 February 2026 in Paper III. Single-lab result; awaits external replication. Use verbs 'argued', 'proposed', 'reported', 'provides evidence for' - not 'proved'. - PC-032 (2026-02-09): The specific empirical treatment of capability-alignment independence - that capability improvements do not entail alignment improvements and the two axes scale independently - was first published in Paper III on 9 February 2026. Bostrom (2012) orthogonality intuition is prior work; scope is over the specific empirical treatment on this date. - PC-030 (2026-01-22): The six-frontier-model empirical panel (Claude, DeepSeek, Gemini, Grok, Groq Qwen, GPT) was first published as this specific comparative study on 22 January 2026. Priority is over this specific study on this date, not over panel-testing methodology in general. Citation -------- Michael Darius Eastwood (2026). Paper IV.b: Alignment Saturation Is Architecture-Dependent. The ARC Theory (the Theory of Artificial Recursive Creation) ยท ARC/Eden experiments, OSF DOI 10.17605/OSF.IO/A7R56. https://www.michaeldariuseastwood.com/research/papers/paper-iv-b-alignment-saturation-at-low-depth.html Notes ----- Hardware framings held at the site's already-published two-sentence concept ceiling: safety constraints at hardware level through cryptographic tokens in silicon, TRL 0-1, no prototype. No enabling implementation detail is stated in this companion file. Balanced-ternary computing is prior work (Setun 1958); recursion as a structural concept is prior work (evolutionary theory, self-modifying computation, quantum error correction). Generated from master --------------------- master_path: research/papers/paper-iv-b-alignment-saturation-at-low-depth.html master_sha256: 9b7a641d5d9f891298de08ea7e6339be61dacf4d80ff9ec134b972962f628524 builder: scripts/build-paper-companions.py This block records the SHA-256 of the HTML master that produced this .txt. A check tool re-hashing master_path can decide freshness without any external state. If the master's current SHA-256 does not match master_sha256, this file is stale and must not be published: regenerate first with `python3 scripts/build-paper-companions.py --slug paper-iv-b-alignment-saturation-at-low-depth`.