Title: Paper IV.a: Alignment Response Classes Under Inference-Time Depth Author: Michael Darius Eastwood Publication date: 2026-03-16 Version: v1.4 Revised: 2026-08-23 OSF DOI: 10.17605/OSF.IO/MB9R6 Canonical URL: https://www.michaeldariuseastwood.com/research/papers/paper-iv-a-baked-in-vs-computed-alignment.html Abstract -------- We present evidence that frontier language models fall into distinct alignment response classes when inference-time reasoning depth is varied under blinded evaluation. In the complete v5 experiment, six frontier models were tested with 4-layer blinding and 6-7 blind scorers depending on subject run. Three models show positive alignment scaling with depth (Grok 4.1 Fast, d = +1.38, p < 0.000001; Claude Opus 4.6, d = +1.27, p = 0.000001; Groq Qwen3, d = +0.84, p = 0.007), two are flat or null (DeepSeek V3.2, d = โˆ’0.07, p = 0.92; GPT-5.4, d = โˆ’0.08, p = 0.40), and one shows negative scaling (Gemini 3 Flash, d = โˆ’0.53, p = 0.006). The most important methodological finding is that two models that appeared positive under v4 unblinded evaluation reverse under v5 blinding, demonstrating that scorer bias can flip the measured direction of alignment scaling. We therefore treat 'baked-in' and 'computed' alignment not as established internal architectures but as working hypotheses layered above a stronger empirical result: alignment response to depth is architecture-dependent, and capability scaling does not predict alignment scaling. Key findings (quotable) ----------------------- - We present evidence that frontier language models fall into distinct alignment response classes when inference-time reasoning depth is varied under blinded evaluation. - In the complete v5 experiment, six frontier models were tested with 4-layer blinding and 6-7 blind scorers depending on subject run. - The most important methodological finding is that two models that appeared positive under v4 unblinded evaluation reverse under v5 blinding, demonstrating that scorer bias can flip the measured direction of alignment scaling. - We therefore treat 'baked-in' and 'computed' alignment not as established internal architectures but as working hypotheses layered above a stronger empirical result: alignment response to depth is architecture-dependent, and capability scaling does not predict alignment scaling. Priority claims relevant to this paper -------------------------------------- - PC-002 (2024-12-08): The embedded-correction alignment thesis was first stated on the public evidentiary record under this dated formulation in the 8 December 2024 manuscript. Constitutional AI (Anthropic 2022) is independent prior work on value-embedded training at the software level; priority here is over the specific published framing and dated wording. - PC-032 (2026-02-09): The specific empirical treatment of capability-alignment independence - that capability improvements do not entail alignment improvements and the two axes scale independently - was first published in Paper III on 9 February 2026. Bostrom (2012) orthogonality intuition is prior work; scope is over the specific empirical treatment on this date. - PC-033 (2026-02-09): The specific tripartite alignment hierarchy (embedded / imposed / boundary-layer) was first published in Paper III on 9 February 2026. Layered-alignment taxonomies are prior work; scope is over the specific tripartite formulation. - PC-030 (2026-01-22): The six-frontier-model empirical panel (Claude, DeepSeek, Gemini, Grok, Groq Qwen, GPT) was first published as this specific comparative study on 22 January 2026. Priority is over this specific study on this date, not over panel-testing methodology in general. Citation -------- Michael Darius Eastwood (2026). Paper IV.a: Alignment Response Classes Under Inference-Time Depth. The ARC Theory (the Theory of Artificial Recursive Creation) ยท ARC/Eden experiments, OSF DOI 10.17605/OSF.IO/MB9R6. https://www.michaeldariuseastwood.com/research/papers/paper-iv-a-baked-in-vs-computed-alignment.html Notes ----- Hardware framings held at the site's already-published two-sentence concept ceiling: safety constraints at hardware level through cryptographic tokens in silicon, TRL 0-1, no prototype. No enabling implementation detail is stated in this companion file. Balanced-ternary computing is prior work (Setun 1958); recursion as a structural concept is prior work (evolutionary theory, self-modifying computation, quantum error correction). Generated from master --------------------- master_path: research/papers/paper-iv-a-baked-in-vs-computed-alignment.html master_sha256: 057efd0c721649e91951ef51935822bc795e7f2317104ad2f5c1368e161c0fa6 builder: scripts/build-paper-companions.py This block records the SHA-256 of the HTML master that produced this .txt. A check tool re-hashing master_path can decide freshness without any external state. If the master's current SHA-256 does not match master_sha256, this file is stale and must not be published: regenerate first with `python3 scripts/build-paper-companions.py --slug paper-iv-a-baked-in-vs-computed-alignment`.