Paper IV-b: alignment does not saturate the same way for every model ==================================================================== Paper IV-b of the ARC/Eden programme reports that alignment saturation with reasoning depth is architecture-dependent, not universal. Canonical path: /research/papers/paper-iv-b-alignment-saturation-at-low-depth.html Author: Michael Darius Eastwood Research programme: https://doi.org/10.17605/OSF.IO/6C5XB Paper IV-b: alignment does not saturate the same way for every model Michael Darius Eastwood · Independent AI alignment researcher Published 3 July 2026 Michael Darius Eastwood · Alignment scaling · 3 July 2026 Michael Darius Eastwood, independent researcher, London: author of the ARC/Eden research programme; the embedded-correction alignment thesis is recorded in a source record dated 8 December 2024 (sent-side SHA-256 f0d1f38f). The obvious next question after Paper IV-a is not just whether an AI system gets more or less aligned with extra thinking, but what shape that curve takes. Does the benefit plateau early, so that spending more compute is wasted? Does it keep climbing? Does it curve downward? Paper IV-b reports that these shapes are not universal. They are architecture-dependent. That has direct consequences for how a deployed system should budget its reasoning tokens. Paper IV-b · Reports that alignment saturation with reasoning depth is architecture-dependent across six frontier models, with saturating, continuing-improvement, and degrading shapes all present in the same dataset. OSF DOI 10.17605/OSF.IO/6C5XB. The question it asks If alignment quality is not a fixed property of a trained model but a function of how much reasoning depth you allow, what does that function look like? Does every architecture follow a Michaelis-Menten style saturation curve, so that the first bit of extra thinking captures most of the available improvement? Does the response continue to climb with a slow logarithmic drift? Or do some architectures produce fundamentally different shapes, saturating early, scaling steadily, or curving downward as depth grows? What it found In the earlier earlier unblinded dataset (four models, unblinded), a saturation curve fitted both Type 2 models cleanly, with half-maximum constants around 18 and 37 reasoning tokens: most of the available improvement showed up in the first increment of depth. That pattern was the paper’s original headline. Under the blinded evaluation across six models, the picture broadened. --- Machine-readable companion. Cite: Eastwood, M. D. (2026). "Paper IV-b: alignment does not saturate the same way for every model". /research/papers/paper-iv-b-alignment-saturation-at-low-depth.html