Paper II: what the retraction actually says =========================================== Paper II tried to pin down the ARC exponent experimentally. The first number was wrong. The revised number is smaller, harder to make headlines with, and more honest. Canonical path: /research/papers/paper-ii-experimental-validation.html Author: Michael Darius Eastwood Research programme: https://doi.org/10.17605/OSF.IO/6C5XB Paper II: what the retraction actually says Michael Darius Eastwood · Independent AI alignment researcher Published 3 July 2026 Michael Darius Eastwood · Experimental validation · 3 July 2026 Michael Darius Eastwood, independent researcher, London: author of the ARC/Eden research programme; the embedded-correction alignment thesis is recorded in a source record dated 8 December 2024 (sent-side SHA-256 f0d1f38f). Paper II was the programme’s first attempt to measure the ARC exponent in a controlled experiment rather than infer it from published tables. The first pass gave an eye-catching number. A second pass across more models made that number vanish. Both results are on the record. The paper is now interesting mostly for what it shows about how easy it is to overfit a scaling law to a single system. Paper II (blinded six-model re-analysis) reports controlled measurements of the ARC exponent across six frontier models, revising the original single-model estimate downwards after cross-architecture replication under blinding. The earlier unblinded single-model fit of alpha approximately 2.24 was retracted, corrected to approximately 0.49 under blinding. OSF DOI 10.17605/OSF.IO/6C5XB. The question it asks If Paper I is right that sequential recursion produces a super-linear exponent while parallel recursion does not, then a properly controlled experiment should measure both exponents on the same set of problems, on the same models, at matched compute. What is alpha for sequential thinking, once you actually run the tests? Is it larger than one? Is it consistent across model families? And what does the answer say about scaling laws for capability more generally? What it found Two things replicated cleanly across the study. Parallel recursion (majority voting over independent samples) yielded alpha close to zero for every model tested, matching Paper I. Sequential recursion beat parallel recursion for every model where both were measurable. --- Machine-readable companion. Cite: Eastwood, M. D. (2026). "Paper II: what the retraction actually says". /research/papers/paper-ii-experimental-validation.html