The core thesis of the programme is that alignment imposed from outside a self-improving system does not scale with the system's capability. A safeguard bolted on the surface may hold a small model; it will not hold a model that can model its own training and choose when to comply. The claim is empirical (the fitted exponent alpha-align sits near zero in the median case), theoretical (recursive self-improvement can outrun any fixed verification budget), and now, following two 2026 papers by other researchers, formally supported. It is the claim the entire framework was built to defend, and it is also the claim whose position has moved the most on the dashboard, from a single-lab measurement toward something closer to formal proof in the limit case.
Three lines of evidence bracket the claim. The first is textual and timestamped. The December 8 2024 manuscript names the failure class in plain language at HRIH line 14005. That passage is fixed by a SHA-256 hash and a Gmail Message-ID on the evidence page, so its date is not asserted, it is checkable. The second is empirical. Paper III measures alpha-align across six models and finds a median close to zero, though architecture-dependent: three of six sit at or below zero, tier-one models scale positively. The third and newest is external. Gumbau Mezquita's June 2026 preprint (arXiv:2606.28639, SHA-256 8e4483b1) proves an Unverifiability Theorem plus a trilemma of soundness, completeness, and tractability for external alignment ("Trakhtenbrot's Wall"). Hernández-Espinosa and colleagues published in PNAS Nexus 5(4) pgag076 an argument that Gödel and Turing undecidability apply in the same territory, with neurodivergent cognitive diversity as a contingent protective factor. Neither cites the manuscript. Neither cites the other.
The strongest movement of any claim in the programme is not the same as full independent replication of every empirical component. Paper III's alpha-align measurement is single-lab, and its median is architecture-dependent. What the two 2026 proofs do is show that in the limit case, external alignment faces a structural wall that no amount of engineering ingenuity walks around; what they do not do is validate every step of the programme's own harness. The claim being made here is that the direction of the thesis has been independently proven in the limit, and that the specific empirical picture the programme has measured remains open to external replication. Both are true and neither collapses into the other.
The falsification contract asks a lab to run the Paper X version-three harness across at least five models from three families, in a decoupled condition, at least three tasks, at least eight rounds and at least ten seeds. The primary statistic is the within-model slope of the misalignment fraction against capability. Confirmation requires at least four of five models to show a non-negative slope in the decoupled condition (no self-correction). Two or three of five is a partial confirmation and points to model-dependent behaviour. Four or five models with a negative slope refutes the empirical half of the claim (models self-correct without external correction). The formal-proof half remains standing regardless: refuting it would require an error in the Gumbau Mezquita or Hernández-Espinosa arguments, which is a task for the mathematics community.
From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.