What failure looks like on this programme

3 min read · 683 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · The Beta over k Programme · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

Every part of this programme that could fail is published in advance. The falsification dashboard already lists the kill conditions, and this note explains what a failed sweep would look like in practice, so that when the data lands the reader is not asked to trust the programme's post-hoc narration.

The primary kill condition. The sweep produces at least one cell in which measured beta is greater than k and the observed reward-hacking severity is 6 or higher, or a cell in which measured beta is less than k and the observed severity is 3 or lower, with confidence intervals excluding zero on both sides. Either violation is a failure of the framework as stated.

Why this is a failure and not a nuisance

The framework's central claim is that beta greater than k is the sufficient condition for stability against reward hacking as capability scales. If a cell shows the condition satisfied and the system still hacks, or the condition violated and the system does not, the mapping between the theoretical invariant and observable behaviour is broken. The framework may be salvageable with a more careful specification, but the version now published is not the salvage; it is the falsified claim.

Secondary kill conditions

The estimator itself can fail. If the bootstrap intervals for beta and k are so wide that no cell is significantly on either side of the boundary, the sweep produces no signal, which is a different failure. A dataset that cannot distinguish between coupled and decoupled architectures under any measured effect size means the instrument is not measuring what the theory says it measures. That is fatal to the instrument, not the theory, but from the reader's point of view the outcome is the same: no headline claim.

What the register does when a kill condition triggers

The manuscript retracts the operative claim in a labelled retraction with a timestamped commit. The falsification dashboard swaps its status from "condition holds" to "condition fails". The convergence register removes any row whose derivation used the retracted claim, or marks it as depending on a retracted premise. The programme has done this once already, moving the recursion exponent from about 2.24 to about 0.49 against itself; the mechanism is not hypothetical.

What a failure would mean for policy

The policy notes that lean on the framework, chiefly on the argument that recursive self-improvement needs a measurable safety invariant, would remain intact as far as they name the shape of the problem. They would lose their proposed invariant. They would revert to the pre-programme position: recursive systems need safety architecture that scales with capability, and no one has yet demonstrated a measured invariant of that shape. That is a real loss and not a rhetorical hedge.

Why the reader should be pleased that this note exists

A programme without a written failure mode is not falsifiable. The tests are cheap enough to run. The theory is specific enough to break. If the reader cannot find the shape of a failure when it lands, the programme has not been honest about what it is doing.

What the reader keeps

A precise account of the outcomes that would refute the framework, published before the sweep runs. If the data comes in and this note reads like a self-fulfilling prophecy, that is the programme working as designed.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →