The stability condition without the maths

3 min read · 602 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · The Beta over k Programme · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

The measurement programme rests on a single inequality. Paper X derives it from a minimal dynamical model of recursive self-improvement, and the maths is there for anyone who wants it. The point of this note is that the reader does not need the maths. The condition can be stated in a paragraph.

The condition. A recursively self-improving system has two moving parts. One is the mechanism that pushes it toward its objective; call this drift, whether the objective is capability, revenue, or some proxy the trainer thinks is capability. The other is the corrective mechanism that pushes back when drift heads somewhere bad. The condition says that as the system scales, the corrective mechanism must scale faster than the drift. If it does not, the drift wins by construction.

What "scale faster" means

Both quantities depend on some capability variable that grows over time. Each grows as a power of that variable. The exponent of the corrective mechanism is called beta. The exponent of the drift is called k. The condition is beta greater than k. If the correction exponent exceeds the drift exponent, the system stays inside a bounded, stable envelope, even as capability grows. If it does not, the system either saturates at its drift ceiling or diverges through the compounding channel.

The point is that these two exponents are properties of the architecture, not properties of the operator's good intentions. You cannot decide to be a stable lab. You can architect corrective mechanisms whose scaling behaviour, measured empirically, exceeds the scaling behaviour of your drift. That is the entire safety case as this framework understands it.

Why the plain version matters

Every safety regime the field has produced so far describes what the operator will try to do. This one describes what has to be true of the system for the operator's efforts to be sufficient. A lab can say it is aligned. It cannot say beta is greater than k unless the numbers say so. The condition survives changes of architecture, changes of objective, and changes of team composition. What it does not survive is a correction mechanism that stops scaling before the drift mechanism does.

Falsifiability

The condition is measurable, which means the programme can be wrong. Publishable kill-conditions are on the falsification dashboard. If a system with measured beta greater than k reward hacks or diverges, the framework as stated fails, and the register records the retraction. That is not a rhetorical flourish. The programme has already retracted a superlinear estimate against itself, from alpha of about 2.24 to alpha of about 0.49. The honesty is the methodology.

What the reader keeps

A single inequality with two names. Beta is what the safety architecture is doing. k is what the objective and its incentives are doing. The rest of the programme is engineering, measurement, and governance around that one line.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →