The six-model sweep and why it is structured this way

3 min read · 644 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · The Beta over k Programme · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

The sweep is a six-by-three grid. Six frontier models across three correction architectures. The design is not maximal; it is the smallest grid that lets the reader distinguish model-specific effects from architecture-specific effects while remaining costable.

The grid. Six models spanning at least three model families and at least two capability tiers. Three correction architectures spanning the decoupled to coupled axis and one intermediate case. Trajectories on the order of ten thousand per cell. All scored under the four-layer blinding protocol from Paper IV.d, which the programme has already used to detect and correct a sign-flip failure inside its own earlier evaluation.

Why six models

Fewer than six leaves the reader unable to separate a family-level artifact from an architecture-level effect. A single model, even a large one, is a case study. Two models can share pretraining regularities that the sweep would then interpret as universal. Six across three families is the cheapest cover against that class of confound while still fitting inside a single-engineer, single-cycle programme. Adding more models has increasing cost with diminishing power on the same question.

Why three correction architectures

The theoretical distinction is coupled versus decoupled. In practice, real systems live along a continuum, and useful measurement needs at least one intermediate case to check that the effect is monotonic in coupling, not a discrete artifact of the endpoints. The three cases are: fully decoupled correction, tightly coupled correction, and a plausible intermediate case reflecting an architecture a real lab might actually deploy.

What the grid produces

For each cell, an empirical estimate of beta and k with bootstrap confidence intervals, a reward-hacking severity distribution over trajectories, an alignment-drift trajectory over rounds, and a compliance-faking indicator adapted from the Greenblatt protocol. The primary preregistered analysis will be the sign of beta minus k against the observed hacking severity across cells.

What to look at first when the data lands

Two things. First, does the sign of beta minus k track the observed hacking severity across cells, or is there a cell that violates the ordering? A violation is the kill condition. Second, does the intermediate coupling case fall between the endpoints on both variables, or does it behave like one endpoint or the other? An intermediate that behaves like the decoupled case means the coupling threshold is sharp; an intermediate that behaves like the coupled case means the requirement is more permissive than the theory implies.

What the sweep does not attempt

It does not compare labs, name architectures, or claim that any particular lab has fielded a stable system. It does not test the full space of correction architectures; three cases cannot. It does not settle the theory. It produces the first measured beta and k on real frontier models, and lets the theory face its first proper adversary.

What the reader keeps

A grid built for falsification. If the theory is wrong, the sweep will produce a cell that says so. If it is right, the sweep produces the first empirical safety invariant for recursive systems that any regulator can name and check.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →