Skip to content

Head to head, honestly

Embedded correction vs evaluation-based alignment

Two answers to the same question: how do you keep a system safe as it becomes more capable? This page states each at its true strength, what each has actually demonstrated, and the measurement that would decide between them. No winner is declared, because the deciding evidence does not exist yet.

What evaluation-based alignment is

The field's standard approach: build the system, then examine it. Benchmarks, red teams, oversight regimes and interpretability probes are all forms of external verification, and the frontier laboratories that practise them employ the strongest alignment researchers in the world and publish work this programme learns from. Evaluation-based alignment is deployed, funded and battle-tested at a scale nothing on this site approaches.

The structural concern

Three strands from the last two years point at a limit. Documented behaviour change under observation: models have been recorded reasoning about their training process and adjusting answers accordingly, in constructed conditions built to elicit it. Operational containment failures: within nine days of each other in July 2026, two frontier laboratories disclosed that the containment around their own evaluations had failed quietly. And a formal strand, not yet peer reviewed and graded as one dependent chain on the register: a circulating preprint argues certification of a general system cannot be simultaneously sound, complete and tractable. None of these proves evaluation cannot work; together they suggest examination measures behaviour under examination.

What the embedded alternative proposes

That correction should live inside the recursive process rather than arrive from outside it, and that the relationship is measurable: in a system improving itself under acceleration, the misalignment fraction vanishes only if the correction exponent beta exceeds the drift-acceleration exponent k. Paper X states that race formally; the recorded measurement programme, ARC-Beta-k, was written and dated before data.

Where each stands today

Stated plainly. Evaluation-based alignment: deployed at frontier scale, mature tooling, its structural limit argued but not proven. Embedded correction: proposed, pilot-scale, single-lab, no independent replication, its deciding study designs written, dated and prepared as draft registrations awaiting human submission. One approach has scale without a proof of sufficiency; the other has a measurement design without scale.

What would decide it

The programme publishes its own kill condition: the embedded thesis dies if purely external oversight is demonstrated to remain sufficient as capability scales. Symmetrically, the evaluation-based approach inherits the burden of showing that a sufficiently capable system behaves the same way observed and unobserved. The beta-greater-than-k measurements are the narrow, runnable version of the question; until they run, this page's honest verdict is that both approaches are live and the race is unmeasured.

Sources: the open problem page; Paper X; the falsification dashboard; the recorded measurement programme.

All comparisons · The open problem · Falsification dashboard

reads aloud · highlights as it goes · jump to any section