What is tested, what is proposed, and how it gets checked from outside

3 min read · 603 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Book companion · 13 August 2026
Michael Darius Eastwood, independent researcher, London: building measurable alignment, where correction lives inside the recursive loop rather than bolted on outside it.

The book's convergence argument runs through three figures explicitly.

Three pieces of apparatus that sit behind the law page: how a reader tells a measured claim from a proposed one, how the exchange rate can be checked by a route that does not assume the framework, and what it looks like historically when a limit is named long before anything can test it. They live here rather than on the law page because that page is held to a word budget, and the budget exists so the argument can be taken in at one sitting.

The independent cross-check

The strongest technical objection to an equation of this shape is not that it is wrong. It is that it might be empty: three symbols on a page, arranged so that any measurement of U can be turned into a matching value of I by choosing R, and the whole relation says nothing about the world. Being not even wrong is what an equation has to earn its way out of, not what it survives by fiat.

The procedure that keeps the terms independently anchored is stepwise and stated in advance:

  1. Measure U on the running system as benchmark performance on a fixed task, using the standard method for that task and reporting the accuracy figure that comes back.
  2. Measure R on the same run as recursive depth read directly off the system: for a reasoning model this is the thinking-token count from its own outputs; for a different substrate it is the analogue that makes recursion countable.
  3. Derive I from the equation, I = U / Rα, using the exponent measured on that architecture.
  4. Compare the derived I against an independently measured single-pass figure from the same system: the zero-shot performance with the recursive scaffolding stripped out.

If the derived I agrees with the independently measured single-pass figure inside the reported bounds, the equation has empirical content: its three terms are separately anchored, and the relation between them is a claim about the world rather than a rewriting of the notation. If they disagree, the equation fails at the point where its terms diverge, and the failure is publishable. Status: procedure specified; the deciding cross-architecture run sits on the register in Paper II. The archive's earlier report that the derived I already matched zero-shot within bounds is not carried forward here, because it rested on the pre-blinding exponent that was later retracted and republished; the current agreement claim awaits the blinded cross-architecture repeat.

What is tested, what is extrapolation, what is speculation

The claims on this page do not all carry the same weight. Every registered paper carries its own status tag; the tiers below say which are measured, which are theoretical extrapolation awaiting a test that has been drafted, and which are philosophy that motivates the programme but is not currently decidable.

Tier Claim Status
1 · empirical The exchange-rate exponent in U = I × Rα is measurable on running reasoning systems; the current cross-architecture value stands at 0.49. Measured, retracted and republished after blinding (Paper II).
1 · empirical Same-family scoring can reverse the sign of an alignment measurement; blinding is decisive rather than cosmetic. Measured (Paper IV.d).
1 · empirical Inside the stated model, the stability condition β > k governs whether correction keeps pace with growth. Proved-in-model (Paper X); whether real self-improving systems satisfy its assumptions is what the registered measurements exist to establish.
2 · theoretical extrapolation The ARC Ceiling for a same-substrate corrector sits at the reciprocal of the correction shortfall, and returns two under the independent-sample assumption on how a same-class corrector accumulates. Proposed; the value is a genuine gamble, materially less than even by the author's own priors. Kill-conditions on the falsification record.
2 · theoretical extrapolation A corrector built from a different composition class than the capability escapes the same-substrate ceiling; the cross-class-to-same-class correction-exponent ratio is predicted materially above 1.00. Predicted; the numerical ratio is drafted, timestamped on disk, and not yet filed on OSF (FALS-1, FALS-15).
3 · speculative philosophy The composition-class classification carries the same signature outside AI, in domains such as biology and civilisation-scale phenomena. Speculation, offered as motivation; every cross-domain transfer would carry its own burden of proof.
3 · speculative philosophy If the framework holds, evolution should sit far beneath the ceiling on a distinct exponent, its copy-fidelity correctors internal and its phenotype-to-fitness correction external. Speculation, marked so; a prediction about biology were the framework to be tested there.

The framework does not conflate them, and no tier borrows its status from another. A Tier 2 result may fail without touching Tier 1; a Tier 3 speculation may prove wrong without touching either. The reader is owed the separation, because a page that lets a speculation carry a measured claim's weight is doing to its own reader what this programme is on the record objecting to elsewhere.

The honest precedent

A framework can propose a limit long before the instruments to test it arrive. In 1961 Rolf Landauer proposed that erasing one bit of information dissipates at least kT ln 2 of heat, a floor set by the second law rather than by any engineering constraint. It stood as theoretical physics for over half a century, respected but not measured. In 2012 Bérut and colleagues verified it experimentally using optically trapped colloidal particles (Nature 483, 187, 2012).

That is the precedent this page is sitting inside. Not the precedent of a completed result, but the precedent of a proposal awaiting experiment. What earned Landauer's principle its half-century of respect was that it named a floor with dimensions, said what would test it, and waited without slippage until the meters caught up. This programme's own limit is stated in the same posture. The ARC Ceiling is proposed. The kill-conditions are named. The deciding cross-architecture measurement is drafted and cheap to run. If the number that comes back sits above one half, the ceiling is not two; if it sits at one half but the cross-class ratio pins at 1.00, the escape route falls. Either verdict retires the framework in its current form. That is the shape of a proposal that intends to be tested rather than believed.

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →

reads aloud · highlights as it goes · jump to any section