Skip to content

Run it yourself

Everything on this estate is self-attested until someone who is not me runs the tests. This page exists to make that as easy as possible: the runnable tests, what each decides, and a dated register of independent attempts that is honest at zero and updates as runs arrive.

The attempts register

This section renders from the replication-attempts register at every build. As of 2026-08-28: 0 independent attempts recorded. The register states its own rule: an independent run outranks the programme’s own in either direction, and a failed replication enters with the same prominence as a success.

No independent attempt has been recorded yet. That is the honest current state, dated above, and it updates the day an attempt with a checkable artefact arrives. The tests below need no permission and no contact to run.

What you can run today

The baked-in versus computed alignment test (Paper IV.a)
Decides: whether a stated value survives deletion of its statement, or was resting on a document.Requires: API access to any frontier model · the test’s own folder
The ARC scaling challenge (the decisive trial, openly licensed)
Decides: the scaling-regime question the ARC Principle stakes itself on.Requires: a training or evaluation pipeline; full protocol in the repository README · the test’s own folder

The decisive trial and the runnable tests are openly licensed: any laboratory or individual can run them without permission and without contact. Completed attempts reach this register by email with the run artefacts; entry requires a checkable artefact, never a claim alone. Send completed runs to michael@michaeldariuseastwood.com.

The kill conditions the tests feed · The registered programme · Research hub

reads aloud · highlights as it goes · jump to any section