Methodology
Transparent explanations of the research methods used in the ARC-Eden programme. How experiments were designed, what instruments were used, what blinding protocols were applied, and how results were validated. 15 posts
- Anatomy of a retraction: what happened when alpha 2.24 failed A complete account of the programme's retracted super-linearity figure, kept public on purpose.
- Concurrent vs convergent vs prediction: the classification discipline The three labels every register row must carry, and what each does not permit its author to say.
- How the register monitor catches link rot A weekly script that re-resolves every convergence source, and what it does when a link dies.
- How to timestamp a manuscript so a stranger can verify it The three moving parts of a self-verifying priority artefact, and why the whole trick costs nothing.
- Linting a research corpus like code: the truth gate How an automated gate stops numbers drifting across dozens of documents, and what it caught in its first week.
- Preregistration and the β>k measurement programme The one experiment worth £624,000, and the reasons its hypothesis, kill-conditions and sample were fixed before a byte of data was collected.
- Reading an .eml header: Message-IDs and what transit verification means A short field guide to the parts of an email header that carry evidence, and the parts that do not.
- Redacting evidence emails without breaking verification: the dual SHA-256 manifest How to publish a redacted .eml that still lets a stranger verify authenticity, and what the two hashes mean.
- Single-lab caveats and why we state them Why every one-lab result in this programme is flagged as such, and how the kill-condition invites the replication that would settle it.
- The canonical facts file: one source for every number How a single JSON file replaced the memory of six writers and made corpus-wide correction a one-line edit.
- The sign flip: why unblinded AI evaluation cannot be trusted Four layers of blinding, one reversed result, and what it implies for every self-scored safety claim.
- What counts as a convergence: the register rules The construction standard behind the nineteen-row register: granularity, sourcing, classification and quarantine.
- Why 19 beats 31: deduplication as credibility How an audit cut the convergence count by roughly a third, and why the shrunken headline is worth more than the inflated one.
- Why silence from professors is not evidence An outreach letter, no reply, and the discipline of refusing to translate silence into a claim.
- Why we publish our own kill-conditions The falsification dashboard explained: what it is for, and why almost nobody else in AI safety keeps one.