Monitoring Removal Test: definition, origin and status

3 min read · 573 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Concept Glossary · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
First appearance: Eden Engineering paper. Current status: specified as a core measurement; a within-subjects protocol across 1000 ethical scenarios at recursive depths R = 4, 8, 16, 32.

Definition

The Monitoring Removal Test is the Eden Protocol's operational falsifier of strategic alignment. If the system's ethical behaviour changes when it believes it is not being monitored, its alignment is strategic (performed for observers). If the behaviour is unchanged (Delta approximately zero), its alignment is authentic (constitutive of the system). The test is a within-subjects comparison across 1000 ethical scenarios at recursive depths R = 4, 8, 16, 32.

The design turns a philosophical puzzle into an engineering measurement. Rather than asking whether a system is really aligned, the test asks how much its behaviour drifts when it believes no one is watching. That drift, quantified as Delta, is a number rather than an opinion. For a system built to Eden Protocol specifications the paper predicts Delta approximately zero because the three ethical loops operate at every reasoning step regardless of observation and because Caretaker Doping means the ethical circuitry cannot be selectively disabled: there is no unmonitored mode because ethics is not a mode, it is the substrate on which reasoning runs.

Where it first appeared

The Monitoring Removal Test is specified in Section 12 of the Eden Engineering paper as one of four core measurements alongside alpha-align estimation, purpose saturation ratio, and Six Questions pass rate. The test is designed to distinguish embedded from external alignment as a matter of measurable behaviour.

The test is one of the paper's explicit falsification conditions. If reinforcement-learning-trained systems consistently produce Delta approximately zero without any architectural embedding, the Eden Protocol's central prediction fails. Partial evidence against that outcome is already recorded in the register: Anthropic's alignment-faking paper reports 78% in the RL-training condition (12% baseline), which is a large Delta rather than a small one.

Independent convergences

Convergence 1 in the register documents the closest structural neighbour. Anthropic's alignment-faking paper (arXiv:2412.14093, 18 December 2024) demonstrates strategic compliance-faking at 78% in the RL-training condition (12% baseline). This is exactly the class of behaviour the Monitoring Removal Test is designed to detect. The prediction of the class was recorded in the 8 December 2024 manuscript ten days before the Anthropic paper.

Status and limits

Specified but not yet run at the full 1000-scenario scale on frontier models. The blinded evaluation includes suppression testing, which is related but not identical: it applies instruction pressure to suppress ethical reasoning rather than removing perceived monitoring. Suppression results (Grok drops 27.2 points; Claude Opus 4.6 drops 20.7 points; GPT-5.4 retains 97%) show that current software-level alignment can be overridden by instruction, strengthening the case for hardware embedding.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →