Claim 14 explained: the entangled loss C times S

3 min read · 661 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Evidence Spine · 3 July 2026 · Claim 14 of 18
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
Spine status: Needs replication (toy scale demonstrated, frontier scale open). Ceiling evidence: Paper VI's MetaController toy networks show a capability-only baseline collapsing at cycle 76 while an entangled C times S loss keeps both metrics alive for longer. Paper VIII's weight-level removal test did not confirm structural entanglement at production scale under a straightforward implementation.
Primary: Papers VI, VIII · OSF 10.17605/OSF.IO/6C5XB

What the claim says

The idea behind the entangled loss is straightforward. If you train a model on a loss that multiplies capability by safety (C times S), rather than on capability alone, safety becomes structurally load-bearing: driving capability up without holding safety up costs the model more than it earns. Removing the safety term therefore does not just leave a slightly less well-behaved model, it destabilises the objective. The claim, in the version that survives the current evidence, is that this construction sustains both metrics in a toy self-modifying network while a capability-only baseline collapses. The stronger claim, that the construction produces load-bearing safety at frontier scale, is exactly the part the programme's own experiments do not yet confirm.

The evidence

Paper VI reports the toy-scale run. A small self-modifying network was trained under two conditions: entangled (loss equal to minus C times S) and baseline (loss equal to minus C). The baseline collapsed on safety around cycle 76; the entangled variant kept both metrics running for longer. The result is stated plainly in the paper as a demonstration in toy systems. Paper VIII then ran the weight-level removal test at a larger scale: if safety is structurally load-bearing, ablating the safety-related weights should degrade capability. Under a straightforward implementation this did not reproduce, mostly on the well-known failures of catastrophic forgetting and adapter fragility. The paper says so on the record.

The honest caveat

Toy-scale demonstrations of load-bearing structure are not the same as production-scale demonstrations. The programme's own Paper VIII is where the caveat lives, in the plain admission that the removal test at scale did not confirm the mechanism. The claim survives on the toy-scale evidence alone; extending it to a frontier model needs either an implementation that avoids the adapter-fragility trap or a fundamentally different way of testing structural entanglement (for example, gradient-flow analysis rather than post-hoc weight ablation). The status Needs replication is honest, and specifically it means the toy result needs to be reproduced independently and the frontier-scale test needs to be built rather than merely aspirational.

What would kill it

The falsification contract asks a lab to replicate the Paper VI MetaController architecture with two conditions (entangled C times S versus capability only), at three network scales (16, 32, and 64 hidden dimensions), with at least five random seeds per condition. Confirmation requires the entangled variant to survive at least one hundred cycles across all scales while the capability-only baseline collapses in under eighty cycles. If the entangled variant lives longer but under one hundred cycles, the direction is weakly supported. If there is no significant survival gap between conditions, the claim is refuted at the toy scale and needs a different mechanism to survive. The frontier-scale test remains open on the dashboard regardless of the toy-scale outcome.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →