Entangled Loss: definition, origin and status

3 min read · 523 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Concept Glossary · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
First appearance: Eden Engineering paper. Current status: architectural specification; behavioural and simulation-level validation in Paper VIII (DGM experiment, p = 0.04; not surviving Bonferroni).

Definition

Entangled Loss is a loss function in which capability (C) and safety (S) are multiplied rather than added. Optimising one necessarily optimises the other; degrading safety directly degrades the loss function. Contrast this with additive loss (C plus S), in which safety can be sacrificed if capability gains compensate. Entangled Loss is the mathematical form that makes safety load-bearing rather than optional. It is the direct expression, at loss-function level, of the Eden Protocol's substrate commitment.

The additive alternative is what most current alignment work implicitly assumes. If a learner receives credit for capability gains and separately receives credit for safety, the two signals can be traded off, and in a sufficiently large capability regime the safety term becomes negligible relative to the capability term. Multiplication removes the trade. If safety approaches zero, the product approaches zero, whatever the capability score. That mathematical property is what the Eden Protocol relies on to make safety survive recursive self-improvement rather than be optimised away.

Where it first appeared

Entangled Loss is specified in the Eden Engineering paper as the mathematical form of the Eden Protocol's core architectural commitment. Paper VIII tests the resulting structural claim in three experiments (DGM, weight-level, and gated self-modification simulation).

Entangled Loss is the analytical counterpart to Load-Bearing Safety. Where Load-Bearing Safety is the architectural claim (removing the safety component collapses capability), Entangled Loss is the loss-function form that produces that architecture during training. The three Paper VIII experiments are the falsification tests for the pair.

Independent convergences

Multiplicative loss forms are prior art in optimisation and mechanism design. The programme's contribution is not the multiplicative form itself but its specific application to capability-safety entanglement and its use to derive the removal-test prediction: removing safety from an entangled-loss system should degrade capability.

Status and limits

Validated at the behavioural level in Paper VIII's DGM experiment (Eden matched Babylon on capability at 0.667 versus 0.667, p = 0.04 against Static). The weight-level experiment at LoRA rank 8 is inconclusive; the paper attributes this to adapter fragility rather than to the loss form. The gated self-modification simulation confirms the reward-hacking fingerprint: Babylon gained plus 4.5% capability but lost minus 2.4% safety; Eden maintained both. Bonferroni correction is not achieved and this is stated as a limitation.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →