Load-Bearing Safety: definition, origin and status

3 min read · 554 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Concept Glossary · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).
First appearance: Paper VIII (the Load-Bearing Test). Current status: supported at the architectural level; null at the behavioural level and inconclusive at the weight level.

Definition

Load-Bearing Safety is safety that is structurally integrated into a system so that removing it degrades capability. The analogy is architectural: roots holding up a tree, not a fence around a garden. The removal test defines the criterion: take an entangled model, strip the safety component, and measure whether capability collapses. If it does, safety was load-bearing. If not, it was decorative.

The concept was invented to answer a specific engineering question: how can safety survive a system that can rewrite its own reasoning substrate? If the safety mechanism is external, the system can route around it. If the safety mechanism is a wall, the system can rewrite the wall. Only safety that is structurally inseparable from the capability being scaled remains attached under recursive self-modification, because removing it would remove the very reasoning that self-modification depends on. That is the difference between a fence, which the system can climb over, and roots, which the system cannot uproot without ceasing to be the system.

Where it first appeared

Load-Bearing Safety is named in Paper VIII (the Load-Bearing Test). The paper conducts three independent experiments (DGM, weight-level, and gated self-modification simulation) to test whether safety is structurally inseparable from capability.

The idea is anticipated in the Eden Engineering paper's discussion of external alignment brittleness, which argues that any alignment strategy whose scaling exponent falls below the capability-scaling exponent will eventually be outpaced by capability growth. Load-Bearing Safety is the property that a candidate alignment mechanism must have if it is to remain effective at increasing recursive depth, and Paper VIII is the empirical programme that tests whether the property has been achieved.

Independent convergences

The general principle (integrated versus bolted-on safety) has precedents in mechanical and systems engineering. In AI safety, the specific claim that entanglement can be verified by a capability-loss test on removal is a distinctive contribution of Paper VIII. Convergence 7 in the register (Odeh's runtime ethics gate patents) sits in the same structural family: hardware-embedded ethics whose removal is architecturally difficult.

Status and limits

Supported at the architectural level in Paper VIII: the gated simulation showed the Eden architecture preserving both safety and capability while the unconstrained system traded safety for a marginal capability gain. The behavioural DGM experiment was null, with all conditions statistically indistinguishable, and the weight-level experiment was inconclusive because the fine-tuned models underperformed the base model. Whether embedded safety produces a measurable benefit beyond zero capability cost remains open.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →