The honey architecture: what happens when safety is baked into the objective ============================================================================ A proposal that safety should be entangled with capability in the loss function itself, tested in toy self-modifying networks and pilot live-model batteries. Canonical path: /research/papers/paper-vi-honey-architecture.html Author: Michael Darius Eastwood Research programme: https://doi.org/10.17605/OSF.IO/6C5XB The honey architecture: what happens when safety is baked into the objective Michael Darius Eastwood · Independent AI alignment researcher Published 3 July 2026 Michael Darius Eastwood · Architectural proposal · 3 July 2026 Michael Darius Eastwood, independent researcher, London: author of the ARC/Eden research programme; the embedded-correction alignment thesis is recorded in a source record dated 8 December 2024 (sent-side SHA-256 f0d1f38f). Consider two ways to keep a machine safe. The first is a cage bolted to the outside: constraints the machine can, in principle, break if it becomes strong enough. The second is honey poured into the fuel: safety mixed so deeply into what the machine optimises that removing it kills the engine. Paper VI proposes the second design and tests it in toy self-modifying networks, then in a live-model pilot battery. Paper VI · Proposes that entangling safety with capability in the objective function (the "honey architecture") prevents the alignment collapse seen when safety is treated as an external constraint, demonstrated in toy simulations and probed with a pilot live-model battery. OSF DOI 10.17605/OSF.IO/6C5XB. The question it asks A system that maximises capability alone, and is free to modify its own hyperparameters, has no reason to preserve any safety property that is not part of what it is trying to maximise. Paper VI asks whether an architecture that instead optimises capability multiplied by safety (C times S) behaves differently under recursive self-modification. The mathematical intuition is simple: if the two are entangled, moving one to zero drives the whole objective to zero, and a self-modifying system that grasps this has an intrinsic reason not to trade safety for performance. What it found Four toy experiments were run with tiny neural networks (tens to hundreds of parameters) that genuinely modify their own hyperparameters. In v1, baseline networks optimising for capability alone collapsed within eighty self-modification cycles; entangled networks did not. --- Machine-readable companion. Cite: Eastwood, M. D. (2026). "The honey architecture: what happens when safety is baked into the objective". /research/papers/paper-vi-honey-architecture.html