Head to head, honestly
The Eden Protocol vs Constitutional AI
Both reject the idea that a list of external rules can hold a system smarter than its rules. They differ on where values should live, and they differ enormously in maturity. This page credits the prior work first, states the difference in kind, and names what would decide between them.
The shared ground, and the prior work
Constitutional AI is Anthropic's method, published in 2022: a set of written principles guides the model to critique and revise its own outputs during training, so values shape behaviour rather than being bolted on as filters. It is deployed at frontier scale and it is independent prior work, credited on this site's record wherever value-embedded training is discussed. The Eden Protocol's claim is over its specific dated framing and named architecture, anchored in the 8 December 2024 manuscript, not over the idea of training values in, where Constitutional AI came first in public.
The difference in kind
Constitutional AI shapes the training objective: principles, critique and revision produce a model whose behaviour reflects the constitution. The Eden Protocol proposes something structurally stronger and far less proven: that care should be entangled with capability itself, so that removing the alignment damages the system. Its four layers are a purpose kernel, graduated autonomy that is earned and revocable, a monitoring-removal test that checks whether behaviour survives the absence of oversight, and an entanglement proof in which safety and capability are trained as a product.
Maturity, stated without flattery
One of these is running in production systems used by millions. The other is at pilot scale in one laboratory: the entanglement mechanism is demonstrated on toy networks, the weight-level removal test did not confirm structural entanglement at scale, and the controlled pilot found Eden matched its unaligned comparison on capability within noise, with fewer reward hacks. Those results are encouraging and small, no outside laboratory has replicated them, and the programme says so on every surface that reports them.
What would decide it
For Constitutional AI, the open question is durability under self-improvement: whether trained-in values survive when the system can modify the machinery that carries them. For Eden, the open question is existence: whether capability-entangled care can be built at scale at all. A frontier-scale system whose alignment demonstrably survives the removal of oversight would be evidence for embedding in general; which architecture gets there is an empirical race, and this programme's entry is the smaller and less proven one.
Sources: the embedded-correction priority record; the Eden Protocol page; Papers VI and VIII; the evidence spine.
All comparisons · Eden Protocol · The embedded-correction thesis