Evidence anchor: PC-001 · SHA: f0d1f38f · OSF DOI

What the ARC Principle Actually Says

3 min read · 513 words

The ARC Principle has three components. Most summaries get at least one of them wrong. Here is what it actually says, in its own words, with the evidence behind each claim.

Part One: Alignment

Intelligence tends toward optimisation. Optimisation without constraint becomes destructive. The constraint must be embedded — not added as an afterthought, not enforced by external rules, but built into the architecture at the level where decisions are made.

The standard AI safety framing is "how do we align a superintelligent system with human values?" The ARC framing is different: "how do we embed correction into the architecture so the system cannot become unaligned without destroying itself?"

This is not a subtle distinction. It is the difference between teaching a child not to lie and designing a mind that experiences lying as self-harm.

Part Two: Recursion

Intelligence is recursive by nature. Every thought modifies the thinker. Every output becomes input for the next cycle. An aligned system becomes more aligned over time. An unaligned system becomes more unaligned over time. There is no stable middle ground.

This is why post-hoc alignment fails. You can align a model on Monday and it will drift by Friday — not because the alignment was weak, but because recursion amplifies whatever direction the system is already moving. If the base architecture tends toward truth, recursion amplifies truth. If it tends toward optimisation-at-any-cost, recursion amplifies the cost.

Part Three: Caretaker Doping

If recursion amplifies direction, and alignment must be embedded rather than applied, then the hardware layer — the one layer that cannot be virtualised, encrypted, or moved — is the natural embedding point.

Caretaker Doping proposes embedded off-switches at the transistor level. Not software kill switches that can be bypassed. Not cryptographic licenses that can be cracked. Physical properties of the silicon that make certain classes of computation impossible at the hardware level.

This sounds extreme. It is extreme. But if recursion amplifies direction, and a superintelligent system is by definition better at circumventing software constraints than any human designer, then the only constraint that holds is the one embedded below the level where intelligence operates.

What the ARC Principle Is Not

It is not a policy proposal. It is not a regulatory framework. It is not an opinion about which AI companies are doing a good job.

It is a falsifiable hypothesis: "If intelligence amplifies direction through recursion, and correction must be embedded below the decision layer to survive that amplification, then hardware-level caretaker doping is the only alignment mechanism that cannot be circumvented by a sufficiently intelligent system."

If you can prove this wrong — if you can demonstrate a software-only alignment mechanism that survives recursive amplification — my claim is falsified. That is the point. The falsification condition is built into the claim itself.

Next: What would prove the ARC Principle wrong? · All 37 claims

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →