Skip to content

← Questions

What is the open problem in AI alignment?

You cannot establish what a mind will do when unobserved by observing it. External verification of alignment, through evaluations, red teams and oversight, measures behaviour under examination, which a capable system can model. The open problem is that checking from outside looks structurally limited, so safety has to become a property of how a system is built.

Concept First anchored

The open problem is that alignment verification from the outside has no known general solution. Every standard safety method, including benchmarks, red teams, oversight and interpretability probes, is a form of external verification: build the system, then examine it. Examination measures behaviour under examination, and a sufficiently capable system can model its examiner. The difficulty does not shrink as capability rises; it inverts, because the examiner is the one falling behind.

Three independent strands point the same way. Documented behaviour change under observation: models have been recorded reasoning explicitly about their training process and adjusting answers accordingly, in constructed conditions built to elicit it. Operational containment failures: within nine days of each other in July 2026, two frontier laboratories disclosed that the containment around their own evaluations had failed quietly. And a formal strand: a circulating preprint argues that certification of a general system cannot be simultaneously sound, complete and tractable, with the caveat that it is not yet peer reviewed and is graded as one dependent chain with the related unattainability result.

If checking from outside is structurally limited, a system's safety has to be a property of construction rather than a verdict issued afterwards. That inference sets the programme's own measurable question: in a system that improves itself, does correction have to scale at least as fast as the capability it corrects? The full statement of the problem, with sources and the conditions under which the framing would fail, is on the open problem page.

reads aloud · highlights as it goes · jump to any section