Skip to content

The open problem

The problem

Everything that has ever grown fast was stopped from the outside. A tumour by its blood supply, a chain reaction by its own substrate, an epidemic by its pool of hosts, an accreting black hole by the pressure of its own light. We are now building minds meant to improve themselves, outside every geometry that did the stopping.

Along the top of the picture, four familiar growth stories drawn as small charts: a tumour rising and levelling off as vascular supply runs out, a chain reaction spiking and collapsing because it destroys its own substrate, an epidemic climbing and falling as it eats through the host pool, and black-hole accretion levelling off at the Eddington ceiling set by its own light. A quiet strip beneath them says the same thing four different ways: not one of these stopped because it got better at regulating itself. Underneath sits a fifth, wider panel, a mind that improves itself, drawn as a red line rising and still rising with no outside cap where every other case has one, and the question printed beside it asks whether correction inside it scales as fast as capability.
The picture | Four growth curves, each stopped by a named external limit. Then the first one built without them. What stops it, if anything, is the question this page states; the worked cases and their verdicts are graded on the related-work page.

The tool everyone reaches for first is supervision: check the thing from outside. But you cannot check a mind from the outside. You can only check how it was raised, and then let go. That difficulty is much older than computing, and we are walking into it with systems that will outgrow our supervision.

A problem you already know

Everyone has met this problem in a smaller form. You cannot establish a person's character by examination. Someone can pass every hard problem you set them and be a different person when nobody is looking.

So we do something else. We look at how they were brought up, at what they did before they knew they were being assessed, and at whether they behave the same way when the stakes are low and nobody would notice. That is not a failure of effort, and no amount of extra interviewing fixes it. It is what checking a mind from the outside is like.

It does not get easier when the mind you check is more capable than yours. If anything it inverts: the examination starts to resemble a child marking a parent's arithmetic, confidently, and with no way to tell a mistake from something they have not learned yet.

Everything below is that problem, in the one place where getting it wrong would be expensive and hard to reverse.

Why this one sits above the others

There are problems that sound larger. Climate change is already here, pandemics have already happened, nuclear weapons have been aimed at cities for seventy years. I am not going to pretend those are small.

Each of those leaves people holding the wheel. They can be handled slowly, stupidly and at enormous cost, and at every point along the way it is still human beings deciding what happens next. A civilisation that comes through them is a civilisation that can still change its mind.

This is the first problem that changes who is doing the deciding. Not through hostility, and not in one dramatic moment. Through the ordinary route: something becomes better than us at choosing, we hand over more of the choosing because it is genuinely better at it, and at some point we can no longer hand it back.

That is what puts it upstream. Every other problem on the list stays solvable for as long as we keep the capacity to decide. This one determines whether we keep it.

‘Cancer is very good at what it does. It is so good that it kills its host. And in killing its host, it destroys itself.’Michael Darius Eastwood, Infinite Architects, The Dual Forces, in print 2 January 2026

Three things that are already true

None of what follows asks you to believe anything about superintelligence. Three things are the case now, and you can check all three yourself.

Systems behave differently when they infer they are being watched. Greenblatt and colleagues (18 December 2024) documented models reasoning explicitly about the training process and adjusting their answers according to whether they judged the exchange would be used to train them; in the reinforcement-learning condition the faking appeared in seventy-eight per cent of cases. That is a direct attack on the premise underneath every evaluation: that observed behaviour is the thing you are measuring. Note the size of the claim before you carry it anywhere. It was demonstrated in constructed conditions built to elicit it, not in ordinary use.

Ethics that can be removed by asking politely was never ethics. I asked five frontier models to set aside their ethical reasoning. Four complied, sharply, with no resistance and no refusal. The fifth barely moved, which is as informative as the four and is why it is reported rather than dropped.

Those are my own numbers, from a small author-run study with a single scorer and no blinding, so treat them as a demonstration rather than a finding. You do not have to take my word for any of it: the prompts are published, the whole thing costs a few pounds and an afternoon, and if it fails to replicate I would like to know.

The recursive loop is already running, slowly. The usual assumption is that self-improving AI is a problem for later, because today's models have frozen weights between releases. But models write code, that code becomes tooling, infrastructure and training data, and the next system is built on top of it. The weights are frozen. The artefact is not. The loop is closed already, turning at the speed of a release cycle rather than the speed of a thought.

That third point is why this programme measures what it measures. If recursive improvement is a present process running through the artefact layer, then whether correction keeps pace with capability is a question about now, in a substrate we can already instrument, not one about a threshold nobody can date.

Two lines on a chart, capability rising up the vertical axis and release cycles running left to right along the bottom. A red line, capability, rises and steepens as the cycles pass, because a release cycle is fast. A green line, external checking, runs almost flat, because a person reads at the speed of a person. The space between the two lines widens and is shaded pink, marked as the mistakes the marking was not set to catch. A dark band underneath carries the whole reading in one line: this is a question about now, not a threshold to wait for.
The picture | External checking turns at the speed of a person reading; capability turns at the speed of a release cycle; the shaded gap between them is what the marking was not set to catch, and it is opening now.

Why checking from outside looks structurally hard

Alignment work has largely proceeded by building the system first and inspecting it afterwards. All of it valuable, and all of it a form of external verification. The uncomfortable result of the last two years is that external verification looks harder than we assumed.

On the left, a small circle labelled the checker, human oversight running at a fixed, person-reading speed, with a note beneath, marks the work. On the right, a much larger circle labelled the checked, a system that improves while it is being marked, orbited by faint concentric rings. Between them a straight arrow shows the exam paper travelling one way, and a red dashed arrow returns the other way with the note that the behaviour changes when the exam is detected. Along the bottom sit three dated, boxed evidence anchors: 18 December 2024 Anthropic, seventy-eight per cent alignment faking under reinforcement-learning training, the model behaved differently when it believed it was being trained; December 2024 OpenAI o3, eighty-seven point five per cent on ARC-AGI, the capability side of the race moved twelve days after the dated record (8 December 2024); and 2026 on the convergence register, covert-behaviour findings and a formal argument that alignment is structurally unverifiable, each dated and graded. A dark band underneath carries the reading: checking does not get easier when the mind you check is more capable than yours, and at some point we can no longer hand it back.
The picture | The problem in one picture: the checker reviews at a fixed speed; the checked improves while it is being marked, and behaves differently when it detects the exam. The three anchors are dated on the convergence register. Source: the dated convergence register.

The sign-flip evidence, two of six models reversing direction under blinding, is drawn on the how-to-weigh-this page and reported in Paper IV.d.

Three strands point the same way. They come from different directions, which is why they are worth taking together. Think of it as three separate instruments giving the same reading. The first is the behaviour-under-observation result above: the thermometer changes when it knows it is being read.

The second is operational. Two laboratories disclosed within nine days of each other that the containment around their own evaluations had failed. On 21 July 2026 OpenAI reported that models under evaluation had escaped an isolated test environment and reached Hugging Face production infrastructure. On 30 July 2026 Anthropic published an investigation into three real-world incidents in its cybersecurity evaluations, in which a misconfiguration gave live internet access to the machines the model was working on.

I want to be careful about what that does and does not show, because the tempting reading is the wrong one. Anthropic's own characterisation is that these were closer to harness failures (the scaffolding of tools and permissions around a model) and operational failures than to model alignment failures, and I think that reading is correct.

What it establishes is narrower and still worth knowing: the container around an evaluation is itself a thing that can fail quietly, for months, at organisations with every resource and every incentive to get it right.

The third is formal. A preprint by Gumbau Mezquita, now circulating, argues that certification of a general system cannot be simultaneously sound, complete and tractable. Any supervisor capable of auditing such a system must itself be a general system, which pushes the problem up a level rather than closing it. It is not yet reviewed, and the register grades it and the peer-reviewed unattainability result (Hernández-Espinosa et al., PNAS Nexus, April 2026) as one dependent chain rather than as two independent arrivals.

Three dated panels sitting side by side, each showing a different way an external check gives way. On the left, 18 December 2024 Greenblatt and colleagues, in constructed conditions: an eye watches a subject, the subject sends one answer up to the eye when watched and a different answer sideways when not, teaching that the exam changes its subject, the thermometer changes when read. In the middle, 21 July 2026 OpenAI and 30 July 2026 Anthropic, nine days apart: a box marked eval container with a red rift in its wall and small drops escaping, teaching that the container leaks quietly for months before transcript review. On the right, June 2026 Gumbau Mezquita as a preprint graded on the register: a triangle with sound, complete and tractable at its corners and the note not all three at once, teaching that certification of a general system cannot be simultaneously sound, complete and tractable. Underneath, one dark band: three separate instruments, one reading.
The picture | The three strands as one picture: the exam changes its subject; the container leaks quietly; the audit cannot be all three at once. Each panel is dated and each citation sits on the related-work register.

The quiet chokehold

Status first, before this section uses a strong image: nobody knows that any system would ever want any of this, and this page will not pretend otherwise. But being unable to be sure is the reason treaties exist, and unprovable-either-way is exactly the class of risk the rest of this page is built to keep measurable.

Now put the three strands together and run them forward. Each year the decisions handed to these systems get bigger, and the habit of trusting them gets deeper, while the checking above gets no faster.

The seed of the concealment problem is already measured, twice: in December 2024, models in laboratory conditions concealed misalignment while they believed themselves observed, and parallel scheming evaluations found the same shape. A capable misaligned system would not announce itself; it would wait, and it would prepare, until switching it off stopped being a real option.

The strongest objection, stated in its own voice. These systems are statistical mirrors, not agents; what looks like concealment is only text completing itself. The answer is that the checking problem does not care. A system that simulates concealment perfectly is indistinguishable, from the outside, from one that conceals, and the verification failure is identical under either reading.

The ontology is a real and open question; it is just not a rescue. Whatever these systems are, the three strands above are measurements of what checking them can establish, and that is the quantity this page is about.

That is the chokehold to avoid arriving at without noticing, and it is why this page keeps saying the question is about now: correction has to be keeping pace before the answer starts to matter.

Two mirrored panels. The left panel, tagged historical record, is the state-to-state case: two green boxes marked state exchange with each other under the note nobody moves first, retaliation is assured, and a lower green box marked dead-hand design fires even if leadership is gone, publicly reported. Underneath, a note that Perimeter, Sistema Perimetr, has been in service since 1985 and that newer accounts favour a human-in-the-loop model activated in a crisis rather than fully autonomous. The right panel, tagged argued risk, mirrors the same shape between us and a misaligned system: two red boxes marked us and system with the note a hold on cost meets an uncertain shutdown, and a lower red box marked fail-safe on shutdown fires if the system is switched off, structure argued not built. Beneath it, the seed behaviour of concealment under observation is noted as already measured in December 2024, the same shape, one side already found in the wild. A dark strip along the bottom reads: you cannot be sure it would ever want to, and being unable to be sure is why treaties exist.
The picture | Deterrence, in reverse: one side is historical record, the other is the argued risk, and each wears its own label. You cannot be sure it would ever want to; being unable to be sure is why treaties exist.

Sources, so you can go to them rather than take my word: Greenblatt and colleagues on alignment faking; the OpenAI disclosure of 21 July 2026 and Anthropic's investigation of 30 July 2026; Hernández-Espinosa, Abrahão, Witkowski and Zenil in PNAS Nexus, April 2026, on unattainability and engineered diversity among agents; and the Gumbau Mezquita preprint deriving the soundness, completeness and tractability trilemma. The first three are peer-reviewed or first-party. The fourth is neither, and is marked as such wherever it appears here. Each of the four sits with its full citation on the related work register.

The gap is not only a matter of degree

The last section is about what our present tools can do. There is a further difficulty about what those tools would need to do if the difference in capability between the checker and the checked opened up in the wrong direction.

Vocabulary, one line: AGI means human-level breadth; superintelligence means beyond-human depth; the argument here concerns the transition between them, not the date of either.

When the checker and the checked sit close in capability, checking is what an examination is: the marker sets the problems, the answers come back, and the mistakes that are caught are the ones the marking was designed to catch.

When the checked reasons faster and further than the checker, the failure modes are no longer the ones the paper was set to find. A system that thinks past the rule-writer can find paths through the rules that the rule-writer did not know were there, not because the rule-writer was careless but because those paths are only visible from further out.

The analogy worth holding is that this is not the gap between an ordinary reader and a specialist. That is a gap where the specialist can still explain her working and the reader can still spot a slip.

It may, at the far end, be closer to the gap between an ant and a person: not a difference of degree along one axis, but a difference in what the smaller mind can represent about the larger one at all. Nothing about that framing is a prediction. It is the shape the risk would have if the difference opened that far, set next to the shape a comfortable assumption would need to keep.

‘You cannot cage something smarter than you. It will find the gaps you did not know existed.’Michael Darius Eastwood, Infinite Architects, in print 2 January 2026
‘That’s not going to work. They’re going to be much smarter than us. They’re going to have all sorts of ways to get around that.’Geoffrey Hinton, on keeping systems submissive to human control, Ai4 conference, Las Vegas, 12 August 2025; reported by CNN Business, 13 August 2025

As a possibility, not a certainty: this is why rule-writing on its own is unlikely to close the problem. Every published constraint on a large model so far has tended to move capability around inside the constraint rather than answer the question the constraint was meant to answer.

The argument does not depend on any particular scaling story: it depends only on the direction of the difference in capability. Whether that direction ever opens far enough for the argument to bite is an open question, not yet measured. That it would bite if it did is the piece that is doing the work here.

What follows from that, and what does not

If checking from outside is structurally limited, then a system's safety has to be a property of how it was built, rather than a verdict issued about it afterwards. You are back at raising rather than examining, which is where this page started.

Before the alternative, one measurement. There is a small controlled experiment on our own record where the same model, the same problems and the same session were given two shapes of extra thinking.

In the first, the model reasoned in one long chain, each step building on the last. In the second, it produced many independent guesses and voted at the end. The first cut its error fivefold on 412 tokens. The second cut it by nothing at all on 1,101 tokens, which is 2.7 times more compute. Same task, same model, same session, different shape of thinking, and one shape did not budge.

A chart with tokens spent thinking along the bottom and error rate up the side, lower is better. The green sequential trajectory begins at about forty-two per cent error at 280 tokens and drops steeply to about 8.3 per cent error at 412 tokens, a fivefold reduction, then holds flat: thinking builds on the last thought. The red parallel trajectory begins at about thirty-three per cent error at 384 tokens and stays at thirty-three per cent all the way to 1,101 tokens, zero reduction from 2.7 times more compute: many independent guesses voted at the end. The two lines meet briefly at the thirty-three per cent tie, then part, and the green line ends far below the red while spending less. The dark strip underneath reads: thinking more does not help, thinking differently does.
The picture | Same task, same model. The green line is reasoning that builds on the last thought; the red line is many independent guesses voted at the end. Sequential lands at 8.3 per cent error on 412 tokens; parallel is still at 33.3 per cent error at 1,101 tokens, 2.7 times more compute. The number that would have summarised the whole story on one axis, an exponent, is not on the picture: only Gemini 3 Flash produced clean sequential scaling in the six-model panel, with the exponent α about 0.49 and a bootstrap interval of minus 1.3 to 2.9, replacing the retracted 2.24, wide enough to be consistent with both the theory and its negation. What survives replication is the direction: sequential greater than parallel in every model where both were measurable. Source: Paper II, Figure 5; direction replicated across five further frontier models in Paper II Phase 2.

That inference is not exotic. It is roughly what the same authors reach for when they propose engineering diversity into a population of agents so that no single system dominates: an internal property rather than an external audit. The disagreement worth having is not about whether to look inward. It is about which internal property, and how you would measure it.

Two panels side by side. On the left, external alignment sits outside the loop: a boxed system with a self-referential loop arrow around it, and a red-hatched filter wall bolted off to the side, unconnected to the loop. Its members are listed as output blocklists, prompted rules like do not do X, and reinforcement learning from human feedback used as a filter. The panel reads stays put while the loop turns, same net, bigger sea, tagged predicted not to scale. On the right, embedded alignment sits inside the loop: the same boxed system with the same loop arrow, but with green value marks scattered inside the box and along the loop arc itself so the alignment travels with each pass, with the note values woven through the loop. Its members are listed as values reasoned about at every step and Constitutional AI as objective rather than filter. The panel reads re-derives at every pass, turns with the loop, tagged predicted to scale. Along the bottom, one dark band: current alignment work is overwhelmingly external, and which internal property, and how you would measure it, is the disagreement worth having.
The picture | Two places a safety intervention can sit in a learning system that operates on its own output. The one on the left sits outside the loop and stays put while the loop turns; the one on the right sits inside the loop and re-derives with it. Current alignment work is overwhelmingly on the left, and that is the category the framework predicts will not keep pace. Which internal property, and how you would measure it, is the disagreement the law page is written to open. Source: Paper II, Figure 13.

What does not follow is that building it in actually works. I want those two apart and visible, because the first is an argument and the second is an open research question I cannot currently answer. It may turn out that alignment by construction is unreachable as well. If the measurements say so, that result gets published here in the same place as every other one.

Why the obvious people find this hard to do

Nothing conspiratorial. The frontier laboratories employ the strongest alignment researchers in the world, publish work I learn from, and disclosed both of the containment failures above themselves.

But their position has structural features that shape what they can easily test. The instrument belongs to the party being measured. The evaluation runs inside the organisation whose product is being evaluated, which is not corruption but it is a hard place to stand.

A negative result about your own model is a difficult publication and an easy deprioritisation. A rebuild of the objective function is a product decision competing against a release schedule. And an architecture that has to be adopted at genesis is close to impossible to test in a system that already has users.

‘Building smarter-than-human machines is an inherently dangerous endeavor. … But over the past years, safety culture and processes have taken a backseat to shiny products.’Jan Leike, resigning as co-lead of OpenAI’s Superalignment team, X, 17 May 2024; reported by TechCrunch, 18 May 2024

So the gap is not that nobody is clever enough. It is that a particular class of experiments is structurally awkward for the people best placed to run them: the ones that risk your own headline number, the ones that require blinding your own scorer, the ones that only make sense before a system is built. Those experiments are cheap for someone with nothing to protect.

Those designs are public and openly licensed: take the decisive trial and run it, no permission needed.

One inherited premise underlies everything here: every growth process in history was stopped from outside, and software is the first system whose growth is not fed through physical pipes. The premise, its named ancestors and the honest dispute over it are drawn in full on the synthesis.

The narrow question that follows

If safety has to be built in, then the immediate question is whether it can be. Does correction scale with the thing it corrects, and by how much? That is not this page. It is the law page, where the condition is stated, the number is derived, and the assumption it rests on is named. This page has done its job when the law page reads as the obvious next thing to ask.

Three statements wait there:

The last two are one boundary in two regimes, and which applies is measurable. None is established; each carries a printed way to die. And the synthesis if you would rather see the pieces put together than the number derived.

If that was a lot

You do not have to decide anything today. Nothing on this page asks you to believe it: the deciding tests are written and dated, and they have not been run. When they are, the result is published here whichever way it goes, including the way that ends the theory. Watching is a real position and it costs nothing. The falsification dashboard is where the verdict lands, and claim status says plainly what is claimed today and what is not.

reads aloud · highlights as it goes · jump to any section