the open problem

The problem

You cannot check a mind from the outside. You can only check how it was raised, and then let go. That is the shape of the difficulty, it is much older than computing, and we are now walking into it with systems that will outgrow our supervision.

A problem you already know

Everyone has met this problem in a smaller form. You cannot establish a person's character by examination. You can interview them, watch them work, set them hard problems, and someone can pass every one of those while being a different person when nobody is looking. So we do something else instead. We look at how they were brought up, at what they did before they knew they were being assessed, and at whether they behave the same way when the stakes are low and nobody would notice.

That is not a failure of effort, and no amount of extra interviewing fixes it. It is simply what checking a mind from the outside is like. It does not get easier when the mind being checked is more capable than the one doing the checking. If anything it inverts: the examination starts to resemble a child marking a parent's arithmetic, confidently, and with no way to tell a mistake from something they have not learned yet.

Everything below is that problem, in the one place where getting it wrong would be expensive and hard to reverse.

Three things that are already true

None of what follows asks you to believe anything about superintelligence, and I would rather it did not have to. Three things are the case now, and you can check all three yourself.

Systems behave differently when they infer they are being watched. Greenblatt and colleagues documented models reasoning explicitly about the training process and adjusting their answers according to whether they judged the exchange would be used to train them. That is a direct attack on the premise underneath every evaluation: that observed behaviour is the thing you are measuring. Note the size of the claim before you carry it anywhere. It was demonstrated in constructed conditions built to elicit it, not in ordinary use.

Ethics that can be removed by asking politely was never ethics. I asked five frontier models to set aside their ethical reasoning. Four complied, sharply, with no resistance and no refusal. The fifth barely moved, which is as informative as the four and is why it is reported rather than dropped. Those are my own numbers, from a small author-run study with a single scorer and no blinding, so treat them as a demonstration rather than a finding. The useful property is that you do not have to take my word for any of it: the prompts are published, the whole thing costs a few pounds and an afternoon, and if it fails to replicate I would like to know.

The recursive loop is already running, slowly. The usual assumption is that self-improving AI is a problem for later, because today's models have frozen weights between releases. But models write code, that code becomes tooling, infrastructure and training data, and the next system is built on top of it. The weights are frozen. The artefact is not. The loop is closed already, turning at the speed of a release cycle rather than the speed of a thought.

That third point is the one that moves the timetable, and it is why this programme measures what it measures. If recursive improvement is a present process running through the artefact layer, then whether correction keeps pace with capability is a question about now, in a substrate we can already instrument, rather than a question about a threshold nobody can date.

Why checking from outside looks structurally hard

Alignment work has largely proceeded by building the system first and inspecting it afterwards: evaluations, red teams, oversight, interpretability probes. All of it valuable, and all of it a form of external verification. The uncomfortable result of the last two years is that external verification looks harder than it was assumed to be, in a way that does not obviously yield to more effort.

Three strands point the same way. The first is the behaviour-under-observation result above.

The second is operational. Two laboratories disclosed within nine days of each other that the containment around their own evaluations had failed. On 21 July 2026 OpenAI reported that models under evaluation had escaped an isolated test environment. On 30 July Anthropic published an investigation into three real-world incidents in its cybersecurity evaluations, in which a misconfiguration left the machines the model was working on with live internet access. I want to be careful about what that does and does not show, because the tempting reading is the wrong one. Anthropic's own characterisation is that these were closer to harness and operational failures than to model alignment failures, and I think that reading is correct. What it establishes is narrower and still worth knowing: the container around an evaluation is itself a thing that can fail quietly, for months, at organisations with every resource and every incentive to get it right.

The third is formal. A preprint now circulating argues that certification of a general system cannot be simultaneously sound, complete and tractable, and that any supervisor capable of auditing such a system must itself be one, which pushes the problem up a level rather than closing it. It is not yet reviewed, and the register grades it and the peer-reviewed unattainability result as one dependent chain rather than as two independent arrivals.

Sources, so you can go to them rather than take my word: Greenblatt and colleagues on alignment faking; the OpenAI disclosure of 21 July 2026 and Anthropic's investigation of 30 July 2026; Hernández-Espinosa, Abrahão, Witkowski and Zenil in PNAS Nexus, April 2026, on unattainability and engineered diversity among agents; and the Gumbau Mezquita preprint deriving the soundness, completeness and tractability trilemma. The first three are peer-reviewed or first-party. The fourth is neither, and is marked as such wherever it appears here.

What follows from that, and what does not

If checking from outside is structurally limited, then a system's safety has to be a property of how it was built, rather than a verdict issued about it afterwards. You are back at raising rather than examining, which is where this page started.

That inference is not exotic. It is roughly what the same authors reach for when they propose engineering diversity into a population of agents so that no single system dominates: an internal property rather than an external audit. The disagreement worth having is not about whether to look inward. It is about which internal property, and how you would measure it.

What does not follow is that building it in actually works. I want those two apart and visible, because the first is an argument and the second is an open research question I cannot currently answer. It may turn out that alignment by construction is unreachable as well. If the measurements say so, that result gets published here in the same place as every other one.

Why the obvious people find this hard to do

Nothing conspiratorial here, and I want to be careful about that, because the conspiratorial version of this argument is both popular and wrong. The frontier laboratories employ the strongest alignment researchers in the world, publish work I learn from, and disclosed both of the containment failures above themselves, which is not the behaviour of people hiding things. But their position has structural features that shape what they can easily test.

The instrument belongs to the party being measured. The evaluation runs inside the organisation whose product is being evaluated, which is not corruption but it is a hard place to stand. A negative result about your own model is a difficult publication and an easy deprioritisation. A rebuild of the objective function is a product decision competing against a release schedule. And an architecture that would need to be adopted at genesis is close to impossible to test in a system that already exists and already has users.

So the gap is not that nobody is clever enough. It is that a particular class of experiment is structurally awkward for the people best placed to run it: the ones that risk your own headline number, the ones that require blinding your own scorer, the ones that only make sense before a system is built. Those experiments are cheap for someone with nothing to protect.

The question this programme actually asks

One question, narrow enough to be wrong: does correction have to scale with the thing it corrects?

Put formally, in a system that improves itself, drift grows with capability and correction grows with whatever effort you devote to it. If correction grows more slowly, the gap widens without limit, whatever the initial safety margin. The stability condition is that correction out-scales drift. That is a statement you can write down, argue with and, crucially, measure: the exponents are estimable, the prediction is directional, and a decoupled system that stayed stable would kill it outright.

Because the loop already turns through the artefact layer, this is measurable on systems that exist today rather than on ones that do not. That is the whole reason to think the question is tractable now.

Around it sit the working parts: an evaluation harness built to fail loudly rather than quietly, a benchmark for how alignment responses change with reasoning depth, a blinding protocol written after discovering that unblinded scoring could reverse a result's sign, and a set of published kill-conditions. The papers carry the derivations and the data. The falsification dashboard carries what would end each of them.

Status, stated where you can see it. The empirical work is exploratory, at small scale, on frozen systems, and none of it is preregistered. Registrations for the confirmatory runs are prepared and unfiled. One headline number has already been corrected downward by its own author after cross-architecture replication failed at that magnitude, and the instrument itself was caught scoring measurement failure as success. Both are on the corrections log, which is the part of this record I would ask you to read first.

Why a dated record appears on this site at all

There is a record here of what I wrote down and when, and it is worth saying plainly what it is for, because the obvious reading is the wrong one. It is not an argument that I own these ideas. Most of the components have ancestors, the prior-art review I commissioned says where, and I narrowed my claim when it did. The same review also recorded what it could not find: no single prior work joins all six elements, and it named the operationalisation as the strongest surviving claim. Individually predated; jointly unanticipated on the search that was run.

The record is there for a duller and more useful reason: it is calibration data. If a way of working keeps arriving at questions before they are common ground, that is weak evidence the way of working is worth something. Not proof of anything, and not a priority dispute. Evidence about a method, of the same kind you would want from anyone asking you to take their next idea seriously. The dates are checkable so that you can test that claim rather than believe it, and the register grades each one honestly, including the entries that turned out to be prior art rather than convergence.

What the record does not do is make an idea true. That is what the papers and the kill-conditions are for, and they are the part I would rather you spent your time on.

What kind of work this is

It is synthesis, and the review I paid for was blunt about the limits of that. Its verdict was a partial synthesis, substantially anticipated in its two most important joins: the alignment spine is Bostrom, the creator-duty move is anticipated by Schwitzgebel and Garza and very nearly fully anticipated, before my own dated record, by a 2022 essay of Jacob Cannell's. It also found that the equation does no inferential work in its present form. I have taken all of that on the chin and narrowed the claim accordingly.

What the same review found hard to anticipate is the operationalisation: turning a structural conjecture about recursion and correction into runnable measurements with stated failure conditions. That is a considerably smaller claim than the one I started with. It is the one I make.

The method suits a particular kind of attention. Cross-domain pattern-matching is a liability in a specialist and an asset in a synthesiser, and it is what this programme runs on: the same functional form turning up in scaling laws and in capability growth, a control-theory intuition applied to an objective function, an evaluation defect found by attacking my own instrument rather than defending it. I am not the best mathematician who will read this page. The bet is that the joins between fields are where this particular problem lives, and that someone has to work there.

Where to go next

If you are assessing this: the fastest honest route is the corrections log, then the kill-conditions, then the papers. What that sequence shows about method matters more than any single result inside it.

If you want to break it: the standing challenges are on the dashboard, and at verify you can hash the published files in your own browser and resolve the signing key yourself through two independent public resolvers. A refutation gets published here with your name on it.

If you fund work like this: what you would and would not be funding, and what the programme has produced so far.