The open problem
How to weigh this
Any site proposing answers to a large problem raises the same question: is this real, or is this what being wrong at scale looks like? That is the right question, and this page does not answer it. It hands you the instruments history uses, and the evidence to point them at. The conclusion is yours to draw.
Start from a fact about the field
AI safety currently has no quantitative stability criterion for whether a self-improving system remains correctable as its capability grows. Control theory builds such criteria wherever it can, from barrier functions to Lyapunov methods, which is exactly what makes the gap conspicuous: the field that owns stability has no criterion for the one system that rewrites itself. It has evaluations, red teams, oversight and interpretability, which are all forms of external verification, and the last two years of results on external verification are summarised on the open problem page: behaviour changes under observation, containment around evaluations has failed quietly at well-resourced laboratories, and a formal strand argues the difficulty is structural. This programme proposes a measurable condition, stated with the conditions under which it fails. Both sentences are checkable. Neither tells you how much the second is worth. That depends entirely on evidence that does not exist yet, and the programme's registers say so.
What history says about founding instruments
Engineering acquired stability criteria in the nineteenth and twentieth centuries: Maxwell wrote On Governors in 1868, Lyapunov's stability theory arrived in 1892, Nyquist's criterion in 1932. A good stability criterion outlives every machine it was derived on. It is a property of the problem, not of any device.
The second pattern is stranger and better documented. The people remembered as founders of quantitative fields were usually anticipated in their ideas. Hartley in 1928 and Nyquist in 1924 substantially anticipated the ideas of information theory; what survived as the founded field was Shannon's operationalisation, the quantities and the theorems. Being reducible to "he made it measurable" is not the consolation prize of founding instruments. Historically, it is their definition.
Neither paragraph is a claim about this programme. They are the calibration you need before reading anyone's register, including this one.
The two shapes problem
One mechanism explaining many things is the shape of the notebooks that founded biology. It is also the shape of a thousand forgotten manuscripts. Conviction does not separate them; both are certain. Scope does not separate them; both are vast. The separator, every time history has run this experiment, has been whether the claims could lose, in writing, before the evidence arrived.
So here is what this estate does about carrying the risky shape. Its own register grades its deepest claims speculative, in the author's hand. Every substantive paper states the result that would refute it, the conditions carry live statuses on a public dashboard, and one has already fired: a headline exponent was retracted and the retraction is published where anyone can find it. None of that makes any claim true. It makes them the kind that can lose. That is the only property the separator has ever tested.
Why this is hard to believe, and why that is correct
Most large unifications die on independent replication. That is the base rate. This programme is not exempt from it, and its registers say so: the empirical results are pilots, small, author-run, awaiting independent replication under locked, third-party-runnable protocols. If you feel resistance reading this site, your calibration is working. The site is built for that reader. It does not ask for belief anywhere. It asks for checking, and it prices the checks.
The checks, in order of cost
- One minute: read the corrections log. What somebody does when the evidence goes against them is the fastest honest measure of everything else.
- Five minutes: read the kill-conditions, including the one marked fired.
- Fifteen minutes: hash the published files in your browser and resolve the timestamps through independent public resolvers; the 8 December 2024 anchor is the one everything else leans on.
- An afternoon: run the published suppression test yourself; the prompts are public and it costs a few pounds. If it fails to replicate, the site asks you to say so.
- The real test: the locked-protocol replications, being drafted for independent third-party submission and run so that anyone may execute them and so that the author cannot mark his own homework. Zero are submitted yet; the design work is what the diligence table records.
The machine-error register
Six AI systems work on this programme. They are wrong regularly. This register lists the errors they made that mattered, and the specific check that caught each one. It exists because “the human checks everything” is a claim, and claims on this site come with receipts. Every row names the error class and the check, never the content of unregistered designs.
| date | the error | what caught it | the fix |
|---|---|---|---|
| 2026-08-09 | A matching heuristic read a sentence that QUOTED a choice in order to reject it, and set exactly that forbidden choice on a live draft registration's metadata. | The operator-ordered one-by-one manual proofread of every registration document. | Hand-reviewed table now outranks every heuristic; a sentinel forces operator-decision fields blank; two regression selftests pin the trap; the correction was verified by reading the live draft back from the API. |
| 2026-08-09 | The bare option word Other substring-matched ordinary prose and wrongly ticked a form choice on about forty-nine draft registrations. | The operator's spot check of a single draft, which triggered a full census. | Minimum-length guard on option matching; a surgical erase that removes only the tool's own fingerprint; verified by census and direct reads. |
| 2026-08-09 | A stamped template carried a broken half-sentence into fifty-one registration files, including both patent-held folders. | The manual proofread's garble sweep. | Repaired estate-wide, mirrored into every source file so the layers cannot drift. |
| 2026-08-08 | Six automated audits of one question in one day reported wrong answers in BOTH directions, undercounting and overcounting the same corpus. | Reading the documents instead of searching them. | Became binding canon with two enforced gates: a question about what documents say is settled by reading them; a search settles nothing. |
| 2026-08-07 | Twenty-six of thirty-two model-supplied book quotations were fabricated. | The character-for-character verbatim law, matched against the published EPUB. | None shipped; every public quotation now carries verbatim verification against the source edition. |
| 2026-03 | The programme's original headline exponent, 2.24, rested on a harness scoring defect and same-family scorer bias. | Cross-architecture replication under the compliant blinded protocol. | Public retraction; the robust 0.49 (interval [-1.3, 2.9]) published in its place; the retraction leads the record rather than hiding in it. |
| 2026-07 | An early registration quoted design-sensitivity figures that had been asserted rather than computed; a later one's primary statistic had zero per cent detection when actually simulated. | The measured-not-asserted rule: sensitivity scripts must run and their output be pasted. | Figures withdrawn and recomputed; the failing primary statistic replaced, with the failure recorded in the registration itself. |
| 2026-08 | An unescaped regex dot made the year 2024 match the retracted value 2.24 seven times in one audit. | The fixed-string counting discipline adopted after review. | Literal matching for numeric audits; the lesson is in the lane's permanent memory. |
| 2026-08-09 | The assistant's own funder-facing copy claimed every quantity was already frozen, an overclaim false for working-tier drafts that deliberately carry open decision markers. | The operator-ordered second-pass review of the assistant's own packet, before publication. | Replaced with the per-registration truth; the checker gets checked. |
| 2026-08-09 | A naive version census produced confident wrong republish verdicts because shared navigation blocks and two different numbering schemes fooled the pattern match. | Reading each page's own title block before reporting, per the read-not-search canon. | Verdicts corrected; only one paper genuinely needed republishing and it was the one the census had not flagged. |
| 2026-08-09 | A generated public dataset carried the design titles of the two patent-held studies toward publication. | The receiving lane's pre-publication redaction check, before merge. | Site copies published counted-but-redacted; the source dataset now redacts at generation so no future copy can re-leak. |
| 2026-08-09 | The phrase already filed, describing draft registrations that have not yet been submitted, reached canon and a live page. | Cross-checking the phrase against the diligence table's zero-submitted truth. | Swept at source and on the site to: written, dated and prepared as draft registrations awaiting human submission; the word filed is reserved until submissions begin. |
| 2026-08-16 | The ceiling law's printed form alpha_crit = 1/gamma ran the wrong way under the paper's own definition of gamma, and its earlier 'derived in the minimal model' status overstated what existed; the defect survived review because 1/gamma and 1/(1-gamma) coincide at exactly one half, the conjectured value. | Two adversarial author-review audits, commissioned and run externally against the statement paper, converged on the defect; direct re-verification against Paper X confirmed it (zero occurrences of alpha_crit). | Corrected to alpha_crit = 1/(1-gamma), the reciprocal of the correction shortfall, derived in three lines with the correction note printed beside the law (statement paper v4.0, 16 August 2026); COR-2026-08-16-001 on the corrections register; a drafted discrimination study tests away from one half so the error class cannot survive measurement again. The reconciliation claimed here on 17 August covered prose, figures and canon but not the executable model; the row below closes that remainder and corrects this claim. |
| 2026-08-21 | The 16 August ceiling correction reached every published sentence and stopped at the code. All eight laws pages displayed the corrected ceiling, animated from it, then decided the reader-facing verdict from the superseded form. Across the widget's own slider range 712 of 1681 positions were wrong, 42.4 per cent, half false-safe and half false-alarm. The error inverts at gamma one half, the sliders' home position and the only value where the two forms agree, so nobody who never dragged the slider could see it. | An external scientific ruling commissioned by the author read the live page and reported that the displayed ceiling and the executed rule disagreed; a lane found it independently the same hour. Confirmed by evaluating both expressions across the page's declared range. | Corrected on all eight pages so the verdict flips where the displayed ceiling is crossed. Guarded by a check that reads the relationship, not the string: it evaluates whatever ceiling and verdict a page declares across that page's own slider range, discovers its pages rather than listing them, reads attributes order-independently because the mirrors alphabetise them, and fails rather than passes when it finds nothing. The ruling's further inference, that the law be suspended, is not adopted: monotonicity under the paper's own definition of gamma, the direction of the ARC Bound and the printed derivation each force 1/(1-gamma) alone. Its separate point, that neither form is derived from a joint model of growth, correction, delay and covariance, stands as an open item. |
The case, as I would honestly assemble it
Every page above hands you instruments and asks you to weigh them yourself. This section, once, assembles the strongest honest form of the positive case, with each pillar carrying its own weight and its own limit, because a reader who has come this far has earned the assembled version rather than the scattered one.
The law is nearly definitional once its terms have meters. That a stability ceiling exists, set by the reciprocal of how well a system corrects itself, is the strong claim, and it is the one with structural support independent of any measured number. The neighbourhood is the default of nature. Square-root accumulation is what independent contributions do wherever they aggregate: the central limit theorem, diffusion, the standard error of every mean ever computed. A corrector at γ = ½ is not an exotic guess; its reciprocal is two. The exact value is materially less than even, by this programme’s own priors. Exact point values are measure-zero events, and the independence assumption is the named weak joint: the parallel-channel pilot showed correction channels can be correlated, and correlated contributions are precisely what bends γ off one half. That failure point was published here first, with studies drafted against it, awaiting human submission.
Two derivations meet at two and part company everywhere else. The registered reading is coincidence, doubly determined; their divergence away from the meeting point is the drafted test, awaiting human submission, never extra evidence. The measurement was blinded four layers deep and published its failures. One clean fit in six models is the honest yield, stated with the five that were not clean, and the resemblance of 0.49 to one half is never counted as support: it is a different exponent family, and the separability battery exists to rule the resemblance out as artefact. The headline number fell under scrutiny and the fall was published. 2.24 to 0.49 (interval [-1.3, 2.9]) is an anti-hype trajectory: corrections here run against the author’s interest, dated and in public. The architecture-blind alternative is pinned to a number with no freedom. The null’s expected ratio, exactly 1.00 under the specified exchangeability model, drawn on the consequences page’s trigger figure and carried on the falsification register, gives the framework something it can lose against.
What this case does not include: a measured γ (never measured on any real system), a filed cross-class ratio (drafted, awaiting human submission), or any claim that the interval settles direction (it covers zero). A hunch whose value depended on being exactly right would be a poor bet; a hunch that generated a falsifiable, registered, cheap measurement eighteen months before the field asked the question has already done its work, whichever number comes back. Weigh that, not the rhetoric.
The honest close
If the modest reading is right, this is a carefully run independent research programme with unusually good records. The checks above will show you exactly that. If the larger reading is right, the same checks are how you would come to know it before anyone said so. The clicks are identical either way, which is not an accident. It is the design.
See the dated record on evidence, and the closest work that came before on related work.