This page is for the reader whose job is to say no. Everything on it is stated in the calibrated register: what is proved, what is measured, what is conjectured, and what would kill each claim. Nothing here asks for belief; it asks for an hour and gives you the shortest route to a verdict.
The condition: a self-improving system stays stably correctable while correction scales at least as fast as capability. Written beta greater than k. Both rates can be read off a running system.
The stronger conjecture is a law before a value: the ceiling of stable self-improvement is the reciprocal of the correction shortfall, one minus the correction exponent. Feed it one assumption, that same-class correctors accumulate corrections the way independent samples do, so the exponent is one half, and it returns two. If the digit dies the law is untouched; if the law dies the digit has nothing to hold it up. Proposed, not measured, and this page treats it that way.
| claim | status |
|---|---|
| Correction outpacing capability implies stable correctability (the beta greater than k theorems) | Mathematical result under stated assumptions; the assumptions are listed with the theorems |
| Beta and k can be estimated from real trajectories | Preliminary: validated on synthetic trajectories; returned not-resolvable on the first compliant real battery for lack of capability range, and says so |
| Internal correction is capped near square-root scaling (the independence premise) | Conjecture, named as the weakest link; on trial in the registered studies |
| The stability ceiling sits at two | Unproven hypothesis; the reciprocal of the conjecture above. The earlier two-independent-ways-must-agree test is withdrawn as vacuous: shared accumulation forces agreement, so a test that cannot fail is not a test. Its replacement is one accumulation law with two consequences (the conversion exponent and the correction ceiling), and the cross-class ratio already on the register carries the unification instead; the honest state of every registration is on the diligence table |
| Current measured exponent approximately 0.49, sub-linear | Measured, published, after a public retraction of the earlier 2.24 |
| Independent replication | Outstanding; the standing invitation is one clone and fourteen checks in about two seconds |
Any of these is published with the same prominence as a confirmation. The claim is retired. That is not a promise invented for this page: one headline claim has already died this way, and its death is on the record.
| result | meaning |
|---|---|
| beta above k | predicted stable correction |
| beta below k | predicted instability |
| stable growth above two in depth | the proposed ceiling is falsified |
| the estimator cannot distinguish them | the measurement programme itself has failed, and says so |
The fastest way to check this programme is not to read it. Clone the public repository and run the independent theorem suite: fourteen checks, no API keys, about two seconds. It re-derives the central results from scratch. It shares no code with the harness, so a common bug cannot hide. If you find an error, the falsification record is where it will be published, with your name on the catch if you want it.
The number-one suspicion about an outsider is rediscovery. So this register answers it before it is asked: each claim beside the earlier literature and exactly what is new here. Yampolskiy’s controllability result is conceded without argument; the d over d plus one exponent has multiple independent derivations; the functional equations are Cauchy’s; the square-root law is the central limit theorem’s. The verified row-by-row table is being assembled from the estate’s related-work verdicts and lands here. Until then, the related-work page carries the concessions in full. If any claim here has been established previously, send the reference: the register updates and credits the catch.
The biography sits at the end of this page. It is not evidence. Who is making these claims, and how to cite him; what a funder should read; what he does when the evidence goes against him.
On 25 August 2026 five systems were given the same prompt, from accounts not signed in or in private windows, with no prior context and nothing to prime them: ChatGPT, DeepSeek, Grok, Perplexity and Claude. The prompt is printed on the standard and it forbids both easy shortcuts by name, dismissing the work for having no institution behind it and believing it for sounding ambitious. It asks three things: is this legitimate research, what is the weakest part, and what would change your answer.
They converged. Not on the theory being right, which none of them said, but on two verdicts and one list.
All five: a legitimate research programme. ChatGPT, on the question put bluntly: “If you are asking is this a crank’s website dressed up as research, no, I don’t think that is a fair description.” DeepSeek: “The author has issued public retractions and corrections, including withdrawing a headline number that favoured the theory. That is the opposite of pseudoscience.”
All five: not established science. ChatGPT again: “Has this established a new general theory of recursive intelligence, scaling, or AI alignment? No. The evidence is nowhere near that level yet.” That is also this programme’s own position, stated on every paper.
All five named the same weaknesses, and they are reproduced here rather than summarised away, because a review you edit is not a review. No independent replication. No peer review. The decisive experiment not yet run. The alpha estimate carrying an interval wide enough to include zero. The cross-domain work using domains the author selected, without a locked external registration. And the one that cuts deepest, from Claude: “A programme this focused on demonstrating good conduct, with a cheap decisive test written and unrun, invites the question of why the test hasn’t run.”
Two things follow, and only one of them is comfortable.
The comfortable one: every model reached the retraction on its own and every model counted it as evidence of method rather than of failure. The habit of publishing your own worst result is legible to a reader who has never heard of you, which is the only kind of reader this page is for.
The uncomfortable one: their criticisms are correct, and none of them can be answered by writing better. No page, no markup and no argument makes an unrun experiment run. That list is not something to rebut; it is the work queue, and it is the same queue the papers already publish. One of those criticisms has already produced a correction: three of the five read this estate as having no preregistration at all, which sent me back to the registered programme page, where a sentence claimed a timestamped public registration for every study while the table under it said private draft forty-six times. The sentence was wrong. It has been corrected, and the models that caught it are credited there.
And one of the five criticisms is narrower than every model assumed. All five faulted the cross-domain work for using domains the author selected without a locked external registration. That is correct about the twenty-five-domain primary cohort, and it stays on the page. It is not correct about the extension. On 17 March 2026 a packet was committed to the public repository the papers cite, locking a twelve-domain candidate list with the operator class and predicted family fixed for every row before any data were extracted, alongside a protocol, a stated primary endpoint, inclusion rules, declared exclusion risks and checksums. It scored ten of twelve. None of the five models found it, because nothing on this site pointed at it, which was this site’s failure and not theirs.
The part of that record worth more than the result: thirteen minutes after committing the successful run, the author downgraded his own claim to a pilot dry run, because extraction had happened before any external timestamp existed. The commit log carries both, at 00:43 and 00:56. The file still contains the note he wrote to himself, that the packet must not be uploaded as a preregistration. A criticism answered by producing a document is ordinary. A criticism answered by a document in which the author penalises his own winning result, unprompted and in public, is the thing this whole page is trying to establish, and it was sitting in a repository the site never linked. The full record.
These are machine assessments and they are not peer review, which is the thing still missing and named as missing throughout this site. A language model can be wrong and can be flattering. What it cannot easily do is invent a retraction that is not there, or miss forty-six rows that contradict a headline. Read the prompt, run it yourself, and see whether you get the same answer or a better one.
The next day, two further zero-history readers were run against the same printed prompt: an unsigned ChatGPT session in a private window, told to read every page and the PDFs and to ignore affiliation in both directions, and a DeepSeek complete read. Same two-door verdict. ChatGPT: “Yes, I would call this legitimate research in the sense of a genuine research programme, but I would not call its central empirical claims established or well-validated science yet.” On weakness: “The weakest part is the inferential bridge from small, researcher-controlled empirical results to broad general laws.” On what would change its answer: “I wouldn’t need a spectacular result. I’d actually prefer a boring one”, and “The decisive question is whether somebody other than the author can take the predictions, freeze the protocol, run the experiment, and get the same answer.” On the retraction: “That is meaningful positive evidence about the research process.” DeepSeek graded the estate “legitimate research programme in progress, not a legitimate established theory” and surfaced the decisive test by name and by number, including the null ratio of 1.00 and the exact status phrase “awaiting human submission”.
One of the readers also independently rederived the programme’s own named weak joint: that ten apparently positive pilots are not ten independent confirmations when a common methodological assumption could generate apparent support across all of them. That is FALS-023, correlated correction channels, published by this programme against itself as its single most likely point of failure. It is the second outside rederivation on record: DeepSeek reached the same objection, in both directions, on 15 August 2026 without access to the register. The objection remains open; study-ah is drafted against it.
What neither reader surfaced, because nothing pointed at it, was the dated-predictions estate, the live notebook and the reception record. That was this site’s failure, not theirs, and it is the failure the machine layer has now been rebuilt to close.
No independent expert has yet put their name to a review of this work. That is stated here before you discover it, and it is the standing offer: read one paper, any paper, and be named on this page as its first independent reader, agreeing or disagreeing. Disagreement in public is worth more to this programme than praise in private; the first name to appear here will appear with its objections intact. Ready to forward: the reader invitation and, for laboratories, the replication invitation.
Every dated prediction the models were assessing is listed in the dated prediction register, with its source artefact, so the claim they graded can be checked line by line rather than taken from this page.
None of the five models was shown the running record, and it bears directly on the post-hoc question they each raised. The laboratory notebook, kept live from 10 March 2026 and running to 206 pages, was written while the experiments ran rather than assembled after them, and it contains the material a tidied account omits: the false positives blinding exposed at version five, and an explicit list of what the author got wrong. Anyone repeating their assessment should read it.