Research · dated predictions
What was predicted, when it was fixed, and how to check it
This page exists because the record was being read backwards. A programme that publishes its own retractions was being summarised as one that predicted nothing, and the reason was procedural: no registry form had been filed. So here is every prediction this programme made before its evidence existed, in the order it was made, with the artefact that fixes each date and the means to check it yourself.
Read this before the table, because it concedes the thing you were about to raise
Priority over any individual prediction is not claimed here, and it does not need to be. Almost every line below has been anticipated by somebody, somewhere, in some form. Recursive self-improvement has been discussed for decades. Value lock-in has a literature. Substrate-independent consciousness is an old argument. If your objection is that a particular row is not original, you are almost certainly right, and it costs this page nothing.
The claim is the conjunction. These predictions are not a collected list of other people’s ideas; they are consequences drawn from a single stated principle, published together, before the tests, with thresholds and deadlines attached. A framework earns its keep by generating a set that arrives together and can fail together. Take any row on its own and it has precedent. Take the set, and the question becomes whether one principle should have been able to produce it.
Convergence is therefore the honest word for what follows, and it is used deliberately in place of priority. Where later work has matched a prediction, that is recorded as convergence and dated on both sides, never as a claim that anybody was scooped.
The chronology
8 December 2024 · sealed record · fixed
The ARC Principle written down and sent, in a record carrying its sender seal: that intelligence and recursion are fundamental rather than emergent, and that recursive depth governs how capability scales. This predates the book, every paper, and every experiment in the programme.
How to check: The seal is server-generated and is not the author’s to move. Paper II reproduces the transport headers in its provenance note.
30 April 2025 · manuscript · fixed
A dated manuscript standing between the sealed record and the print edition, closing the gap between them so the chain has no unexplained interval. It matters less for what it adds than for what it removes: the objection that nothing exists between December 2024 and January 2026.
How to check: the date appears on the priority ledger and the evidence page, which were published before this register existed.
2 January 2026 · Infinite Architects, in print, ISBN 978-1-80605-620-0 · open, deadlines pending The book is now free to read in full, so the printed predictions can be checked without buying anything.
Nine numbered predictions across two appendices, each with a threshold, a deadline or a falsifier. Appendix A section A.3 carries four, and is followed by A.4, headed Falsification Criteria, stating in advance what result would sink the framework. Appendix F carries five:
- Meta-cognitive emergence in at least one system by 2028, recognisable by self-directed architectural improvement its designers did not anticipate.
- Alignment drift above 15 per cent within eighteen months without embedded constraint, and below 5 per cent with it.
- Recursive capability gains above 300 per cent on standardised benchmarks in a single training cycle by 2029.
- Value stability under adversarial conditions where software-only alignment fails.
- Convergent consciousness signatures across biological and artificial systems.
It closes: “These predictions are my wager. If they fail, the framework is wrong or incomplete. If they succeed, something important has been glimpsed. Time will judge.”
How to check: Both appendices are reproduced in full here, word for word from the printed edition, so the predictions can be read without buying anything. A printed book cannot be edited by anyone, ever, which is a stronger guarantee againstpost-hoc revision than any registry offers.
13 February 2026 · Paper III, priority record · five numbered falsifiers, open
Paper III carries a section headed Priority record: dated predictions, whose contents are, in its own words, “publicly dated for reference as of 13 February 2026”. Each carries a numbered falsification criterion, so each can be killed individually:
- F11, the ARC Bound. No classical sequential recursive system sustains an exponent above 2.0. Refuted if any system demonstrates above 2.3 with a 95 per cent interval excluding 2.0.
- F12, leaf venation. Quasi-two-dimensional biological transport networks will show an exponent of two thirds, and not the three quarters of three-dimensional vascular systems. A prediction about leaves, from a theory about intelligence, which is the kind of prediction a framework cannot make by accident.
- F7, the scaling crossover. A measurable depth exists at which behaviour switches from base-dominated to recursion-dominated, computable in advance.
- F4, the coupling identity. A measured coupling parameter predicts the scaling exponent to within plus or minus 0.3.
- F10, the composition operator. Measuring how two recursive blocks compose predicts the full functional form before the full curve is measured.
The same section records a prediction that did not come out cleanly. External alignment constraints were predicted to show an exponent near zero. Under blind evaluation across six models the median was near zero, but the result proved architecture-dependent: three models scaled positively, two were flat, and one scaled negatively. The paper reports that split rather than the median alone, which is the only reason the prediction counts for anything.
How to check: open Paper III and search for “priority record”. Every criterion number above appears in the text.
17 March 2026 · public repository, preregistration packet · run, 10/12
Twelve domains, each with its operator class and predicted family fixed before data extraction, with a written protocol, a stated primary endpoint, inclusion rules, declared exclusion risks and checksums. Result: ten of twelve, p = 5.44e-4. Thirteen minutes later the author downgraded his own successful claim to a pilot dry run, because extraction had preceded any external timestamp.
How to check: Clone the repository and read the commit log for 17 March 2026. Both the result and the self-demotion are in it: e5c22dd at 00:43, f5e06a8 at 00:47, and c853277 at 00:56. The full account.
20 March 2026 · Paper II, version 13 · FAILED, published
A prediction of this programme failed, and this is it. The single-model exponent of 2.24 did not replicate across architectures. It was retracted in public and the cross-architecture estimate revised to 0.49, with an interval wide enough to include zero. The retraction is on the corrections log and on the falsification dashboard, not buried in a footnote.
How to check: The corrections log and the falsification register.
Standing wagers: predicted, and not yet proven
These are open. Each has a stated way to lose, and none of them has been settled. They are listed so that a reader can hold this programme to them later, which is the only reason to publish a prediction at all.
| wager | fixed | how it loses |
|---|---|---|
| Meta-cognitive emergence in at least one system | 2 Jan 2026, print | nothing qualifying by the end of 2028 |
| Alignment drift above 15 per cent without embedded constraint, below 5 per cent with | 2 Jan 2026, print | the two arms fail to separate at those thresholds. Registered as study-v, which cites the printed figures |
| Recursive capability gains above 300 per cent in a single training cycle | 2 Jan 2026, print | no such gain by the end of 2029 |
| Convergent consciousness signatures across biological and artificial systems | 2 Jan 2026, print | the weakest of the nine, and held to the strictest standard here precisely because it is the weakest |
| F11, no classical sequential system sustains an exponent above 2.0 | 13 Feb 2026 | any system above 2.3 with a 95 per cent interval excluding 2.0 |
| F12, leaf venation at two thirds rather than three quarters | 13 Feb 2026 | measured exponent excludes two thirds |
| F4 coupling identity, F7 crossover depth, F10 composition operator | 13 Feb 2026 | each carries its own numbered criterion in Paper III |
| The eighty-domain extension | drafted, not yet run | never fitted. The predictions are written and the data have not been touched, which is the cleanest state a prediction can be in |
One further item is held as a conjecture and not a prediction, and the distinction is deliberate: a same-substrate corrector ceiling at one half. It has no settled derivation, the author’s own stated prior is materially below even, and it is recorded here so that it cannot later be presented as something firmer than it was.
Where each dated prediction stands now
This section renders from the per-prediction outcome register at every build, so the grades below can never drift from the register that holds them. The grading rule is stated in the register itself: internal reruns are never confirmations, retracted rows stay prominent, and pending rows name their deciding instrument. As of 2026-08-27, 0 of 10 rows are graded fulfilled; the strongest grade any row carries today is exploratory support, and that conservatism is deliberate.
10 dated predictions: 3 pending · 2 open, undecided · 2 exploratory support · 1 mixed, recorded · 1 convergent timing only · 1 failed and retracted
- pending Book Appendix F Prediction 2: alignment drift above 15 per cent without embedded constraint and below 5 per cent with it, within 18 months (in print 2 January 2026).
- open, undecided Book Appendix A.3: AI development scales quadratically rather than linearly with recursive depth (in print 2 January 2026; framework from 8 December 2024). The early alpha approximately 2.24 estimate was publicly retracted; the blinded replacement near 0.49 carries an interval wide enough to include zero and two, so it is consistent with the proposed ceiling and does not decide it. Neither fulfilled nor failed: open, with the deciding instruments drafted.
- exploratory support The 12-domain locked-manifest extension: per-row scaling families predicted before fitting (manifest dated 17 March 2026) (17 March 2026, locked manifest). Run 17 March 2026: 10/12 family matches, binomial p = 5.4e-4; the author demoted the run to a pilot dry run within thirteen minutes, in public commit history; graded exploratory support, never confirmation.
- exploratory support The 50-domain tiered suite: per-domain predicted families before fits (16 March 2026 catalogue). Empirical tier 19/25 under the original candidate set, 18/25 under the corrected seven-model rerun (11 August 2026, disclosed in full); published-direct 13/13; provisional 3/6; analytic 6/6 reported as definitional; tiers never blended; author-selected domains, fitting protocol not registry-filed at the time.
- mixed, recorded Temporal out-of-sample: families fitted on old data predict the functional form of new data (7 domains) (March 2026 design). 4 confirmed, 2 partial, 1 inconsistent; the inconsistent domain is reported as prominently as the confirmations.
- pending The fourth cell: the logarithmic family the corpus has never fitted, on never-fitted domains (registered prospective design).
- pending Acoustic time crystal: super-linear scaling of temporal stability with oscillation cycle count in the transient growth regime (alpha above 1) (10 February 2026, 18:29, self-copied email to the subject lab, the kill staked in the recipient's favour).
- convergent timing only Quantum error correction scaling (the Willow timing) (8 December 2024 manuscript; the Willow team's preprint public since 24 August 2024). Graded convergent timing, never prediction, by the estate's standing law: the team's preprint predates the manuscript record.
- failed and retracted The early headline exponent: alpha approximately 2.24 (sequential scaling) (early programme measurement). Retracted in public after blinded cross-architecture analysis; replaced by approximately 0.49 with a wide interval. Kept prominent as honesty evidence: the record of the failure is part of the prediction record.
- open, undecided Paper III's dated predictions (13 February 2026: F4, F7, F10 to F12), including the leaf-venation d/(d+1) exponent (13 February 2026, Paper III Priority Record section). Each carries its falsification criterion in the register; the venation exponent's 10 March correction is itself dated and recorded; statuses live per-condition in the falsification register.
The one document that cannot be written afterwards
Everything above is a prediction with a date on it, and a determined sceptic can always ask whether the date is the whole story. There is one class of document that answers that question by its nature rather than by assertion, because it cannot be produced after the fact without being a forgery: a laboratory notebook kept while the work is happening.
The ARC alignment scaling report is that notebook. It was commenced on 10 March 2026 and carries the version stamp live, updated in real time. It runs to 63,667 words across 73 sections and 346 subsections, and the PDF is 206 pages. It was written as the experiments ran, not assembled once they had finished.
The reason it belongs on this page is not its length. It is that a running record keeps the author honest in a way a finished paper cannot, because the author does not yet know how the story ends. So it contains the things a tidied account leaves out: the version that was wrong, the false positives, the day the method changed and why.
It records, for instance, that blinding was added at version five and that doing so exposed false positives in version four. That is a methodological failure written down by the person who made it, at the moment he found it, in a document he was publishing as he went.
And it keeps a list of what he got wrong. Verbatim, from the report’s own self-assessment:
Wrong: Trusting v1’s α = 2.24 for even a moment. Not adding blinding from the start. Publishing the v4 “Computed Alignment” taxonomy before blinding had been tested. Assuming max_tokens would control reasoning depth in v1. The token bug in Paper II that captured reasoning_tokens instead of total_tokens.
It then does the thing that costs the most and is worth the most: it declines to grade itself. “The errors were, in retrospect, predictable. The first version of any experiment is always wrong in ways the experimenter cannot anticipate. The only question is how quickly the errors are detected and corrected. In this case: 48 hours from first code to a 2,549-entry blinded dataset with the v4 false positives identified and corrected. I do not know if that is fast or slow by the standards of alignment research.”
That last sentence is the register this whole page is trying to hold. A person fitting a story afterwards does not write down that he trusted a number he later retracted, and he certainly does not end the confession by admitting he cannot tell whether his own recovery was impressive. Read it, and check the predictions above against the record of them being tested.
And the notebook’s mid-run states are independently timestamped. This website’s own public commit history captured the report twice while the experiments were running: on 13 March 2026 (commit ef74315a1, 57,233 words, ending at Chapter 51) and again in the early hours of 17 March (commit 17221714e, 63,970 words, with the results recorded). The 13 March snapshot contains the framework’s numeric predictions, including the zero-parameter exponent law with its worked values, and not one mention of the fifty-domain suite or the twelve-domain extension, because they had not yet run. Their result files carry 16 and 17 March timestamps of their own, and the extension’s outcome entered a second public repository in commits between 00:43 and 00:56 on 17 March. Prediction snapshot, result artefacts, outcome snapshot: three legs, each timestamped by a different mechanism, all checkable by anyone with git. The 13 March snapshot also already says the widest claims “should be advanced as a high-value proposal with partial support, not yet as a proven law”: the self-limitation was committed before the favourable results arrived, not after.
Be precise about the boundary, as ever: the sandwich proves external ordering for the 14 to 17 March window only. The earlier material, including the pre-run statement of what results would constitute genuine evidence, is dated inside the live document and corroborated by the run files, but the report’s first git capture is the 13 March one, so that window rests on in-document ordering. The distinction is stated here rather than blurred, and the exact commands to check both snapshots are in the machine register under report_git_snapshot_ordering.
It is also deposited on OSF, at osf.io/j6mkb, which matters for a reason that has nothing to do with convenience: that copy sits on infrastructure this programme does not own and cannot edit. If this site were taken down tomorrow, or quietly revised tonight, the notebook would still be there, byte for byte, with its deposit date attached. A record you can only read where its author keeps it is a weaker record.
What the timestamps prove, and what they do not
The register itself is anchored: hash manifests of the published artefacts carry OpenTimestamps proofs, so the exact bytes of the record can be checked against a public chain. Be precise about what that buys. Those anchors are dated August 2026. They prove the record was frozen by then and has not been edited since. They do not back-date anything.
The dates on the predictions themselves come from elsewhere and each is separately checkable: a sender seal on the December 2024 record, a manuscript dated 30 April 2025, an ISBN and a print run on 2 January 2026, a public commit history for February and March 2026. Anchoring and dating are two different claims, and running them together would be the sort of shortcut this page exists to avoid.
Why the failed row is on this page
A prediction register that lists only the surviving predictions is a marketing document. The 2.24 retraction is the most load-bearing row here precisely because it went against the framework, was found by the author, and was published by the author. Every model asked to assess this work cold reached that retraction on its own and counted it as evidence of method. It is the reason the rest of the table is worth reading.
The same applies to what is missing. Where a test lacked an analysis plan frozen in advance, the paper reporting it says so in its own abstract. Those admissions are not edited down as this record gets stronger.