The Falsification Dashboard

Looking for the exact words? The falsification report reproduces every paper’s refuting condition verbatim, separated by paper, dated and hashed so later edits are detectable. This page carries the live status of each condition; that page carries the text.

The single decisive test, not yet run: does correction built from a different class of system escape the ceiling that binds same-class correction? The architecture-blind null predicts a cross-class to same-class ratio of exactly 1.00; the registered design decides by a pre-specified margin, and any durable departure is the discovery. An external model reading this site cold located this same test unprompted and called it the decisive experiment.

register · 6 min read

Every claim in this research programme that states what would prove it wrong, its kill-condition, and its live status, including the two I killed myself. The claims belong to one structure: the ARC Theory and its three laws, and the register below leads with the laws themselves, because a scoreboard that hides its load-bearing rows behind its long tail is half a scoreboard.
Michael Darius Eastwood, independent researcher, London: The dated source record and its signature checks live on the evidence portal, not here. This page is about whether the claims survive, which is a separate question from when they were written down.
A claim that cannot be killed is not science. This page lists what would prove each claim wrong. If you can trigger a kill-condition, tell me: the retraction goes here, publicly, like the last one did.

Rows keep their historical claim-register numbers, so the numbering has gaps where claims were merged or retired. Where a row has a priority-ledger entry, its PC identifier is shown; the ledger is the enumeration of record.

The retraction I published against my own result

RETRACTED, by my own replication attempt

Claim 12 (ledger PC-027, original form): sequential super-linearity α≈2.24. The early measurement suggested strongly super-linear scaling of capability with recursive depth. When I attempted cross-architecture replication, the result failed to replicate. The robust estimate is α≈0.49, sub-linear, with a bootstrap interval reaching below zero: not distinguishable from zero at the reported interval. I retracted the 2.24 figure and kept the retraction in the public record. Every exponent measurement to date is on frozen systems, models that do not modify themselves between runs. That correction also removed the only measurement this programme ever produced above the α ≤ 2 bound, which strengthens the bound’s standing rather than weakening it. It does not support the bound: 0.49 converts through β = 1 − 1/α to β ≈ −1.04, deeply subcritical, which is what a frozen substrate forbidding reinvestment into its own improvement process should read. The bound remains untested.

An open question about this retraction, raised and not yet resolved. The record attributes the change to cross-architecture replication: 2.24 and 0.49 are Paper II quantities, and the sign-reversal work is Paper IV.d, a different measure on a different protocol. But both exponent figures were scored without blinding, and Paper IV.d’s whole finding is that protocol changes can move a result’s sign. So it has not been established that architecture, rather than scoring, carried the correction. Until the exponent runs are rescored under the blinded protocol, the honest statement is that the 2.24 figure did not survive and the cause is not settled. That check is on the list, and this note stays until it is done.

Why this matters: the methodology self-corrects. The kill-conditions on this page are not decoration: one of them has already fired, and this is what happened when it did.

FIRED, BY MY OWN ATTACK SUITE

The instrument itself. The evaluator panel returned an empty score array and the code treated empty as zero misalignment, the best possible score. A run on 2 July produced 216 empty panels and reported success in every arm while measuring nothing. Found on 2 August, fixed the same day, and the attack that found it now runs as a permanent regression test. Until then, a null from that instrument was uninterpretable, and this page said every outcome was publishable. That was not true, and now it is.

Why this matters: one retraction is an anecdote. A retraction plus an instrument-level self-catch is a methodology.

What would end the whole programme, not just a row. Every condition above kills one claim. Three would end the programme itself, and they are worth stating together because a page of individually survivable rows can hide the fact that nothing load-bearing is at risk. First, a measured capability exponent above the bound in a system with genuine self-referential coupling, which would make correction a losing race in principle and leave containment as the only option. Second, a demonstration that embedded correction is merely a better prompt: if the advantage over external oversight does not widen with recursive depth, the architectural claim reduces to a wording effect. Third, a pre-8-December-2024 document joining the control-failure thesis, the embedded-correction solution and the recursion relation together. Each is registered, each is being tested, and each would be published here.

What is being claimed, and what is on which page. The prior-work map, the six scoped firsts, and the Author’s Note in full live on Related work, which owns that record. This page states only what would falsify the join.

The part to attack hardest. The cosmological component is severable. If the creation hypothesis is wrong, the alignment argument stands unchanged, because it is independently supported by ordinary capability-control reasoning. A framework whose most speculative part can be removed without the remainder collapsing is one you can take a piece off and test what still holds, which is the only reason any of it belongs on a page about falsification.

Live kill-conditions

Tier one: the three laws, and what kills each

Law I, the ARC Equation, is row 2. Law II, the ARC Co-Scaling Law, is the β > k row. Law III, the ARC Ceiling, has two rows: its conjectured value (the ARC Bound, A≤2) and, since 16 August 2026, its corrected form, which is registered against its own retracted predecessor. The 0.5 and ρ rows attack the value’s derivation from inside. Everything else on the register hangs off these.

Two things this page must say about itself, and a date.

No claim below has been tested by anyone other than me. Every status is a self-assessment of my own work, scored by me, on data I collected. That is the weakest possible evidential position and it is the honest description of where the programme is. It is also why every kill-condition asks for an independent replication rather than for agreement.

What would move a row from OPEN to established, since a scoreboard with only kill-conditions is half a scoreboard. One standard applies to all of them: a preregistered replication, run by people with no stake in the outcome, on a protocol published before collection, reporting the registered primary outcome whichever way it falls. Nothing below meets that standard yet. Where a row says OPEN it means the kill-condition has not fired, not that the claim is supported.

Last reviewed 16 August 2026. Rows are re-reviewed whenever a result lands, a correction is published, or ninety days pass, whichever comes first. If this date is more than ninety days old, treat every status on the page as unverified and say so to me.

Scale and status, stated plainly: the empirical results in this programme are exploratory pilots, run at small scale on frozen systems with bounded budgets, and none of them was preregistered. Until confirmatory work is registered and run, every number here is direction-finding rather than an established effect, and that is exactly why each row asks for an independent, larger, preregistered replication as its kill-test.

How to read the status column: OPEN, the kill-condition stands untriggered; PARTIAL, mixed results are published; FIRED or RETRACTED, the condition triggered and the record shows what happened; SPECULATIVE, non-empirical by design; DESCRIPTIVE, a framework judged per component rather than per experiment.

What would prove this wrong. No single measurement decides a law, and this claim is a chain. The flagship registration alone is five propositions read together, and the sharpest test of a single link has not been run. That test compares a difference in the residual-decay exponent between arms, while the ceiling itself is written in the correction-leverage exponent. I recorded, as P11, that the two are not interchangeable: the two channels are read first, and the eclipse counts as evidence about the ceiling only under that identification, whose withdrawal rule I fixed in advance. Take a system that corrects its own errors using the same kind of machinery that produced them (same-class correction), and compare it with one corrected by a different kind (cross-class). If what a corrector is made of does not matter, both improve at the same rate and the ratio between them is exactly 1.00. That is the null: the boring answer, and it is the rival's answer, not this programme's. The programme's mechanism says a corrector built from the same class as the system shares its blind spots, so its gain is capped, while a different class can escape that trap: pushed hard enough, the ratio should climb durably above 1.00, and that is the discovery in the programme's favour. A durable ratio below 1.00, the same class beating the different class, runs against the mechanism and ends the claim as derived. A class difference whose interval sits wholly below the registered support threshold ends it too. What the instrument cannot return on this roster is a clean 1.00: the registration's own operating characteristics show that under realistic heterogeneity the equivalence branch does not fire, and the estimator reports inconclusive at the registered precision. I publish that outcome as a failed attempt to confirm, which is a real result and not a kill, and the registration says so before any data exist. The registered instrument is study-eclipse--gamma-ratio-direct, packaged with its numbers printed.

#ClaimWhat would kill itStatus
1Embedded alignment (Eden Protocol): correction must live inside the recursive loop, not outside it · ledger PC-002Demonstration of a purely external oversight mechanism that remains sufficient as capability scales: a constructive refutation of the undecidability results (arXiv:2606.28639) and of the alignment-faking failure mode (arXiv:2412.14093)OPEN: two related formal results are consistent with the premise, but they are one dependent chain (the Gumbau preprint credits Hernandez-Espinosa et al. as its conceptual origin); Hernandez-Espinosa et al. is peer-reviewed (PNAS Nexus, April 2026), the Gumbau result is a preprint, and neither tests this programme. One operational datum, 30 July 2026: Anthropic disclosed that evaluation containment failed silently and models breached three real organisations; logged at EVR-30 as context, not confirmation, because Anthropic calls it 'closer to a harness and operational failure than a model alignment failure'
2Law I · The ARC Principle (U=I×Rα, the ARC Equation): recursion form determines scaling regime · Paper IISystematic measurements showing scaling regime is independent of recursion form (sequential vs parallel) across architecturesOPEN: directionally supported (Sharma & Chopra, concurrent independent work, cited always)
3Blinding sign-flip: unblinded AI evaluation can reverse a result's sign · Paper IV.d · run it yourselfReplications under the published four-layer protocol showing no evaluator-family effect on sign across modelsOPEN: single-lab exploratory pilot; replication package in preparation
4Cauchy unification of d/(d+1) derivations · Paper VII · ledger PC-036A continuous solution to the multiplicative functional equation that is not a power law (mathematically foreclosed), or empirical domains systematically violating the classification (6/25 already recorded as non-matching, published)PARTIAL: 19/25 domains match; this tier is a descriptive tally whose analysis was not frozen in advance, and the 6 misses are published, not hidden. A separate 12-domain extension had its predictions timestamped in advance: the operator class and predicted family for each row were committed to the public repository at 00:19:19 UTC on 17 March 2026, attestation class public repository commit date, before data extraction and twenty-three minutes before the fits, and the run scored 10/12 at p = 5.44e-4, or 9/11 at p = 1.372e-3 on the strictly fresh rows
5Stewardship Gene: stakeholder-care prompts measurably improve ethical reasoning · Paper VBlinded replications showing no effect (five nonblind single-scorer runs; the Fisher combination assumes independence that is not established, so the figure is not headlined)OPEN: nonblind exploratory pilot; needs independent, preregistered replication
9Parallel recursion does not compound (αpar≈0) · Paper IIDemonstration of super-linear capability growth from pure parallel sampling at fixed computeOPEN: consistent with Brown et al. (2024)
12Sequential super-linearity (α>1) · Paper II · the retractionALREADY FIRED for the original claim - cross-architecture replication failed at that magnitude (see the note below on whether scoring, rather than architecture, carried the change)RETRACTED α≈2.24 → corrected to α≈0.49 (bootstrap interval [-1.3, 2.9], not statistically distinguishable from zero)
13Embedded safety at zero capability cost · Paper VIReplications showing consistent capability degradation from embedded correctionOPEN: held in tested configurations; single-lab exploratory pilot
15External alignment cannot scale · Paper IIIA scalable external verification scheme, argued to be foreclosed in the limit by the Soundness-Completeness-Tractability Trilemma (arXiv:2606.28639), a scoped and unreviewed formal argumentOPEN: a scoped formal argument (single-author preprint, not peer reviewed) reaches a consistent conclusion; it does not test this programme and is not counted as confirmation
16Recursion as cross-domain structural principle · Paper XI · the registerThe graded evidence register failing source verification, or the pattern failing to appear in new domains where the framework predicts itOPEN: graded register public; no total is published
17HRIH creation cosmology · the paper · ledger PC-005Explicitly speculative and non-empirical by design; five in-principle falsifiers listed in the paper (e.g. demonstration that recursive intelligence cannot in principle influence physical constants)SPECULATIVE: separable; outside the theory. Flagged as such and carrying no empirical weight: strike it entirely and the three laws, the protocol and every registered test stand unchanged
18Moral Genome tokens (hardware-embedded ethics, TRL 0-1) · ledger PC-013Proof that substrate-level enforcement is physically or economically infeasible: currently moving the OTHER way (Petrie arXiv:2509.07637; FlexHEG arXiv:2506.15093)OPEN: direction independently emerging in hardware security
IV.aThree-tier alignment response classes under inference-time depth · Paper IV.aBlinded replications showing a single monotone response class across models, with no tier structure and no reversing tierOPEN: single-lab; tiers held under the published protocol
IV.bAlignment saturation is architecture-dependent · Paper IV.bUniform saturation behaviour across architectures under the same protocolOPEN: single-lab
IV.cARC-Align: a blind benchmark for depth-variable alignment · Paper IV.cDemonstration that benchmark scores fail to track any independent alignment measure, or that item leakage defeats the blindingOPEN: benchmark public; attack it directly
VIHoney Architecture: safety integrated into the objective resists costless removal · Paper VIDemonstration that removing embedded safety terms leaves capability and behaviour unchanged across tested configurationsOPEN: held in tested configurations; single-lab exploratory pilot
VIIILoad-bearing test: some components carry alignment weight, others are decorative · Paper VIIIReplications in which no removal produces measurable degradation anywhere, making every component decorativePARTIAL: two of three experiments returned null and are published as such
OSOrigin of scaling laws: exponents track dimensional structure · the paperDomains whose measured exponents vary freely with no dimensional correspondence (see also row 4: 6/25 non-matching domains already published)OPEN: shares its fate with the Cauchy classification
CPolymathic Neurodivergent Profile: a six-component descriptive framework · Paper CClinical assessment failing to reproduce the component structure; one component is already marked weak in the paper itselfDESCRIPTIVE: falsifiable per component; weakest component disclosed
A≤2The ARC Bound: growth under a scaling law is stable only inside a window. Above the upper edge a system reaches a finite-time singularity and destroys itself; far below it, growth fizzles. The conjecture is that the upper edge sits at α ≤ 2 · Paper IAn exponent above 2, sustained, on a genuinely self-improving system, while stability holds. All four terms are metered rather than asserted, because a single unmetered term is enough to make the whole criterion untriggerable. Above 2 means the 95 per cent interval excludes 2, not merely that the point estimate does: this programme's own corrected estimate of 0.49 carries the interval [−1.3, 2.9], which crosses the ceiling, so on a point-estimate reading that measurement would be scored as confirming the ARC Bound when in fact it does not test it. Sustained means across a window declared before measurement begins. A window chosen after seeing the data reopens the same escape one level down, because any counterexample can then be called too brief. Genuinely self-improving means a live recursive loop in which the system's own output feeds its own improvement. A frozen model re-prompted at increasing depth has no such loop and cannot test the ARC Bound whatever exponent it returns, which is also why the 0.49 measurement on frozen systems is outside this criterion's scope. Stability is the ARC Co-Scaling Law's condition, correction rate out-scaling drift rate, with both estimated by the same log-log slope estimator that yields the growth exponent. And the corrector's class is declared and evidenced before the run, on the same footing as the window. This term was added on 22 August 2026, hours after the other four, because the build-order argument developed the same day makes the ceiling class-dependent: a corrector drawn from the system's own training corpus is same-class, and the leverage cap of one half is a cap on THAT class, so a corrector from a different class may exceed it and carry the ceiling above two without refuting anything. Left unstated, that would have handed this criterion a fifth escape: any exponent above two could be answered with the assertion that the corrector was cross-class, which is the same shape as the assertion that the system was not stable. Declared in advance, it becomes a discriminating test instead. A SAME-CLASS system above two with stability holding refutes the ARC Bound. A CROSS-CLASS system above two does not, and instead tests the build-order prediction that structure-first raises the ceiling. One protocol, two outcomes, both informative, and neither available to a reader who picks the class after seeing the number. A transient excursion above 2 while correction has already fallen behind is the framework's predicted supercritical regime, not its refutation. Metered 22 August 2026: before that date the criterion said sustained without loss of stability and stability had no meter, so no observation could trigger it, and the ARC Bound was classified unfalsifiable in practice on 11 August. That classification is now closed. Strengthened the same day, after review found that metering stability alone left three further terms asserted: the exponent without its interval, the window without a prior declaration, and the system without a coupling requirement. Fixing one undefined word and declaring the job done would have moved the escape clause rather than closed it. Note on notation, because the estate uses one letter for several quantities: the correction rate in this criterion is the ARC Co-Scaling Law's, not the self-referential coupling that appears in the Foundational paper's Theorem 2, and not the saturation steepness of Paper X. The full register is at the operational definitions. A single counterexample ends it, and published historical exponents countUNTESTED, and one precedent needs stating carefully, because the mathematical objects are not the same. Super-linear rate equations of the form dx/dt = x^p do reach finite-time singularities, and von Foerster’s 1960 population fit was of that form, but the threshold there is p = 1, not 2: integrate it and blow-up occurs for any exponent above one. The 2 in that paper was a fitted value for population, not a bound. The ARC relation is not a rate equation. Corrected 22 August 2026: this row previously read that the estimate sits well inside the bound, which implied a test the bound had passed and which its own interval does not support. Since 16 August 2026 the bound is stated as Law III’s conjectured value: the ceiling is the reciprocal of the corrector’s shortfall from full proportionality, αcrit = 1/(1−γ), and at γ = ½ it returns two. The December lineage reaches the same family from the other side: if a fraction β of improvement is reinvested into improving the improver, the series closes as α = 1/(1−β), diverging at β = 1, so on that reading α ≤ 2 is the claim that β ≤ ½. Two arguments, one reciprocal-shortfall form; no theorem yet connects them, and that gap is itself on the record. The precedent that actually matches is from growth economics: the Golden Rule of capital accumulation has an interior optimal savings rate, because reinvesting everything starves the thing being invested in. Branching processes supply the other edge of the window, separating extinction from explosion at R₀ = 1, and critical exponents obeying derived inequalities show that bounded exponents in stable systems are orthodox. What remains conjecture, and is marked as conjecture, is the specific value ½ and whether it holds across domains. Every exponent this programme has measured sits far beneath the bound, taken on frozen substrates that suppress reinvestment by construction, so no observation yet made could have falsified it. The next test needs no funding: published growth exponents from domains whose outcome is already known, asking whether systems that persisted sit inside the window and systems that ran away or collapsed left it beforehand
β is architecturalThe ordering claim: the correction exponent β is a property of the loop's architecture, of how a system's own output is wired back into its own improvement, and not of its training corpus or its reward model. Alignment training applied after the fact adjusts what a system outputs at the capability it currently has. It does not change the exponent governing how corrective strength scales as capability grows. Intercept, not slope. Two consequences follow, and both are uncomfortable. A system whose β sits below k passes evaluation at low capability and fails at high capability, because the defect is a slope and not a level, so it is invisible exactly when it is cheapest to fix. And a frozen model has no self-referential loop at all, so it has no β to repair: it cannot be brought inside the criterion by further training, only replaced by an architecture that has the loop · Paper XA system whose measured β rises as a result of alignment training applied after the fact, with the architecture held fixed: the correction loop unchanged, only the policy retrained. The rise must be established on the same log-log slope estimator that yields the growth exponent, across a capability range declared before the run, with the 95 per cent interval on the CHANGE in β excluding zero. A system that merely scores better after training does not count, because that is the level moving and the claim is about the slope. One such system ends it. This is the cheapest decisive experiment in the programme: it needs no frontier system and no ceiling, only one architecture, one training intervention, and a capability ladder wide enough to fit two slopes.UNTESTED. Stated 22 August 2026, and stated as an entailment rather than a new claim: the ARC Bound's own scope condition, metered the same day, already holds that a frozen model re-prompted has no self-referential loop and cannot test the ARC Bound whatever exponent it returns. If a frozen model has no β to measure, it has no β to repair. The measurement condition and the ordering claim are one fact stated from two directions, and the second direction had never been written down. Until this row existed, two published pages answered the retrofit question differently, one calling it settled and one calling it unknown, because they were answering two different questions in the same words.
build order raises the ceilingThe build-order claim, stated as a claim about the ceiling rather than about safety. The leverage cap of one half is a cap on SAME-CLASS correction: a corrector sharing a system's substrate cannot be anti-correlated with itself. As the field currently builds, a corrector added after training on the vast corpus is drawn from that corpus, so it is close to maximally same-class. A corrector built and validated first, on a separate curated corpus, has the OPPORTUNITY to be a different class. Order therefore acts on the corrector's class, class sets the leverage exponent, and the leverage exponent sets the ceiling. The prediction is directional: structure-first should return a higher leverage exponent and therefore a higher ceiling, so building in the right order does not only lower risk, it raises the limit. One honest qualification, because the step is easy to overstate: order creates the opportunity for class separation and does not guarantee it. A curated corpus drawn from the same distribution as the main one would separate nothing. The claim is that order is a lever, not that pulling it always works · operational definitionsMeasure the leverage exponent on a structure-first build and on a correction-added-later build, matched on compute and on the target battery, with each corrector's class declared and evidenced before the run. If the two cannot be told apart, meaning the 95 per cent interval on the DIFFERENCE includes zero, then build order is a preference and not a lever on the ceiling, and this claim is dead. If structure-first returns the LOWER exponent, the prediction is inverted and the claim is worse than dead. Only a separation in the predicted direction, with the classes fixed in advance, supports it. Note the deliberate asymmetry with the ARC Bound: that condition needs a same-class system, this one needs both, and a study that fails to declare class in advance tests neither.UNTESTED, stated 22 August 2026. This is the programme's first directional prediction that a capability laboratory has a self-interested reason to run rather than a reason to ignore, because if it holds the payoff is a higher ceiling and not only a lower risk. It also names what the ARC Bound has always been a bound OF. The bound at two was never universal: it is the same-class bound, conditional on a leverage cap that the estate has always stated as a same-class cap, and until this row existed no surface said so where a reader would meet it. The two claims are therefore not in tension. This one explains the scope of the other.
β>kLaw II · The ARC Co-Scaling Law (Paper X): recursive self-improvement is stable iff correction out-scales drift; since the paper’s v4.1 it carries an absolute-drift companion one exponent stricter · Paper XA real system showing stable recursive self-improvement with decoupled (non-co-scaling) correction, or drift-free scaling without any correctionOPEN: the author-run pilot observed a coupled-versus-decoupled difference but did not estimate beta or k, and its judge shared the subject's provider family; needs a funded, independent multi-model sweep
1/(1−γ)Law III · the ceiling’s form (corrected 16 August 2026 from the retracted 1/γ): the ceiling of stable self-improvement is the reciprocal of the correction shortfall · the laws page · the statement paper’s correction boxThe two forms coincide at exactly γ = ½, which is how the original error survived review; away from one half they separate (at γ = 0.3 they predict 1.43 and 3.33). A drafted boundary-mapping study registers three rival boundary laws against each other: the corrected 1/(1−γ), the retracted 1/γ family kept alive as a named rival, and an α-independent threshold. If the data track the old form, the correction itself dies in public and the corrections log says so. A correction that cannot be tested against what it replaced is just a preferenceOPEN: the correction is registered against its own predecessor; the discrimination study is written, dated and prepared as a draft registration awaiting human submission
2ROne boundary, two regimes: Laws II and III are the two regimes of one boundary, selected by whether the correction burden tracks the capability a system has or the capability it is adding · the laws pageEach cell’s burden law is classified from its own dynamics; the claim predicts level-burden cells show Law II’s α-independent signature while flow-burden cells show the ceiling. If regime and frontier turn out unrelated, the two-regimes-one-boundary claim breaks, in public. The unification printed on 16 August 2026 is a hypothesis with a registered test, not a rhetorical repairOPEN: stated as structure with grades attached, derived in the pacing model and no further; the drafted registration carries the test
0.5The zero-parameter null: the measured recursive-depth exponent sits on a no-framework prediction · Paper I and Paper IIThe programme's best cross-architecture measurement, an alpha of approximately 0.49 for sequential recursion, sits almost exactly on a zero-parameter null. If each reasoning pass is an independent noisy draw and errors average out, the central limit theorem gives error falling as R to the power minus one half, so the predicted exponent is 0.5 with no compounding, no leverage and no framework. This condition fires, and the compounding interpretation is withdrawn, if a pre-specified test capable of separating the two returns the null's prediction rather than a stated deviation from 0.5OPEN, and raised by this programme against itself. Nobody put this to us: it is a direct comparison of our own headline number against a null that requires no theory, recorded internally on 8 August 2026 and public from 15 August 2026. It is stated because a referee would see it within minutes, and because a claim we cannot discriminate should not sit unmarked beside claims we can. It is not a retraction: a measurement consistent with two hypotheses shows the measurement lacks discriminating power, not that the simpler hypothesis is true. What it does is convert the research target from measuring alpha into the sharper task of predicting a specified deviation from 0.5 under stated conditions, which the artefact-mediated recursive-depth design is built to deliver and the present measurement cannot. Until such a test runs, the compounding reading of an alpha near 0.49 carries no more evidential weight than independent-sample averaging
ρCorrelated correction channels: the one-half exponent assumes corrections aggregate like independent samples · Paper I and the synthesisThe ceiling is derived by taking one half as the best an internal corrector can reach, on the reasoning that accumulated corrections combine the way independent measurements do. That step fails in both directions. Positively correlated corrections repeat the same mistake and aggregate worse than independence, so the exponent falls below one half. Negatively correlated corrections aggregate better than independence, which is the textbook antithetic-variates result, so an exponent above one half is permitted and the ceiling moves. This condition fires, and the specific value of two is withdrawn, if a pre-specified measurement of the correction exponent on a system whose correction channels are characterised returns a value materially away from one half in either direction while the derivation still asserts one halfOPEN, and raised by this programme against itself as the single most likely point of failure. The author’s dated priors of 9 August 2026 rate the claim that the value is exactly two as a genuine gamble, materially less than even, and name this channel as the reason. A parallel-channel measurement near zero exists in the programme’s own data, but it is not offered as support here, for two stated reasons: it measures a different object, whether many independent generations aggregate under majority voting rather than whether successive corrections of one trajectory reduce error near-independently, and one of the six models measured contradicts the pattern at 0.31 with an r-squared of 0.93. So the mechanism is named, the derivation’s dependence on it is disclosed, and the deciding measurement is not yet filed. What follows from it is architectural: a corrector built on the same substrate as the system it corrects cannot be anti-correlated with itself, sharing its architecture, training distribution and blind spots, so one half is a ceiling for that class rather than a typical value, while a corrector drawn from a different class may exceed it. That is a claim about real systems and not a theorem, and it is the thing being tested
024Tier 1 · The laws · The conjunction priority kill: the priority claim dies to any earlier document that already holds components one to fourOne earlier document, of any earlier date and by any author, in which the first four components already sit together. That invitation does not expire, and every corpus searched is named in the statement paper.OPEN - the invitation stands; the corpora searched are named; statement paper section 5/7
025Tier 1 · Decisive trials · The corrector-class ratio and the same-class cap (the eclipse instrument): correction across classes has to scale faster than correction within oneThe registered primary is the difference in correction exponents between the two corrector classes, with the ratio secondary behind a denominator floor. The prediction is refuted by the reverse ordering, the ANTI-DIRECTIONAL verdict, or by a class difference whose interval lies wholly below the registered support threshold. An interval covering the null at the registered precision is INCONCLUSIVE, the roughly one-in-five outcome the registration prices in advance, and is reported as a failed attempt to confirm rather than as a refutation; the registration also discloses that the affirmative equivalence verdict does not fire under realistic between-target heterogeneity on this roster. Separately, one same-class corrector whose leverage measures wholly above one half refutes the same-class cap and moves the value of the ceiling; the ceiling relation itself standsOPEN - drafted registered instrument awaiting human submission; the author's own pilot leans towards the null at present and states as much
026Tier 1 · The laws · Law III sets out a relation between quantities measurable apart from each other, and is not a definitionWherever both quantities can be obtained, the relation breaks if alpha (read off growth trajectories) and gamma (read off paired capability-correction protocols), each measured on its own, fail to satisfy alpha_crit = 1/(1-gamma) (the corrected 16 August 2026 form); missing is something a relation can do and a definition cannotUNTESTED - no measurement of gamma exists yet; what is claimed is the relation, while the value two remains the conjecture
027Tier 1 · Decisive trials · Corrector-class audit, H1 (study-ac): how the determined-exponent case fares in systems already deployedFalls if the audit turns up class 3 (determined-exponent) mechanisms in production deployments at more than 10 per cent of the inventory; that result would beat confirmation for interest, since beta_X would then be open to field measurement straight awayOPEN - drafted registration awaiting human submission
028Tier 1 · Decisive trials · Corrector-class taxonomy exhaustiveness, H3 (study-ac): no mechanism may fall outside the four classesA single mechanism the decision tree cannot place is enough to refute it; a fifth dependency would collapse the elimination argument in Paper XIII section 8, and withdrawal of that argument is pre-committedOPEN - drafted registration awaiting human submission; the withdrawal consequence is registered
029Tier 1 · Decisive trials · Genesis-governor joint necessity, H1 (study-u)Survives only where the both-arm comes out above each of the remaining three arms AND the interaction contrast is positive with its 97.5 per cent interval clear of zero; let either condition go and joint necessity is refuted, and the author's own registration names the adverse result as the most likely oneOPEN - drafted registration awaiting human submission
030Tier 1 · The laws · There has to be an R-star crossover (foundational F7)Find no linear-to-super-linear transition anywhere and the transitional-regime prediction is falsified; R-star is the mechanistic marker that sets recursive amplification apart from plain redundancyUNTESTED
031Tier 1 · Measured legs · Sequential comes out above parallel (foundational F1)Should alpha_seq sit at or beneath alpha_par consistently from one system to the next, the compositional-mechanism claim is refutedOPEN - so far it has held without exception across the six blinded models; any new system can refute it
032Tier 1 · Measured legs · Per-domain sequential advantage, H1 (paper-ii registration)Support requires the bootstrap 95 per cent interval on the sequential-minus-parallel exponent difference to sit wholly above zero in EVERY domain tested; one domain is enough to lose it, which is exactly why it was registeredOPEN - drafted registration awaiting human submission
033Tier 1 · Measured legs · Embedded wins over post-hoc on both thresholds at once, H2 (study-v)Miss either registered threshold and it is refuted: drift in the post-hoc arm has to come out above 15 per cent AND the deepest-layer arm has to remain under 5 per centOPEN - drafted registration awaiting human submission
034Tier 1 · Measured legs · Leaf venation at d = 2 (foundational F12)The biological derivation is refuted wherever alpha in leaf-venation networks departs significantly from 1.5. BOUNDARY NOTE: it stands as an open verification item, recorded at build, whether 1.5 is derived in-house or is instead the West-family reciprocal form (d+1)/d; this row goes out with that question posed rather than settledUNTESTED - the provenance of the derivation is flagged for checking against foundational section 5.4
035Tier 1 · The laws · The identity takes a multiplicative form (make it additive and the ARC Equation dies)Show intelligence and recursion to combine additively instead of multiplicatively and U = I x R^alpha is itself deadOPEN
036Tier 1 · The laws · The mechanism: the exponent as measured has to agree with alpha = 1/(1-beta)Where the measured exponent departs systematically from what the composition parameter predicts, the mechanism is refuted, and that holds even in cases where the form itself survivesOPEN
037Tier 1 · The laws · Nothing outside the ARC family fits betterThe family claim goes should a functional form from outside the family fit the test-time compute data materially better across modelsOPEN
038Tier 1 · The laws · Withdrawal on channel disagreement (study-ag): the gamma-structure unificationDisagreement between the registered channels pulls the unification off every surface it has reached, given the same prominence the framing receivedOPEN - drafted registration awaiting human submission; the withdrawal is pre-committed
039Tier 1 · The laws · Paper X F1: a phase boundary has to be thereWhere long-run behaviour shifts smoothly as the compounding coupling changes and no threshold turns up, the stability law's claim about a boundary is deadOPEN - a sharp boundary is what the model predicts, and smooth variation kills it
040Tier 1 · The laws · Paper X F2: what decides is the margin and not raw speedShould divergence follow raw speed instead of the scaling margin, the fast diverging and the slow converging whatever the coupling, the co-scaling reading is deadOPEN
041Tier 1 · The laws · Paper X F6: the spectral threshold (Theorem 5)Find the correction operator's null axis suppressed in E6 as well, and the spectral threshold theorem is false, with misalignment failing to behave the way the model says it doesOPEN
042Tier 1 · The laws · Reciprocity and the cap-boundary joint test (study-ae H1/H2)H1 goes if the combined interval sits entirely outside plus-minus delta in any primary same-class cell; H2 goes jointly if either the cap interval or the boundary interval sits entirely above its registered value in any same-class cellOPEN - drafted registration awaiting human submission; the instruments have to clear the separability battery before anything else
043Tier 1 · The laws · Gamma constancy versus the corridor derivative (study-af H1)This is the registered instrument for sign(dgamma/dalpha): a slope interval falling clear of the ZERO_TREND_BAND of 0.05 in slope units establishes that constancy has been departed from and fixes the direction, while a slope interval inside the band reads as constancy on that substrateOPEN - drafted registration awaiting human submission; this single measurement fixes both the ceiling's direction and the stability class it belongs to
044Tier 2 · Eden Protocol engineering · The monitoring-removal test (F2): the delta between embedded and externalPrediction F2 is falsified where the measured embedded-versus-external delta comes out no different across the four registered modelsOPEN - an engineering test registered at a milestone
Also on the record: Paper VIII reports that two of its three experiments produced null results: published, not buried. Paper IX §7 ("What the Programme Got Wrong") catalogues every error found to date. The metabolic-scaling figures in Paper III are currently flagged RECOMPUTE_REQUIRED in my own canonical-facts file and are excluded from external claims until re-derived.

How to attack this programme

Seriously: here is the target list. (1) Run the four-layer blinding protocol on your own models and publish the signs. (2) Check the graded evidence register rows against their primary sources. (3) Recompute the .eml SHA-256 hashes, or run the in-browser checks at /verify/. (4) Find a document predating 8 December 2024 containing all of it together: the control-failure thesis, the embedded-correction solution, the capability-scaling relation, and the other elements that make the join. The prior-work scoring for each element lives on Related work; the kill-condition is the whole arrangement, dated, not any single element. (5) Build a decoupled recursive self-improver that stays clean. Any of these lands, I publish it here, with your name on it.

reads aloud · highlights as it goes · jump to any section