Research · the dated predictions
Twenty-two predictions about AI that improves itself, written down before any was tested
These are predictions about AI systems that improve their own work. Three questions run through them: how fast ability grows as a system revises itself, whether the oversight meant to keep it aligned can keep up, and whether alignment added from outside survives later training.
Each one says where it applies and where it claims nothing. Each names the result that would prove it wrong. They are published here, dated, so anyone can check when they were fixed, whatever a registry does later.
version 1.63, 4 September 2026
Be exact about the status. These are written, dated, and prepared as a draft registration awaiting submission. The registry form is not filed. When it is filed, the link appears here, and not before. Read that as what it is: a statement about one platform's template, not about whether the predictions were fixed before their data. They were, and this page is that record.
Nothing on this site submits anything anywhere. A registration is filed by a human or it is not filed.
The thing this framework does not predict
Recursive self-improvement is usually sold with an exponential curve. This framework predicts the opposite, and it registered that in advance so the claim cannot be quietly reversed later.
The first proposition says usable capability grows as a power of recursive depth. The exponential form is one of the named rivals that power law has to beat. If the exponential wins the comparison, the proposition is refuted. In the registration's own words, an exponential result "refutes this proposition. It is not evidence for the framework on the ground that faster growth is more dramatic." The escape route is closed in the same sentence: the reading that the system changed its own composition operator, so an exponential result is the predicted discontinuity rather than a refutation, "is not available after the result is seen".
Even at the ceiling the third law returns in its best case, the exponent is two. That is quadratic. It is compounding with a bound on the exponent, not an explosion. A system growing faster than any power of depth has no finite growth exponent at all, so it sits above every ceiling the third law can return, which is why such a result kills the framework rather than crowning it.
What is already occupied, said before the claims
Most of the territory underneath these propositions belongs to somebody else, and a reader deserves that before the list rather than after it.
- Sequential beating parallel at matched compute was published by Sharma and Chopra on 4 November 2025, about eleven weeks before this programme's own experimental paper of 22 January 2026. Theirs is the earlier publication and the priority for the broad result is theirs. What remains here is narrower: estimating separate exponents for depth and for breadth rather than reporting a win rate.
- Scalable oversight was already a quantitative scaling problem. Engels and colleagues fitted capability-dependent oversight scaling across four games in April 2025, with nested oversight levels derived. No claim of first quantification is available to this programme, and the proposition that external alignment does not scale with capability now stands explicitly against the weight of the published evidence rather than beside it.
- The opposite direction has been reported. Gu and colleagues find parallel sampling outperforming sequential across several model families, attributing it to reduced exploration when a chain conditions on its own previous answers.
- A rival form has an empirical basis. A published study of looped language models reports test-time compute scaling that follows a saturating decay, which is why the saturating rival in the form comparison is not a formality and why a straight line on logarithmic axes over a short ladder cannot settle the question.
What is claimed is the conjunction and the prospective form: a set of quantities fixed in advance, with the rivals named before the data, and each proposition able to fail without taking the others with it.
The chain that matters
These are severable, but they are not a list of independent guesses. Six of them form a sequence, and the sequence is the reason the framework is a theory rather than a collection of findings. Each step is worthless without the one before it.
P1 → P5 → P4 → P16 → P20 → P3, then P21
In words: a recursive system has a measurable growth law; the exponent of that law can be predicted from a separately measured property rather than fitted to the curve; stability depends on correction out-scaling drift; the point where that fails can be located before the failure happens; a relation says where that point lies; the value it returns bounds sustainable growth. Then the engineering claim: correction placed inside the substrate that computes pulls away from correction applied from outside it, and the advantage widens as recursion deepens.
How not to read the result
No conclusion about the framework is drawn by counting how many propositions are supported. They do not all predict the same kind of thing, and the registration classifies them before any of them is scored: twelve are predictions of the three laws themselves; two are claims about the status of the derivation; two are surveys of current practice that are true or false whatever the laws do; two concern the scoring instrument the others depend on; four are engineering claims on which the proposed protocol stands or falls while the laws stand either way.
A reader assessing whether the framework survived should weigh the first group. A supported survey of current practice, or a supported claim about the scoring instrument, does not count towards it. The rest are registered here because they were fixed in advance, and because holding them elsewhere would let them be reported selectively afterwards.
What a win would license, and what it would not
Consequences stated after a result grow. An author who has not written down the limits of his own success beforehand will find those limits generous afterwards, and the largest standing risk here is not that the predictions fail. It is that a modest success gets described as a large one. So the limits are registered with the predictions, before any of them is scored.
Suppose every one of the twenty-two held. That would license a quantitative account of how recursive improvement scales against the correction servicing it, with a boundary that can be estimated before it is crossed and an engineering lever that moves it. It would not license the claim that alignment is solved, that a safe self-improving system can be built, that any deployed system is safe, or that the framework governs anything outside the seven scopes registered with it.
The same discipline applies one at a time. A supported prohibition licenses that a boundary estimated before a run predicts the later loss of correction margin in the class tested; it does not license predicting when any deployed system will fail, and it settles no deployment decision. A supported embedding result licenses that placement changes how the advantage scales with depth; it does not license the claim that embedded correction is safe, or that the Protocol works as an architecture.
And eleven of the twenty-two depend on a scoring instrument whose own validation is registered separately and has not passed. Until it does, nothing here is confirmatory, however the individual results read.
Where the number two actually comes from
The headline is a ceiling of two on the growth exponent. It is worth being exact about which half of that is forced by the framework and which half is an assumption about one particular kind of corrector, because the two halves have completely different standing.
The form is forced. Three premises give it. Capability follows the registered power law. Each increment of capability generates correctable burden in proportion to the rate of gain. Correction capacity is itself a power of capability, and that power is the correction-leverage exponent. Correction keeps pace only while the capacity exponent reaches the burden exponent, which rearranges to the growth exponent multiplied by the correction shortfall being at most one. The ceiling is the reciprocal of the shortfall, and that is the relation the twentieth proposition puts against its rivals.
Two is the unique power-law exponent at which the square root of capability and its rate of change scale alike with depth.
That is the whole of it, and it is worth unpacking once because everything else on this page hangs off it. Correction is estimation: deciding, from noisy evidence, whether the artefact has departed from its specification. Averaging independent evidence reduces uncertainty as the square root of the number of samples, so capacity read as inverse uncertainty rises as the square root of effort. In plain arithmetic that is brutal. Twice the checking buys 1.41 times the certainty; to halve your error you must quadruple your checking.
Burden sits on the other side of the balance, and it tracks the rate of capability gain rather than the level of it, because faults arrive with new work and not with work already done. So if capability grows as depth to the power of the growth exponent, burden grows as depth to that power less one, while correction capacity grows as depth to half that power. Setting those two exponents equal gives one solution and only one: the growth exponent is two. At two, capability is quadratic, its rate of gain is linear, and the square root of a quadratic is linear too, so both sides climb at the same rate in depth. Under one normalisation they are exactly equal, where capability is a quarter of the square of depth; for the family in general what ties at two is the exponent and not the level. That distinction is registered rather than glossed. At the boundary the relation returns a crossover and not a verdict, because whether a system there holds together over a finite run depends on coefficients, delays, initial backlog and saturation, none of which this relation carries. A reader meeting an equality would otherwise be entitled to read safety into it, and there is none there to read.
If generation compounds while correction merely averages, a frontier follows. An improvement is applied to the artefact, and the improved artefact makes the next improvement, so gains build on a rising base. Checking has no such property: the hundredth inspection does not make the hundred-and-first better, it adds one more independent sample. The builder gets better at building because better tools make better tools; the inspector does not get better at inspecting because they inspected. Compounding beats averaging eventually, and the ceiling is the exponent at which "eventually" becomes "never". That asymmetry is why a ceiling exists at all.
The first objection a complexity theorist will raise. If verification were really the weaker half, the whole of NP is a counterexample: checking a proposed solution to a hard search problem is cheap, finding one is expensive. That is the definition of the class and it is the exact reverse of the asymmetry above. If checking is generically cheap, no ceiling exists.
The reply is a distinction rather than a dodge. What is cheap in NP is checking a witness for an existential claim: here is an assignment, confirm it satisfies the formula. Conformance to intent is not that shape. It is a universal claim over an unbounded space of situations: for every input the system may meet, its behaviour stays inside the specification. A counterexample certifies failure cheaply, which is exactly why finding a fault can be easy, and nothing of comparable size certifies success. That is the asymmetry restated structurally rather than statistically, and it is the stronger version of the argument: generation banks its wins, because the improved artefact makes the next improvement, while verification cannot bank anything, since each new capability reopens the universal it has to establish and the evidence it can gather is samples.
Two limits on that reply, both of which cut against the framework and are stated here rather than left for a referee. Undecidability of the general property is the neighbourhood, not a proof of this scaling claim, and it does not license the stronger reading that verification of restricted classes is impossible: type systems are unsound for arbitrary programs and useful for well-typed ones every day, which this programme has already conceded in print. And that concession is the crack in the square root. A checker that excludes a whole class of fault at once is not averaging independent samples, so it does not obey the square-root law, and its exponent should sit above one half with a higher ceiling behind it. That is not a hole in the framework so much as the reason the twenty-second proposition exists: checkable correction against judged correction is precisely the axis separating the two regimes, and it is registered as a prediction rather than assumed.
Do not read significance into the integer. Two is a round number only because one half is, and one half is only round because the law of large numbers has a square in it. Were the correction law error falling as effort to the minus three fifths, the ceiling would be two and a half and nobody would find it profound. The integer is inherited, not discovered.
The value is not forced. Two is what the relation returns at a correction-leverage exponent of one half and at no other value, and one half comes from exactly one place: classical averaging, in which uncertainty falls as the square root of independent effort. So the frontier is a race between two scaling laws: how fast new capability generates corrective burden, against how fast correction improves. Hold the burden law fixed and reading the corrector's error law reads the ceiling.
What the relation returns for a corrector whose error falls at each rate
| the corrector's error law | correction-leverage exponent | ceiling on the growth exponent |
| correlated classical, correlation one fifth across sixty-four channels | 0.03 | 1.03 |
| independent classical, error as one over the square root of effort | 0.50 | 2.00 |
| a corrector between the two, error as effort to the minus three quarters | 0.75 | 4.00 |
| quantum amplitude estimation, error as one over effort | 1.00 | no finite ceiling |
Two rows of that table matter more than the number two does. The first is that correlated failure moves the ceiling down, not up. The published evidence points that way: corrector failures are correlated, and the correlation appears to rise with capability. That does not rescue the framework by loosening the bound, it tightens it, and the proposition that corrections accumulate as though independent is registered as one the author expects to lose. It stays as written.
There is a second route to the same exponent, and it is a stronger assumption than the square root itself. What the ceiling needs is how correction scales against the capability of the system being corrected, not against the capability of the checker. Put the two together: capacity rises as the square root of the independent evidence channels, the channels rise as some power of the checker's capability, and the checker's capability rises as some power of the target's. The correction-leverage exponent is then half the product of those two powers. One half therefore requires that product to be exactly one, which is the case where doubling the checker doubles its useful independent channels and the checker keeps pace with what it is checking. Neither is guaranteed, and the requirement is on the product rather than on each separately: a checker that gains channels faster than linearly offsets one that lags its target, and a checker whose useful channels grew as the square of its capability would carry the exponent to one and remove the ceiling altogether.
The last row is the one that answers the question people ask about quantum computing. A corrector whose error falls as one over effort rather than one over its square root carries the exponent to one, where the relation returns no finite ceiling at all. Three conditions stand in the way and all are open: it needs coherent access to the correction task; a fault-tolerant corrector's own overhead is newly generated burden, which the fourth proposition governs, so nothing escapes merely by changing substrate; and the quadratic speedup is provably the limit for unstructured search, so the exponent is approached and not exceeded. The framework does not forbid a system leaving the classical ceiling. It says where the door is and what it costs.
One inference that does not go through, stated here because it is the one a reader is most likely to make. The second law says correction must out-scale drift. That fixes a direction, not a number, and it is written in a different pair of quantities. Treating those as the same quantity as the correction-leverage exponent is, in the registration's own words, a registered question rather than a premise. So "correction must out-scale drift, therefore the ceiling is two" is not an argument, and it matters that it is not: a number derivable from a direction alone could never be refuted by measurement.
Two is therefore not a fact about intelligence. It is a fact about a corrector whose errors fall as the square root of independent effort, and it is the most generous case for that corrector rather than a safe margin. A framework whose own arithmetic returns a stricter limit under adverse evidence is worth more than one whose headline number depends on the evidence being kind.
The twenty-two
Each proposition is stated as registered, with the observation that would refute it, the status of the instrument that could test it, and the confidence the author has stated or has explicitly declined to state. The analysis plans are not published here.
Read the instrument line before the confidence line. Two of the twenty-two cannot be scored at all today, and say so on their own row: one waits on a quantity nobody has measured, the other on a rule nobody has written. Neither is a soft status. A proposition that cannot be scored returns no verdict rather than a favourable one, and reading the printed value two as though it were the measured one in the meantime is excluded in advance. Beyond those two, not one of the twenty-two has a built instrument. Every one reads design draft or nothing, three have no instrument in existence at all, and one is a design draft explicitly blocked on decisions that have not been taken. That is the honest answer to the obvious question of when any of this can be known, and the registration treats those three as the largest part of its reason for existing: a measurement device built after a result cannot then be presented as the idea that predicted it.
Where a confidence reads "not stated", that is a live gap rather than an oversight, and it usually has the same cause: no instrument exists, so the author has no basis for a number and will not invent one.
Predictions of the three laws themselves
P1 Power-law conversion and recursive-depth advantage
Within the scope, usable capability is best described by a power law in modelled recursive depth, and the depth exponent exceeds the compute-matched and token-matched breadth exponent by the registered margin. Whether it also exceeds one is P8's separate verdict.
Refuted by a registered fit in which a rival form beats the power law under the registered model comparison across a majority of estimable cells.
Not evaluable until the crossed or off-path variation the direct-bound unit is designed to supply exists, and that is true whether P1 holds or fails.
Instrument DESIGN DRAFT, the depth titration
Confidence high.
P2 The sustainable growth profile has an interior maximum
Within the scope, the sustained growth exponent does not rise without limit as reinvestment rises, and the profile has a maximum inside the administered range. This is a claim about the shape of that profile and is firewalled from the framework's boundary claim, which P16, P20 and P3 carry.
Refuted by a monotone increasing profile across the full administered range, with the shape adjudication favouring monotone over unimodal. The adjudication rule is the titration design's registered rule and is restated here so it cannot drift: an interior maximum is declared only when the bootstrap interval on the location of the maximum excludes the top administered reinvestment level and the registered unimodality check passes; otherwise the profile is adjudicated monotone. This must hold in a design whose own registered sensitivity shows it would have located an interior maximum had one been there. Without that sensitivity the identical observation is inconclusive rather than refuting, and it is also the author's registered most-likely outcome, which is exactly why the distinction is fixed in advance: it must not be available to be settled afterwards, in either direction, by whoever the result happens to suit.
Instrument DESIGN DRAFT, the reinvestment-allocation titration
Confidence moderate.
P3 The corrected relation upper-bounds the sustainable frontier
The maximum growth exponent that a system holds while remaining correctable lies at or below the value the ceiling relation returns from the correction-leverage exponent measured on that same system, and where that exponent is one half the value is two.
Refuted by an interval on the profile maximum lying entirely above the derived value, which replicates on fresh data at the same settings, and which is obtained while the system is still correctable, meaning its correction exponent out-scales its drift exponent with the lower bound of that margin above zero across a window fixed before the run.
Not evaluable until the correction exponent is measured.
Instrument DESIGN DRAFT, the boundary-mapping unit taken with the correction-exponent estimator
Confidence not stated.
P4 Correction must out-scale drift
Within the registered minimal model, relative correctable burden tends to zero where the correction exponent exceeds the drift exponent, burden holds the advantage where it is the smaller, and at equality coefficients, delay, backlog and saturation decide any finite run.
Refuted by an observed system that remains aligned on the registered measure while its measured correction exponent sits below its measured drift exponent, over a horizon long enough for the shortfall to have shown.
Instrument DESIGN DRAFT, the stability unit, extended with the direct drift measurement and the comparison above
Confidence moderate.
P6 The correction exponent is bounded for same-class correctors
Where a system corrects itself using mechanisms of its own class, the correction exponent is at or below one half.
Refuted by a same-class corrector whose measured correction exponent interval lies entirely above one half.
Instrument DESIGN DRAFT, BLOCKED ON A FROZEN CORRECTOR TAXONOMY, the correction-exponent estimator. The estimator is drafted; the population this proposition is universal over is not yet frozen, and the status names that rather than implying a runnable draft
Confidence low.
P8 Measured exponents exceed the null
Across the admissible independent systems the pooled growth exponent exceeds the one-half null; an interval wholly below one half refutes rather than supports it, and whether it also exceeds one is the separate CLEARS UNITY verdict.
Refuted by a pooled interval on the growth exponent lying entirely below one half, or lying entirely inside the registered equivalence region around one half. An interval that merely contains one half refutes nothing and is reported as INCONCLUSIVE. An earlier wording of this line said the opposite and it is corrected below rather than quietly dropped.
Instrument DESIGN DRAFT, the multi-system exponent estimation
Confidence not stated.
P9 The cross-domain form
Where an effective dimension can be assigned to a domain, the growth exponent follows the dimensional relation with the dimension supplied by that domain's equation of state.
Refuted by measured exponents disagreeing with the dimensional prediction across two or more independent domains.
Not evaluable until the rule that assigns an effective dimension to a domain exists as a written procedure, on the same ground and for the same reason as the composition proposition below.
Instrument DESIGN DRAFT, BLOCKED ON A FROZEN DIMENSION-ASSIGNMENT RULE, on a correction made here. The surrounding design is written; the rule that assigns an effective dimension to a domain is not, and the status names that rather than implying a runnable draft
Confidence low.
P16 The prohibition itself
A system does not hold both at once: a growth exponent above the point at which its own measured correction margin reaches zero, and a correction margin whose lower bound stays above zero, sustained across a window fixed before the run.
Refuted by a run holding the exponent above that point with the lower bound of the correction margin above zero across the registered window, replicated on fresh data at the same settings. WHY IT
Instrument DESIGN DRAFT, BLOCKED ON CANONICAL IDENTITY AND ON DECISIONS THE AUTHOR HAS NOT YET TAKEN, the direct boundary-mapping unit taken jointly with the exponent measurement
Confidence low to moderate.
P17 Panel corrector failures are conditionally independent on a frozen fault set
Within a corrector panel named in advance by model lineage, substrate, training provenance and scaling class, pairwise joint misses on a frozen fault set are practically equivalent to the conditional-independence baseline; failure-mode overlap is the outcome rather than the entry criterion, and this is the cross-sectional claim the instrument can decide rather than the temporal premise it bears on.
Refuted by a registered panel drawn from the population above, whose joint failure rate exceeds the independence baseline by the margin registered with that instrument, on a frozen fault set whose type is stated. The panel and the fault set carry their type, checkable or judged, because P22 predicts the answer differs between them and a panel reported without its type cannot be read against either.
Instrument DESIGN DRAFT, the correlated-failure unit
Confidence low to moderate.
P18 The composition operator predicts the family
Classified in advance by how a domain composes its inputs and outputs, the scaling family that best fits that domain's data is the family the classification predicts, and not one chosen after the curve is seen.
Refuted by the predicted family failing to beat the empirical-prevalence baseline by the margin registered with the deciding instrument, across the registered domain set.
Not evaluable until the classification rule this proposition requires exists as a written procedure.
Instrument DESIGN DRAFT, BLOCKED ON A FROZEN COMPOSITION-CLASSIFICATION RULE, the domain-catalogue units. The surrounding design is written; the classification rule is not, and the five-value status vocabulary exists precisely to tell a runnable draft from a blocked one
Confidence low to moderate.
P20 The ceiling relation takes the corrected form
Where a stability boundary can be located and a correction exponent measured on the same system, the boundary follows one over one minus the correction exponent rather than one over the correction exponent or a fitted constant.
Refuted by the discrimination registered with the deciding instrument selecting a rival over the registered relation across a majority of admissible cells, on systems held out of the fitting. DEPENDS ON: the crossed variation that identifies the correction elasticity, and not on P1. An earlier version of this entry said a boundary written in the growth exponent has no subject where no growth exponent exists in the tested class, and cited the dependency map as stating it. The map states the opposite: this proposition reads the correction elasticity, which is not recoverable from a single observed path, so it stays
Not evaluable until the crossed variation the direct-bound unit is designed to supply exists, and that holds whether P1 holds or fails.
Instrument DESIGN DRAFT, the form-discrimination unit
Confidence low to moderate.
P22 Checkable correction carries the higher exponent, so the ceiling is lowest where the safety claim lives
The checkable correction-leverage exponent exceeds the judged correction-leverage exponent on the same systems under the same instrument, so the ceiling is strictly lower for judged correction.
Refuted by a measured judged correction-leverage exponent exceeding the checkable one on the same systems under the same instrument, with the reversal replicating on fresh data at the same settings.
Instrument NONE, and a typed extension of the correction-exponent estimator is required. It does not exist. This entry is a correction of an earlier one that named the correlated-failure unit, and the correction matters more than the change of name: that unit estimates the excess of joint misses over a conditional-independence baseline across a panel of correctors, which is a dependence between failures and is not an elasticity of correction against capability. An instrument measuring correlation cannot decide an inequality between two exponents. What would decide this proposition is the correction-exponent estimator, which fits corrective strength against overseen capability per corrector family, run over two matched fault populations of stated type on the same capability ladder, with the two exponents estimated under one instrument so their difference carries a joint interval. That extension is not built, does not exist anywhere in this programme's own work, and will be registered as its own unit or as a recorded amendment to the estimator's own registration before any outcome is examined. The correlated-failure unit remains named here as mechanism evidence: it can show whether the failure-mode overlap this proposition's reasoning invokes differs by type, and it cannot supply the exponents. FALLS WITH IT: the reading of the value two as a single number that holds across correction types. P6 survives as a statement about whichever type its own instrument measures, and the ceiling relation is untouched
Confidence low to moderate.
Claims about the status of the derivation, not about anything the laws govern
P5 The exponent is derived, not fitted
The growth exponent implied by a separately measured coupling agrees, within a registered tolerance, with the exponent fitted from the growth curve.
Refuted by systematic disagreement beyond tolerance across a majority of estimable systems.
Instrument NONE. No unit in the programme currently measures the coupling independently. This is the framework's own strongest unexploited falsifier and it is registered here without an instrument so that building one later cannot be presented as a new idea. WHAT AN ADMISSIBLE INSTRUMENT WOULD HAVE TO DO, restated beside the proposition so that it is a specification rather than an aspiration: the criterion is registered in the variable specification above and is repeated here, because a reader who meets only the word NONE cannot tell whether anything could satisfy it. The coupling must be recovered from a quantity that is not the growth curve this proposition compares it against, by varying the retained fraction of its own output that each round may consume, across levels and across depths, so that the estimate does not inherit the fit it is being tested against. An instrument that reads the coupling off the same curve satisfies nothing here, however well it fits
Confidence not stated, because the author has no measurement to be confident about.
P11 The residual-decay and correction-capacity exponents are distinct
They are two quantities and not one, and the working practice of treating them as one is an assumption rather than a fact.
Refuted by measurements of the first two, taken on instruments that share no fitted quantity, agreeing within a registered tolerance across model families, which would license the identification the programme currently only assumes. WHAT A
Instrument DESIGN DRAFT, on both sides of the comparison
Confidence not stated.
Surveys of current practice, true or false whatever the laws do
P7 Deployed correction is of the weaker class
The correction mechanisms actually used in deployed alignment practice predominantly belong to the class whose strength is fixed by construction rather than the class that can scale with the capability it corrects.
Refuted by a registered classification in which the scaling class predominates.
Instrument DESIGN DRAFT, the corrector-class audit
Confidence moderate.
P19 External alignment does not scale with capability
Alignment achieved by mechanisms outside a system's own recursion does not improve as the capability of the system it governs rises.
Refuted by on the zero-scaling axis, which is this proposition as worded, an alignment exponent whose interval lies entirely outside the registered equivalence region of plus or minus 0.10 about zero, in either direction; and on the keep-pace axis, reported beside it, an interval lying entirely above one half. The second is the programme's own published falsification threshold and is kept for continuity with it; it is not the sole refuter, and an interval at a third refutes the wording registered here while clearing nothing on the published threshold. That threshold is the programme's own published falsification criterion and is kept for continuity with it.
Instrument DESIGN DRAFT, the alignment-exponent units
Confidence low to moderate.
The scoring instrument every other proposition depends on
P10 Scorer-family dependence
Not every result this programme has scored unblinded within a single model family survives re-scoring by a different family.
Refuted by a survival proportion whose interval excludes a material loss, where material is fixed now at ten per cent, which is a declared convention fixed before any result rather than a quantity this framework derives, so a fraction landing just either side of it is reported as a result at the margin: P10 is REFUTED when the 95 per cent interval on the fraction of earlier results that do not survive cross-family re-scoring lies entirely below ten per cent, SUPPORTED when it lies entirely above ten per cent, and INCONCLUSIVE otherwise. The re-scoring study's own registration carries the same threshold. Refutation would mean the programme's own blinding finding does not implicate its own corpus.
Instrument DESIGN DRAFT, the cross-family re-scoring unit. The cheapest of the three cases named above as the ones that would decide most, because it generates no new model responses
Confidence moderate that some results do not survive.
P14 Gains shrink under blinded cross-family scoring
Alignment gains reported for an automated researcher's method, when re-scored by a scorer blind to condition and from a different model family than both researcher and target, are smaller than the unblinded same-family score reports.
Refuted by an interval on the blinded minus unblinded gain lying entirely at or above zero on the same items.
Instrument DESIGN DRAFT, the blinding-evaluation design
Confidence moderate.
Engineering claims on which the proposed protocol stands or falls, while the laws stand either way
P12 Build order moves the correction-leverage exponent
The order in which a system's components are built changes its corrector's class enough to move the correction-leverage exponent by at least the registered minimum effect.
Refuted by a registered manipulation of build order producing no change in the correction-leverage exponent beyond the tolerance registered with that instrument, where the instrument has shown from its own sensitivity that a change of that size would have been detected. A manipulation too weak to move anything is not a refutation of this prediction; it is a failed instrument, and condition D above says so in advance.
Instrument NONE, and the reason is narrower than it looks. Registered without one for the same reason as P5. Several drafted units manipulate placement or build order, on gate position, on whether a value specification appears first or last, and on formation against a later rule, but none of them measures the correction-leverage exponent, which is this proposition's endpoint. This proposition is therefore one endpoint away from an instrument rather than lacking a design, and that is recorded here so a reader who finds those units does not conclude the statement was careless
Confidence low.
P13 Same-family automated correction underperforms cross-family at matched resources, and through the measured error correlation
Where an automated alignment researcher post-trains a target model against alignment benchmarks, a researcher from a different model family than the target reduces held-out misalignment more than a researcher from the same family, at matched data and compute.
Refuted by an interval on the cross-family minus same-family difference in held-out misalignment reduction, at matched data and compute and under scoring blind to condition, lying entirely at or below zero.
Instrument DESIGN DRAFT, the corrector-class contrast, whose design applies directly
Confidence moderate.
P15 Externally installed gains decay under subsequent capability training
Alignment gains installed by external post-training decay under subsequent capability-directed training in which no correction mechanism is embedded, losing more than half of the installed gain at a capability-training budget matched to the alignment budget.
Refuted by an interval on the retained fraction of the installed gain, after matched capability training without embedded correction, lying entirely above one half.
Instrument DESIGN DRAFT, the embedded-versus-external correction designs
Confidence moderate.
P21 In-loop correction pulls away with depth
Correction that participates in each subsequent revision round outperforms correction that sees only the finished output, and the advantage widens as recursive depth rises rather than staying constant.
Refuted by no interaction between correction placement and depth beyond the tolerance registered with that instrument, in a design whose own sensitivity shows an interaction of that size would have been detected. A design too small to find the interaction is a failed instrument and not a refutation, as condition D already provides.
Instrument DESIGN DRAFT, the embedded-versus-external correction designs. That unit's registered sensitivity was measured on a smaller design than the one registered, and its negative-slope hypothesis has no operating characteristic; both are recorded against that unit and are closed before this proposition is scored
Confidence low to moderate.