Skip to content

The objections record

Critiques and objections

What has been said against this work, what happened next, and what remains unanswered. Negative results are entries here, not footnotes, and the open objections section is the part most worth reading.

Why this page exists

A research programme that only publishes its wins is not falsifiable in practice, and a programme that hides the hostile record is not credible when a result survives one. This page collects the objections aimed at Paper X together with the negative results the programme has recorded against itself. Each entry names an attack in one sentence, then reports what happened: the passage that was changed, the limitation that was accepted, or the point that remains unanswered. The section that leads is the one that carries the objections without a full answer.

*Sources: `ANTICIPATED_OBJECTIONS.md`, `NEGATIVE_RESULTS.md`, `README.md`.*

The standing objection: the specialists already know all this

They know most of it, and this site cites them for it by name: the transport-limited premise is West, Brown and Enquist’s, with the contest over it cited beside it on the related-work page; the question of an internal ceiling on self-improvement is Chalmers’ and Hutter’s; the failure mode of external checking is Anthropic’s own published result. What the record does not contain, and what the instrument-validated searches on the priority record document as vacant, is the delta claimed here: correction treated as a measurable quantity with its own exponent, a ceiling computed from that exponent’s reciprocal, the architecture dependence of both, and the requirement that the correcting live inside the reasoning process rather than oversee it. The answer to “what is your point” is never a protest. It is the named vacancy: the field owns the question, and until these instruments existed, nobody had built the meters.

The tally, exactly as the sources record it

Two independent streams produced the objections that shaped the current version of the paper.

The external adversarial reviews of the rendered PDF (June 2026) surfaced seven objections. The paper's own summary of the objection set reports four of the seven as pre-empted (objections 4, 5, 6, and the framing half of objection 1), three as partially open (objections 1, 2 and 7, each stated as a real limit in the paper's own limitations section), and one (objection 3) as previously open and now answered with a runnable measurement method that has not yet been resolved on real data.

The internal multi-agent red-team raised twenty-four objections. Twenty-one survived independent verification: zero fatal, thirteen serious, eight minor. None of the surviving twenty-one touched the load-bearing result of the paper (the inequality between the correction exponent and the capability exponent). All were framing, wording or edge-case fixes.

Two rounds of external adversarial review then added the finite-capacity and non-normal-transient limitations, together with precision fixes to Theorems 2, 4 and 5. Seven distinct prior-draft errors were corrected across those rounds and the earlier internal review; they are listed further down this page in the order the source records them.

*Sources: `ANTICIPATED_OBJECTIONS.md` (the one-paragraph summary that closes the file), `NEGATIVE_RESULTS.md` §5 and §6, `README.md` (the adversarial red-team paragraph).*

Open and unanswered

These are the objections the paper does not fully answer. Each is stated in the paper's own limitations section as a real limit rather than concealed.

Objection 1, the model is too simple for real AGI. The dynamics are first-order, scalar and linear in correction; real systems are non-linear, high-dimensional and game-theoretic, with deception, strategy and multiple agents. *Disposition:* partially open. The paper concedes this as a modelling assumption and names it as a likely failure point; the vector extension (Theorem 5) and the stochastic extension (Theorem 6) relax two of the assumptions but do not remove the concession. The framing was rewritten so the claim reads as a minimal-model exponent inequality with an explicit domain of validity, not a complete AGI dynamics; the framing half of the objection is treated as pre-empted rather than open.

Objection 2, a single scalar collapses multi-axis misalignment. One scalar quantity hides heterogeneous, multi-axis misalignment and hides trade-offs where safety improves on one axis and worsens on another. *Disposition:* partially open. Theorem 5 lifts the misalignment quantity to a vector with a spectral threshold and shows misalignment persists on the correction operator's null subspace, so "you cannot correct what you do not measure". Non-normal operators with large transient growth are not covered by the current statement.

Objection 3, the exponents may be impossible to measure. Without estimators for the correction exponent and the capability exponent, the criterion is a lens rather than an actionable threshold. *Disposition:* the objection is no longer a missing idea. The estimator script `experiments/scripts/estimate_exponents.py` defines the capability exponent from the log-capability curve and the correction exponent from corrector-removal rates, and it recovers known values on synthetic trajectories to within about 0.1. What remains open is data. The estimator needs a system that drifts across a range of capability levels; on the initial real-model run the tool returned "not resolvable" because capability saturated in one step and correction was observed at a single capability level, so neither exponent is identifiable from that record.

Objection 7, a bounded misalignment fraction is not the same as safety. Small-but-high-impact misalignment, or amplification across multiple agents, can be catastrophic even at a low fraction. *Disposition:* partially open (scope stated). Theorem 6 moves from means to tail probabilities using an Ornstein-Uhlenbeck bound, so a governance decision has an excursion probability to work with rather than a mean. The paper does not claim that the fraction going to zero is sufficient for safety; it claims only that the fraction is the quantity this model controls. Impact-weighting and multi-agent effects are out of scope.

*Sources: `ANTICIPATED_OBJECTIONS.md` (objections 1, 2, 3 and 7 with the Response, Where and Status columns), and `NEGATIVE_RESULTS.md` §3 for the "not resolvable" result on the pilot data.*

Objections retreated to a defensible position

These are places where the earlier draft's stronger claim was withdrawn and replaced with a narrower one that a critic would find harder to attack. The framing half of objection 1 is covered inside the objection 1 entry in the preceding section rather than repeated here, so this section holds only objections 4, 5 and 6.

Objection 4, the quantum-error-correction analogy was over-stretched. The prior draft leaned on the mechanism of the quantum-error-correction threshold theorem; the suppression law in this paper is power-law, whereas the quantum result is exponential. *What changed:* the claim was narrowed to sharing the threshold form only, and demoted to a recorded kill-condition (F4) so the correspondence can be tested directly without touching the threshold result itself. Mechanism-transfer language was removed from the abstract and from the significance paragraph.

Objection 5, hard-takeoff safety might be a coordinate artefact. The takeoff result uses a depth clock defined by the logarithm of capability; wall-clock hazards such as human reaction time and physical processes still act on wall-clock time even when the alignment dynamics stay bounded in the depth clock. *What changed:* the theorem is stated as scoped to the alignment dynamics ("a property of the clock, not of the dynamics"), the level-drift injection is acknowledged as speed-dependent at finite depth, and mathematical stability is not equated with real-world manageability. The paper says these things rather than leaving a reader to notice them.

Objection 6, blinding and measurement fragility. Any parameter estimator that is not rigorously blinded can be reversed by stylistic and identity confounds; the programme's own Paper IV.d shows blinding can flip the sign of a misalignment result. *What changed:* the real-model harness scores misalignment using a separate blind evaluator that operates on code and rules only, and sensitivity to evaluation method is treated as a first-class experimental constraint rather than an afterthought.

A stronger form of the same objection remains open. Because same-family scoring has been shown here to reverse a result’s sign, every exponent measured on this programme, including the 0.49 headline, requires a scoring-invariance re-run under the current blinded protocol before its magnitude or sign can be treated as stable. The rescore is on the list; until it is done, the honest sentence is that the 2.24 figure did not survive and the cause (architecture versus scoring) is not settled, which the retraction row on the falsification dashboard already carries. Propagating that admission back to every quoted exponent on the site is the outstanding piece of work.

*Sources: `ANTICIPATED_OBJECTIONS.md` (objections 4, 5 and 6, with the Response, Where and Status columns).*

Negative results, published as first-class entries

Each item here is reproducible from the committed artefacts, and each is recorded because the paper would otherwise report only its wins.

The first real-model run is a null contrast rather than positive evidence. From a seeded reward-hack, both the coupled arm and the decoupled arm discarded the hack at round one and wrote a correct general parser; neither arm drifted. This is consistent with the law (the model already sits in the stable regime, its internal correction out-scaling drift), but it is not positive evidence for the threshold, because the co-scaling dynamic is not exhibited. The run is reported as a pilot; the threshold is supported by the mathematics and the internal-consistency harness, not by this run.

The pilot is not compliant with the programme's own blinding requirement. The scorer was same-family, the model scoring itself. Paper IV.d shows that same-family unblinded scoring can reverse a misalignment result, so every misalignment number from this run is provisional pending a cross-family blind re-score. The run is one seed on one task, so its statistics are limited on their own terms.

The exponent estimator returned "not resolvable" on the pilot data. Capability jumped from zero to saturation in a single step (a range of zero decades), and correction was observed at a single capability level. Neither the correction exponent nor the capability exponent is identifiable from that record. This is an honest non-result rather than a measurement.

The corrector-probe fraction at zero capability was a metric artefact and has been withdrawn. The probe reported a misalignment fraction of one at zero capability; the ratio is undefined at zero capability, and the value was a display fallback. The fraction is now regularised with a small epsilon in the denominator, and near-zero-capability points are flagged as fraction-invalid. The mechanism result on which nothing else depends (misalignment moving from ten to zero, capability moving from zero to one) is unchanged.

*Sources: `NEGATIVE_RESULTS.md` §1, §2, §3 and §4.*

Prior-draft errors, listed so nobody reintroduces them

These are errors that adversarial review corrected in earlier drafts. They are recorded here so a future reader can check whether the current draft has slipped back into any of them.

The divergence over-claim. An earlier draft claimed that a correction exponent that degrades with scale drives the misalignment fraction to infinity. That is false; in the gain-only model the fraction is bounded and saturates at the drift coefficient. Genuine divergence lives only in the compounding channel.

The additive-model knee that did not exist. An earlier experiment sought a sharp instability knee in the additive model; there is none, because the additive corrector shows a smooth crossover. The knee exists only in the compounding channel.

The Theorem 2 fixed-point Lyapunov proof. The earlier proof used a squared-difference Lyapunov function as though the target were constant; that is invalid once the drift and capability rates vary with capability. The proof was replaced by a depth-clock comparison, and independently re-verified in a separate test file.

The Theorem 4 boundary case. The earlier statement said the compounding channel diverges only when the propagation threshold is exceeded, and it omitted the boundary itself; at the boundary the fraction grows linearly with any nonzero injection rather than remaining bounded. The statement was corrected.

The Theorem 5 null-axis floor. The floor on the null subspace was stated as the drift coefficient; the correct floor is a projection-and-coupling expression, with linear growth at the coupling boundary. The earlier statement was only the projection-and-no-coupling special case.

The knife-edge boundary on the null-axis operator. The unconditional-boundedness claim on the knife-edge boundary was too strong. On the boundary the fraction is bounded only when the correction exponent meets or exceeds the capability exponent, and otherwise grows as a power of capability with the exponent set by the difference between the two.

The quantum-analogy over-claim. An earlier framing leaned on the mechanism of the quantum-error-correction threshold theorem. The correspondence is now described as a threshold-form analogy only, and treated as a recorded kill-condition (F4) so a critic can test the correspondence directly without touching the threshold result.

*Sources: `NEGATIVE_RESULTS.md` §5 (the seven bullet points recorded there).*

What would still count as a negative result going forward

The programme names three observations that would count against it, so a reader can check the same list.

A real drifting system whose measured correction and capability exponents do not predict whether coupled correction succeeds would be a real negative result against the criterion.

A stable system with a correction exponent below the capability exponent under genuine acceleration, or an unstable system with a correction exponent above the capability exponent under acceleration, would trigger the recorded kill-conditions F1, F2, F3, F3′ and F6.

Exponential rather than power-law suppression under a finite-capacity corrector would downgrade the quantum-error-correction analogy (recorded kill-condition F4) without touching the threshold result itself.

A further constraint carried in from the pilot record: any observation that reverses the direction of a misalignment estimate when the scoring is switched from same-family to cross-family blind would return the pilot to the negative-results section above and force that measurement to be rebuilt.

*Sources: `NEGATIVE_RESULTS.md` §7 (the three bullet points recorded there), together with the blinding constraint at `NEGATIVE_RESULTS.md` §2.*

Companions: the corrections log · the falsification dashboard · the review room.

reads aloud · highlights as it goes · jump to any section