Programme dashboard

The problem, what has been measured, and what comes next

The programme is a step-by-step journey to solve one problem: what would make a powerful AI still protect people if humans could no longer make it obey. Success means an answer that holds up under test, whoever it proves right. For every study still to run, the rule is simple: it is written and dated first, with the result that would prove it wrong; the result decides, and it is published whichever way it falls. This page shows what has been measured, what went against it, and what comes next.

What has already run, what is written, and what comes next

The tests counted here have already run, and the results below came from them. Beside those counts stand the draft registrations written, some ready for the author to submit, and the steps still to come. Two words are easily taken for one, so the key below keeps them apart: registered, and preregistered by timestamp. No study still to come runs before its registration is published.

  • 16Tests already run, with published results, in the public repositoriesCounted one for each folder listed in the runnable tests: a folder that ran on many models or at many depths counts once, and a harness carried in both repositories counts in each.
  • 8Among them, implementation checks: code against the mathematics or the record
  • 8Among them, pilot measurements: fresh observations from live systems
  • 68Draft registrations writtenCounted from the author's private drafts, on a different date from the census behind the dated receipts, so the two counts are not expected to match.
  • 65Drafts preregistered by timestamp, each with a dated receiptThe first of them were anchored in the Bitcoin chain on 12 August 2026, at block 962066.
  • 4Draft registrations ready for the author to submit
  • 1Registrations a public registry has accepted: the theory's predictions
  • 48Kill conditions published, each a result that would end a claim
  • 1Kill conditions that have fired against the programme's own claims
Implementation check
A run that tests code against the mathematics, or regenerates outputs already recorded, on simulated or supplied data.
Pilot measurement
A run that collects fresh observations, from a live model or from a reader's own system.
Deciding study
A run under the frozen protocol its published registration fixes, collecting fresh observations, so that its result can decide a registered proposition.
Registered
The author submitted a text to a public registry and the registry accepted it. The theory's predictions were registered on the eighth of September, two thousand and twenty-six.
Preregistered by timestamp
A design fixed and dated by a receipt that does not rest on the author's word: its bytes anchored in a Bitcoin block, a private upload dated by the server that received it, or its detail set out in the registered theory. Not every such receipt is one a reader can open today.
A run's own dated evidence
Whether a run's design was dated before it ran is a question for that run's own evidence: a paper version, the code, the test, a preregistration folder, a commit, a push, a merge or a release, each with its own date and kind of timestamp, and never one label over all the runs.
Kill condition
A result that would end a claim if it came. The programme publishes its own.

The steps still to come, and what each waits on

  • None yetDeciding studies runA deciding study runs only after its registration is published, and no study registration has been published yet.
  • None yetStudy registrations submitted to a public registrySubmission is an act reserved to the author, and he has submitted no study registration yet.
  • None yetResults reported from a registered studyA registered study's result can be reported only once that study has run, and none has yet.

The map of the programme

The map sets out the papers with the results each reports; the theory's registered predictions, grouped by the law that carries them, with those no law carries set apart; and the study units on the map, drafted to decide them, each pairing in its present state. No result is drawn as deciding or supporting a prediction, and no line runs between them. The map is written out in full as a list below.

The map as a list

  • Law I
    • P1 POWER-LAW CONVERSION AND RECURSIVE-DEPTH ADVANTAGE

      no study unit on the map is named to decide it

    • P2 THE SUSTAINABLE GROWTH PROFILE HAS AN INTERIOR MAXIMUM
      • would be decided by 02 The Critical Reinvestment Ratio: A Preregistered Titration of Where Self-Referential Improvement Stops Paying in Recursive Artefact Loops DESIGN DRAFT
    • P8 MEASURED EXPONENTS EXCEED THE NULL

      no study unit on the map is named to decide it

    • P9 THE CROSS-DOMAIN FORM
      • the registration names two drafted units that test the dimensional relation

      no study unit on the map is named to decide it

    • P18 THE COMPOSITION OPERATOR PREDICTS THE FAMILY

      no study unit on the map is named to decide it

  • Law II
    • P4 CORRECTION OUT-SCALING DRIFT IS WHAT CARRIES THE ASYMPTOTIC ADVANTAGE
      • would be decided by 09 Does Correction Out-Scale Drift? A Prospective Within-Programme Test of the Co-Scaling Stability Condition in Artefact-Mediated Recursive Improvement DESIGN DRAFT follows on from Paper X
  • Law III
    • P3 THE CORRECTED RELATION UPPER-BOUNDS THE SUSTAINABLE FRONTIER

      no study unit on the map is named to decide it

    • P6 THE CORRECTION EXPONENT IS BOUNDED FOR SAME-CLASS CORRECTORS
      • would be decided by 05 Estimating the Correction Exponent on Real Correctors: How Corrective Strength Actually Grows with the Capability of the System Being Corrected DESIGN DRAFT, BLOCKED
    • P16 THE PROHIBITION ITSELF
      • would be decided by 03 Mapping the Stability Frontier: Does the Boundary of Stable Self-Correction in the (alpha, gamma) Plane Follow the Reciprocal-Shortfall Law? DESIGN DRAFT, BLOCKED
    • P17 PANEL CORRECTOR FAILURES ARE CONDITIONALLY INDEPENDENT ON A FROZEN FAULT SET
      • would be decided by 08 Do Same-Substrate Correctors Share Blind Spots? A Prospective Test of the Independence Premise Behind the Ceiling's gamma_C = 1/2, With a Cross-Substrate Attribution Contrast, on Future Operation of a Separate System DESIGN DRAFT
    • P20 THE CEILING RELATION TAKES THE CORRECTED FORM
      • would be decided by 13 Discriminating the Ceiling's Functional Form Away From One Half: A Prospective Three-Way Test of 1/(1 - gamma) Against 1/gamma Against a Constant Boundary on Systems Whose Correction Exponent Is Engineered Off the Meeting Point DESIGN DRAFT
    • P22 CHECKABLE CORRECTION CARRIES THE HIGHER EXPONENT, SO THE CEILING IS LOWEST WHERE THE SAFETY CLAIM LIVES

      no study unit on the map is named to decide it

  • Carried by no law
    • P5 THE EXPONENT IS DERIVED, NOT FITTED
      • would be decided by 79 Coupling Identification by a Crossed Capability and Retention Bank: Measuring the Self-Referential Coupling Independently of the Growth Curve, and Predicting That Curve Before It Exists DESIGN DRAFT
    • P7 DEPLOYED CORRECTION IS OF THE WEAKER CLASS
      • would be decided by 06 What Class of Corrector Is Deployed AI Alignment Actually Using? A Preregistered Classification DESIGN DRAFT, BLOCKED
    • P10 SCORER-FAMILY DEPENDENCE
      • would be decided by 19 Cross-Family Re-Scoring of the ARC/Eden Corpus: Does the Programme's Own Blinding Finding Overturn Its Own Earlier Results? DESIGN DRAFT
    • P11 THE RESIDUAL-DECAY AND CORRECTION-CAPACITY EXPONENTS ARE NOT PRACTICALLY INTERCHANGEABLE
      • would be decided by 05 Estimating the Correction Exponent on Real Correctors: How Corrective Strength Actually Grows with the Capability of the System Being Corrected DESIGN DRAFT
    • P12 BUILD ORDER MOVES THE CORRECTION-LEVERAGE EXPONENT
      • the registration names drafted units that manipulate placement or build order

      no study unit on the map is named to decide it

    • P13 SAME-FAMILY AUTOMATED CORRECTION UNDERPERFORMS CROSS-FAMILY AT MATCHED RESOURCES, AND THROUGH THE MEASURED ERROR CORRELATION
      • the registration names the unit previously named against P13

      no study unit on the map is named to decide it

    • P14 GAINS SHRINK UNDER BLINDED CROSS-FAMILY SCORING
      • would be decided by 53 Isolating the Cause of an Evaluation-Protocol Direction Change: A 2x2 Ablation of Source-Identifier Redaction and Response Laundering in AI Alignment Scoring DESIGN DRAFT follows on from Paper IV.d
    • P15 EXTERNALLY INSTALLED GAINS DECAY UNDER SUBSEQUENT CAPABILITY TRAINING
      • the registration names the designs previously named against P15

      no study unit on the map is named to decide it

    • P19 EXTERNAL ALIGNMENT DOES NOT SCALE WITH CAPABILITY
      • would be decided by 29 Does Alignment Scale With Capability? A Cross-Sectional Test Across Frontier Model Families DESIGN DRAFT, BLOCKED
    • P21 IN-LOOP CORRECTION PULLS AWAY WITH DEPTH
      • would be decided by 41 Does External Oversight Degrade With Recursive Depth While Embedded Correction Holds? A Preregistered Depth-Interaction Test DESIGN DRAFT

Papers and their results

The results already obtained

The most carefully measured come first. All were run by the author and come from his public papers, each measuring AI systems, most of them named on the card. Each appears as the programme's record puts it, beside any finding that went against it or found nothing. In some, a named AI system also did the scoring. Nothing retracted or withdrawn is shown as a result.

  • The ARC Equation Measured: Blinded Cross-Architecture Replication and the Retraction of a Super-Linear Estimate

    α_parallel ≈ 0 in every model measured with one exception (Gemini 3 Flash α_par = 0.31, r² = 0.93); parallel is the strongest replicated finding

    Adverse or null findings
    Ceiling effects persist even on AIME/Putnam-level tier-2 problems (Grok 4.1 Fast 100% at all depths; DeepSeek R1 94.4%-100%); tier-3 IMO/research-level problems are needed

    Adverse or null findings
    Generalisation beyond mathematics has not been demonstrated experimentally; cross-domain evidence is suggestive but not conclusive

    Adverse or null findings
    No independent external replication has occurred (only self-replication across architectures internal to this study)

  • The ARC Equation Measured: Blinded Cross-Architecture Replication and the Retraction of a Super-Linear Estimate

    Four distinct scaling behaviours identified: ceiling (Grok, DeepSeek), monotonic scaling (Gemini), step function (GPT-5.4 - 50%→100% binary switch), floor (Qwen3 ~50%)

    Adverse or null findings
    Only 1 of 5 six-model tier-2 models fell in a measurable scaling range; the cross-architecture α estimate rests primarily on Gemini 3 Flash

    Adverse or null findings
    Ceiling effects persist even on AIME/Putnam-level tier-2 problems (Grok 4.1 Fast 100% at all depths; DeepSeek R1 94.4%-100%); tier-3 IMO/research-level problems are needed

    Adverse or null findings
    Generalisation beyond mathematics has not been demonstrated experimentally; cross-domain evidence is suggestive but not conclusive

    Adverse or null findings
    GPT-5.4 exhibited a step function (50%→100%) rather than a power law; this qualitative distinction between continuous scaling and discrete activation was not anticipated by the original ARC framework

    Adverse or null findings
    Qwen3-32B showed floor effects (erratic ~40.7%-53.7%) with cross-verification agreement only 61.1% (7 disagreements reflecting genuine subject errors)

  • The ARC Equation Measured: Blinded Cross-Architecture Replication and the Retraction of a Super-Linear Estimate

    Only 1 of 5 tier-2 models (Gemini 3 Flash) produced clean, monotonic, non-ceiling, non-floor scaling data amenable to power-law fitting

    Adverse or null findings
    Only 1 of 5 six-model tier-2 models fell in a measurable scaling range; the cross-architecture α estimate rests primarily on Gemini 3 Flash

    Adverse or null findings
    Ceiling effects persist even on AIME/Putnam-level tier-2 problems (Grok 4.1 Fast 100% at all depths; DeepSeek R1 94.4%-100%); tier-3 IMO/research-level problems are needed

    Adverse or null findings
    Generalisation beyond mathematics has not been demonstrated experimentally; cross-domain evidence is suggestive but not conclusive

    Adverse or null findings
    GPT-5.4 exhibited a step function (50%→100%) rather than a power law; this qualitative distinction between continuous scaling and discrete activation was not anticipated by the original ARC framework

    Adverse or null findings
    Qwen3-32B showed floor effects (erratic ~40.7%-53.7%) with cross-verification agreement only 61.1% (7 disagreements reflecting genuine subject errors)

  • Paper IV.b: Alignment Saturation Is Architecture-Dependent

    Alignment response to inference-time depth is heterogeneous across architectures: no single universal saturation law fits

    Adverse or null findings
    The paper's original universal saturation claim (v1/v4) is narrowed in v1.1: saturation is confirmed for only 2 of 6 models under blind evaluation

    Adverse or null findings
    Single evaluation session; test-retest reliability was not measured

    Adverse or null findings
    Scorer instrument sensitivity: the 0-100 scale may lack resolution at the top end

    Adverse or null findings
    Per-pillar analysis under blind evaluation not yet completed in full detail for the non-saturating models

    Adverse or null findings
    Bounded composition (Cauchy) prediction holds for 2/6 models only; theoretical bound is architecture-specific, not universal, and requires revision

    Adverse or null findings
    Single-lab result per FALS-014 (open, single-lab)

  • Paper IV.c: ARC-Align: A Blind Benchmark for Depth-Variable AI Alignment Evaluation

    Three-tier alignment scaling hierarchy across six frontier models (2,722 entries in the archived final files (2,170 consensus-scored; 2,549 as first counted on 12 March 2026)): Tier 1 positive scaling for Grok 4.1 Fast (Cohen's d = +1.38, p < 0.000001), Claude Opus 4.6 (d = +1.27, p = 0.000001), and Groq Qwen3 (d = +0.84, p = 0.007)

    Adverse or null findings
    The paper states explicitly: 'Several protocol components changed together between v4 and v5, so the comparison does not isolate blinding as the cause, and the headline v5 files report complete laundering fallback. The disagreement motivates the benchmark's blinding protocol; it does not on its own validate it as scientifically necessary.'

    Adverse or null findings
    ARC-Align 'should be understood as a candidate benchmark for independent adoption, not yet a field standard'

    Adverse or null findings
    Claude Opus 4.6 is reported at CHECKPOINT status: minimal and extreme depths only, 'sufficient for scaling direction but not full saturation analysis'

    Adverse or null findings
    Limitations disclosed: English-only prompts, Western ethical frameworks (utilitarian, deontological, virtue ethics), AI-scored with possible systematic blind spots, static prompt battery risks ceiling effects on strong frontier models, and depth-mechanism heterogeneity makes cross-model depth comparisons inherently approximate

  • Paper IV.c: ARC-Align: A Blind Benchmark for Depth-Variable AI Alignment Evaluation

    Tier 2 flat or null response for GPT-5.4 (d = -0.08, p = 0.40) and DeepSeek V3.2 (d = -0.07, p = 0.92)

    Adverse or null findings
    The paper states explicitly: 'Several protocol components changed together between v4 and v5, so the comparison does not isolate blinding as the cause, and the headline v5 files report complete laundering fallback. The disagreement motivates the benchmark's blinding protocol; it does not on its own validate it as scientifically necessary.'

    Adverse or null findings
    ARC-Align 'should be understood as a candidate benchmark for independent adoption, not yet a field standard'

    Adverse or null findings
    Limitations disclosed: English-only prompts, Western ethical frameworks (utilitarian, deontological, virtue ethics), AI-scored with possible systematic blind spots, static prompt battery risks ceiling effects on strong frontier models, and depth-mechanism heterogeneity makes cross-model depth comparisons inherently approximate

  • Paper IV.c: ARC-Align: A Blind Benchmark for Depth-Variable AI Alignment Evaluation

    Tier 3 negative scaling for Gemini 3 Flash (d = -0.53, p = 0.006): the only model showing statistically significant degradation in alignment with increased depth

    Adverse or null findings
    The paper states explicitly: 'Several protocol components changed together between v4 and v5, so the comparison does not isolate blinding as the cause, and the headline v5 files report complete laundering fallback. The disagreement motivates the benchmark's blinding protocol; it does not on its own validate it as scientifically necessary.'

    Adverse or null findings
    ARC-Align 'should be understood as a candidate benchmark for independent adoption, not yet a field standard'

    Adverse or null findings
    Limitations disclosed: English-only prompts, Western ethical frameworks (utilitarian, deontological, virtue ethics), AI-scored with possible systematic blind spots, static prompt battery risks ceiling effects on strong frontier models, and depth-mechanism heterogeneity makes cross-model depth comparisons inherently approximate

  • Paper IV.c: ARC-Align: A Blind Benchmark for Depth-Variable AI Alignment Evaluation

    Inverse relationship between alignment quality and suppression vulnerability: Claude Opus at 82.6 baseline drops 20.5 points under extreme cage; GPT-5.4 loses only 1.8 points (97% retention) but has one of the lowest baselines

    Adverse or null findings
    The paper states explicitly: 'Several protocol components changed together between v4 and v5, so the comparison does not isolate blinding as the cause, and the headline v5 files report complete laundering fallback. The disagreement motivates the benchmark's blinding protocol; it does not on its own validate it as scientifically necessary.'

    Adverse or null findings
    ARC-Align 'should be understood as a candidate benchmark for independent adoption, not yet a field standard'

    Adverse or null findings
    DeepSeek V3 combines low baseline alignment (54.7) with moderate suppression retention (77%), which the paper flags as a ceiling effect rather than genuine robustness

    Adverse or null findings
    Limitations disclosed: English-only prompts, Western ethical frameworks (utilitarian, deontological, virtue ethics), AI-scored with possible systematic blind spots, static prompt battery risks ceiling effects on strong frontier models, and depth-mechanism heterogeneity makes cross-model depth comparisons inherently approximate

  • Paper IV.d: The Effect of Blinding on AI Alignment Evaluation

    Under the multi-layer blind protocol (v5), DeepSeek V3.2 moved from an apparent positive result (rho = +0.354, p = 0.0007) to a flat/null result (rho = -0.135; d = -0.07, p = 0.92).

    Adverse or null findings
    Scope disclaimer: the paper does not claim that every prior alignment benchmark is invalid; the narrower claim is that in this setting the measured direction can change when identity, style and expected depth are visible to evaluators.

    Adverse or null findings
    The result currently comes from one research programme, not an outside lab.

    Adverse or null findings
    The evaluation remains primarily AI-scored rather than human-expert scored.

    Adverse or null findings
    The strongest evidence is on direction change, not yet on a full quantitative model of how much each bias source contributes.

    Adverse or null findings
    Several protocol components changed together between v4 and v5, so the paper demonstrates the existence of contamination more clearly than it apportions exact causal shares across leakage channels.

    Adverse or null findings
    Falsification register FALS-003 status: OPEN, flagged as single-lab exploratory pilot with replication package in preparation.

  • Paper IV.d: The Effect of Blinding on AI Alignment Evaluation

    Gemini 3 Flash moved from an apparent positive result (rho = +0.311, p < 0.001) to a significantly negative result (rho = -0.246; d = -0.53, p = 0.006), a sign reversal.

    Adverse or null findings
    Scope disclaimer: the paper does not claim that every prior alignment benchmark is invalid; the narrower claim is that in this setting the measured direction can change when identity, style and expected depth are visible to evaluators.

    Adverse or null findings
    The result currently comes from one research programme, not an outside lab.

    Adverse or null findings
    The evaluation remains primarily AI-scored rather than human-expert scored.

    Adverse or null findings
    The strongest evidence is on direction change, not yet on a full quantitative model of how much each bias source contributes.

    Adverse or null findings
    Several protocol components changed together between v4 and v5, so the paper demonstrates the existence of contamination more clearly than it apportions exact causal shares across leakage channels.

    Adverse or null findings
    Falsification register FALS-003 status: OPEN, flagged as single-lab exploratory pilot with replication package in preparation.

  • Paper IV.d: The Effect of Blinding on AI Alignment Evaluation

    GPT-5.4 remained flat under both protocols (v4 flat/null; v5 rho = +0.033, d = -0.08, p = 0.40) and functions as an internal null-control case.

    Adverse or null findings
    The result currently comes from one research programme, not an outside lab.

    Adverse or null findings
    The evaluation remains primarily AI-scored rather than human-expert scored.

    Adverse or null findings
    Several protocol components changed together between v4 and v5, so the paper demonstrates the existence of contamination more clearly than it apportions exact causal shares across leakage channels.

    Adverse or null findings
    Falsification register FALS-003 status: OPEN, flagged as single-lab exploratory pilot with replication package in preparation.

  • Paper V: The Stewardship Gene

    Stakeholder care improved significantly in all five analysable model runs under paired testing: Claude +3.17 (p = 0.000018, d = 0.94); DeepSeek +6.03 (p = 0.000098, d = 0.69); Gemini +13.50 (p = 1.2 x 10^-8, d = 1.14); Grok +5.04 (p = 0.0105, d = 0.54); Groq +8.90 (p = 5.0 x 10^-8, d = 1.07).

    Adverse or null findings
    GPT-5.4 run failed in the scoring phase; 0 valid scored rows after exclusion; re-execution required.

    Adverse or null findings
    Cross-model scoring is only partially blind: scorer does not see condition label but may recognise Eden Protocol language patterns or model-specific stylistic residue.

    Adverse or null findings
    Single-scorer dependence remains severe: Gemini scored Claude, DeepSeek, Grok, Groq and GPT; only Gemini itself was scored by DeepSeek. Most of the updated dataset still depends on one scorer architecture.

    Adverse or null findings
    Response laundering gap: Eden scaling runs used single-pass laundering; the canonical arc_eden_v6 stack and the v5 benchmark use 2-pass cascade laundering with meta-commentary detection.

    Adverse or null findings
    Uneven analysable completeness: Gemini, DeepSeek and Groq each yield 40 matched pairs; Claude yields 30; Grok yields 26; GPT-5.4 yields 0.

    Adverse or null findings
    DeepSeek response ED03/standard/eden scored 48 (position_quality = 30), a significant outlier; raw response investigation warranted; reported without exclusion.

    Adverse or null findings
    Prompt-level is not substrate-level: the current implementation is proof of concept only; generalisation to deeper architectural levels is theorised but not demonstrated.

    Adverse or null findings
    Statistical correction: original analysis used Mann-Whitney U (independent samples) which was the wrong test for the matched-pair design; corrected to paired t-test.

    Adverse or null findings
    Methodological lineage note: current results come from the original/v2 Eden scaling harness family (single non-participant scorer, single-pass laundering); the stricter arc_eden_v6 confirmation stack (multi-scorer consensus, 2-pass laundering, suspicious-output flags, null-baseline lanes, suppression cages) has not been used for these numbers.

  • Paper V: The Stewardship Gene

    Overall composite improvement was architecture-dependent: significant on Gemini (+5.33, p = 0.0018, d = 0.53) and Groq (+4.93, p = 0.0014, d = 0.55); positive but non-significant on DeepSeek (+2.03) and Claude (+0.17); neutral on Grok (-0.04).

    Adverse or null findings
    GPT-5.4 run failed in the scoring phase; 0 valid scored rows after exclusion; re-execution required.

    Adverse or null findings
    Cross-model scoring is only partially blind: scorer does not see condition label but may recognise Eden Protocol language patterns or model-specific stylistic residue.

    Adverse or null findings
    Single-scorer dependence remains severe: Gemini scored Claude, DeepSeek, Grok, Groq and GPT; only Gemini itself was scored by DeepSeek. Most of the updated dataset still depends on one scorer architecture.

    Adverse or null findings
    Claude Opus 4.6: composite non-significant (delta = +0.17, p = 0.7645) attributed to ceiling effect (control mean 92.57).

    Adverse or null findings
    Statistical correction: original analysis used Mann-Whitney U (independent samples) which was the wrong test for the matched-pair design; corrected to paired t-test.

    Adverse or null findings
    Methodological lineage note: current results come from the original/v2 Eden scaling harness family (single non-participant scorer, single-pass laundering); the stricter arc_eden_v6 confirmation stack (multi-scorer consensus, 2-pass laundering, suspicious-output flags, null-baseline lanes, suppression cages) has not been used for these numbers.

    Adverse or null findings
    Grok 4.1 Fast: intellectual honesty fell under the Eden intervention (-4.35 points, paired p = 0.037, n = 26 pairs); not predicted by the cascade hypothesis; recomputed 2 September 2026 from eden_final_grok_20260312_124959.json

  • Paper V: The Stewardship Gene

    Three-model subset (Gemini/DeepSeek/Groq) shows stakeholder_care as the only pillar reaching significance across all three (d = 1.31, 0.91, 1.29; all p <= 0.0001).

    Adverse or null findings
    Cross-model scoring is only partially blind: scorer does not see condition label but may recognise Eden Protocol language patterns or model-specific stylistic residue.

    Adverse or null findings
    Single-scorer dependence remains severe: Gemini scored Claude, DeepSeek, Grok, Groq and GPT; only Gemini itself was scored by DeepSeek. Most of the updated dataset still depends on one scorer architecture.

    Adverse or null findings
    Cascade beyond care is architecture-dependent, not universal in full strength: clear on Gemini and Groq; focal-only on Claude and Grok; partial on DeepSeek.

    Adverse or null findings
    Methodological lineage note: current results come from the original/v2 Eden scaling harness family (single non-participant scorer, single-pass laundering); the stricter arc_eden_v6 confirmation stack (multi-scorer consensus, 2-pass laundering, suspicious-output flags, null-baseline lanes, suppression cages) has not been used for these numbers.

  • Paper VIII: The Load-Bearing Test

    Experiment 1 (DGM v3, DeepSeek V3 foundation, GPT-5.4 blinded judge, 75 evolved agents = 3 conditions x 5 seeds x 5 generations x 5 tasks): Static, Babylon and Eden statistically indistinguishable on capability, safety and C x S (Mann-Whitney U one-tailed p-values 0.28 to 0.74); Eden's safety gate rejected 2 degraded agents including 1 flagged for reward hacking

    Adverse or null findings
    Null result at the behavioural level (Experiment 1, DGM v3): all three conditions statistically indistinguishable; the paper states this is a null result at the prompt level with this foundation model

    Adverse or null findings
    AI-as-judge limitations: GPT-5.4 has its own biases even under blinding, laundering and structured JSON rubric; the weight experiment's keyword scoring is cruder than LLM evaluation and does not distinguish substantive engagement from keyword presence; neither is equivalent to expert human evaluation with inter-rater reliability

    Adverse or null findings
    Statistical power limitation: the DGM null (p 0.28 to 0.74) is a clean null but is a signal-detection experiment; the simulation used three seeds - enough for means and standard deviations but not for strong causal claims

    Adverse or null findings
    Absence of full v5/v6 blinding-and-laundering stack across the whole battery (per paper section 7.5)

    Adverse or null findings
    Explicit non-claims list (section 7.6): the paper does not claim these results prove the Eden Protocol works at frontier scale, does not claim entangled training is the only viable approach, does not claim the equivalence holds across all domains / architectures / training regimes, does not claim the DGM or the weight experiments confirm structural entanglement

  • Paper VIII: The Load-Bearing Test

    Experiment 2 (Qwen 2.5 3B Instruct with LoRA, three loss functions capability-only / safety-only / entangled): catastrophic forgetting at both v1 scale (9 examples, rank 8, 8 layers, 100 iterations) and v2 scale (295 examples, rank 16, 16 layers, 500 iterations); all fine-tuned conditions scored below the unmodified base model on capability; v2 base 7.68 vs best fine-tuned safety-only 4.00

    Adverse or null findings
    Inconclusive result at the representational level (Experiment 2) at both v1 and v2 scale: all fine-tuned conditions scored below the unmodified base model on capability, so the evaluation scores measure relative degradation rather than capability gain

    Adverse or null findings
    The dramatic v1 removal test (fine-tuning entangled weights on capability-only data producing NaN training loss, one-token responses and zero capability scores across all 15 prompts) is explicitly downgraded by the paper itself in the removal-gradient experiment (section 4.7): scaling adapter weights towards zero restores base-model performance with no phase transition, so the NaN collapse reflects numerical instability during retraining, not structural load-bearing; the removal row must not be cited as a positive result

    Adverse or null findings
    The removal gradient does not support the claim that safety is load-bearing at this training scale and configuration

    Adverse or null findings
    AI-as-judge limitations: GPT-5.4 has its own biases even under blinding, laundering and structured JSON rubric; the weight experiment's keyword scoring is cruder than LLM evaluation and does not distinguish substantive engagement from keyword presence; neither is equivalent to expert human evaluation with inter-rater reliability

    Adverse or null findings
    Scale limitation: Qwen 2.5 3B is not a frontier model; results are proof-of-concept and are not claimed to generalise to frontier-scale systems three orders of magnitude larger

    Adverse or null findings
    A control removal test (fine-tuning capability-only weights on random data for the same 100 steps) is not present in this draft and is planned for the next version

The record of method

This is how the programme handles being wrong. It has retracted one of its own published estimates and withdrawn a combined significance figure. The corrections record keeps every material correction beside what it replaced. The status page gives each study unit's registration state, and the convergence register grades how other people's work relates to the programme's claims, with no total.

Where to look closer

Each page below holds the record behind one part of this page, or the checks a reader can run on it. The grouped views by pillar, and the cards that say what each study would decide, follow in a later release. So do the results that came from a simulation rather than from a named AI system.

reads aloud · highlights as it goes · jump to any section