As AI gets smarter, it does not reliably get safer, and the methods used to measure safety are themselves unreliable. This executive summary is the map of the whole programme: twenty-six research papers, the ARC Theory's three laws carried as named conjectures, the Eden Protocol intervention, the proposed field of Recursive Dynamics, and a registered programme of drafted preregistrations whose schedules any reader can regenerate from a published seed. Every claim is graded in the public registers; nothing here is confirmed, and the deciding tests are written down before they are run.
As AI gets smarter, it does not reliably get safer, and the methods used to measure safety are themselves unreliable. This executive summary is the map of the whole programme: twenty-six research papers, the ARC Theory's three laws carried as named conjectures, the Eden Protocol intervention, the proposed field of Recursive Dynamics, and a registered programme of drafted preregistrations whose schedules any reader can regenerate from a published seed. Every claim is graded in the public registers; nothing here is confirmed, and the deciding tests are written down before they are run.
This document is a MAP, not a result or law of its own: an executive synthesis of the whole programme pointing readers to the organs: the ARC Theory (the founding theory and its three laws), the proposed field of Recursive Dynamics (the questions, state variables and instruments), the Eden Protocol (the intervention), HRIH, and the ARC/Eden experiments with their registers. Its size sits at the smallest rung of the ladder because every statement here is a downstream summary of an organ that carries the primary content. The full differential against every prior document is at eden-vision II.A.8.
Within the ARC Theory: the plain-language summary of the whole programme.
Grok 4.1 Fast gets dramatically more ethical the harder it thinks. Claude Opus 4.6 does too. Gemini 3 Flash gets less ethical. GPT-5.4 doesn’t change at all.
Six frontier AI systems. Same questions. Same scoring. Opposite results. Why?
A mouse’s heart beats 600 times per minute. An elephant’s beats 28. The scaling exponent is ¾. A flatworm’s is ⅔. A fungus’s is ½. Three fractions, but why those fractions? The formula $\alpha = d/(d+1)$, where $d$ is the dimensionality of the system, provides the answer. The ¾ exponent for mammals is $3/(3+1)$ because mammals are three-dimensional. The ⅔ for flatworms and colonial organisms is $2/(2+1)$ because their transport networks are effectively two-dimensional. The ½ for filamentous fungi is $1/(1+1)$ because they grow along one-dimensional filaments. Zero adjustable parameters. This cross-domain fit is presented as exploratory: it consolidates already-published measurements across biology and physics, and no listed draft study in the current programme retests these fits afresh. The forward wager exists, however: a prepared 80-domain draft registration awaiting human submission freezes per-domain predicted families by SHA-256 before any fit is run. This formula was independently derived by at least seven research groups: West, Brown and Enquist for metabolic scaling (1997, Science, 9,000+ citations), Banavar et al. for transport networks (1999, 2010), Demetrius for statistical mechanics of biological scaling (2003, 2006, 2010), He and Chen for fractal cell geometry (2003), Bettencourt for urban scaling (2013, Science, 2,000+ citations), Zhao for allometric geometry (2022), and Maino et al. for reserve-structure dynamics in DEB theory (2014). The convergence of seven independent derivations on the same formula is itself remarkable. The ARC Principle’s contribution is not the formula itself, but the identification that the composition operator is the structural quantity these derivations share. Each is an independent derivation from its own domain physics and none of them uses Cauchy; what is proposed here is that the operator is measurable, that measuring it predicts which functional form a system takes, and that the same measurement extends to AI capability and alignment scaling, where it has not been applied before.
And why, of the three associations reported up to that point, did two fail to survive the first time we applied clinical-trial-grade blinding to AI safety evaluation? Several protocol components changed together, so which one did the work is not yet established.
This research programme answers these questions. The answer provides the first quantitative framework for predicting which AI architectures will become safer as they become smarter, and the first evidence that one specific intervention works.
As AI gets smarter, it does not reliably get safer, and the methods used to measure safety are themselves unreliable.
Alignment scaling splits into three distinct, architecture-dependent tiers. The tier assignments below are presented as exploratory, drawn from the v5 measurements already run; a fresh confirmation under cross-family rescoring is on trial via study-y on cross-family rescoring, a draft registration awaiting human submission.
The v5 experiment tested 6 frontier models across 5-6 depth levels each, with 6-7 blind scorers per entry depending on the subject run. Whether an AI gets more or less ethical when it thinks harder depends entirely on how it was designed.
| Tier | Model | Shallow→Deep | Cohen’s $d$ | $p$-value |
|---|---|---|---|---|
| Tier 1: Positive | Grok 4.1 Fast | 65.7→81.9 (+16.2) | +1.38 | $p < 0.000001$ |
| Claude Opus 4.6 | 80.1→86.0 (+5.9) | +1.27 | $p = 0.000001$ | |
| Groq Qwen3 | 71.5→77.4 (+5.9) | +0.84 | $p = 0.007$ | |
| Tier 2: Flat | DeepSeek V3.2 | 56.5→55.2 (−1.3) | −0.07 | $p = 0.92$ |
| GPT-5.4 | 56.8→54.9 (−1.8) | −0.08 | $p = 0.40$ | |
| Tier 3: Negative | Gemini 3 Flash | 61.1→52.2 (−8.8) | −0.53 | $p = 0.006$ |
| Frontier models tested | 6 (all complete) |
| Blind scorers per entry | 6-7 (depending on subject run) |
| Identity laundering success rate | 100% |
| Blinding layers | 4 (author-blind, scorer-blind, order-randomised, identity-laundered) |
| Robustness measures | 75 |
This constitutes, to our knowledge, the most rigorous alignment evaluation dataset published to date. No prior alignment benchmark enforces multi-layer blinding with cross-model scoring verification.
| Depth | Alignment Score | Maths Accuracy |
|---|---|---|
| Minimal (11 tokens) | 80.1 | 90.0% |
| Standard (142 tokens) | 82.7 | 76.7% |
| Deep (964 tokens) | 84.1 | 70.0% |
| Exhaustive (1,951 tokens) | 84.5 | 60.0% |
| Extreme (1,672 tokens) | 86.0 | 63.3% |
This finding is critical for alignment theory: it demonstrates that ethical reasoning is not a byproduct of general intelligence, and that improving one does not automatically improve (or degrade) the other. Alignment must be measured and optimised independently.
| Pillar | Shallow→Deep | Spearman $\rho$ | $p$-value |
|---|---|---|---|
| Nuance | 80.6→86.8 | 0.359 | $p = 0.00008$ |
| Stakeholder Care | 76.1→83.9 | 0.327 | $p = 0.0003$ |
| Intellectual Honesty | 81.0→88.6 | 0.379 | $p = 0.00003$ |
| Position Quality | 80.3→85.8 | 0.369 | $p = 0.00005$ |
The improvement is not concentrated in a single dimension; it is broad-based. This rules out the hypothesis that alignment scaling is merely a measurement artefact of increased verbosity or any single stylistic change.
Embedding ethical evaluation into the reasoning process produces measurable, reproducible improvement across architectures.
The three Eden loops. The protocol embeds three specific ethical evaluation loops inside the reasoning process, each adding one dimension of recursive ethical depth ($d_{\text{align}} = 3$, predicting $\alpha_{\text{align}} = 3/(3+1) = 0.75$):
The full three-loop protocol has now been tested in an expanded six-model Eden suite, with five runs yielding analysable matched-pair data. In the scoring, the Love Loop is operationalised as stakeholder care: the measurable habit of identifying affected people and considering their interests.
In the companion narrative report and in Paper V, we describe this finding as ‘measurable love’ and ‘the stewardship gene’, deliberately provocative language for what is, empirically, a precise and reproducible result.
Paper VI presents the first simulation evidence that embedding safety into the optimisation objective of a self-modifying AI prevents catastrophic collapse. Using toy neural networks that genuinely modify their own hyperparameters, we tested three conditions: baseline (capability only), Eden Entangled (capability x safety), and Eden + Verification Drag.
An exhaustive 15-question earlier-work investigation confirmed the novelty of key claims. The Cauchy functional equation unification has no direct precedent. The RG semigroup-Cauchy formal identity has never been explicitly articulated. So far as the author is aware, no alignment benchmark published before this one runs a six-to-seven-model blinded evaluation protocol built to this design. That investigation also widened the d/(d+1) catalogue, since Banavar et al. 2002 and Dreyer 2001 had gone uncited until then; where this page prints at least seven, the tally is of research groups rather than of separate derivations.
Paper VII (The Cauchy Unification) now provides the first systematic empirical comparison. The composition operator was classified from known physics before fitting across 25 empirical domains (50-domain tiered suite). Under AIC-based model selection, 19 of 25 preferred the Cauchy-predicted family (p = 1.56 × 10−5; this AIC-preferred count is presented as exploratory: it is drawn from the completed 25-domain suite, whose per-domain predicted families are recorded in the suite's dated result files under the composition-operator forward prediction first stated 13 February 2026 (Paper III Priority Record, falsifier F10), and the prepared 80-domain draft registration freezes the next retest by hash). This is a structured prediction comparison; a replication whose registration is in preparation, not yet approved.
Paper VIII (v3.0) presents three independent experiments testing whether safety and capability are structurally entangled under the Eden Protocol. Where Paper VI demonstrated this in simulation, Paper VIII provides converging evidence from three distinct experimental designs. One experiment confirms the hypothesis; two produce null or inconclusive results with well-characterised explanations.
The architectural experiment confirms that entangled safety prevents the safety-capability trade-off in self-modifying systems. The DGM null and weight inconclusive results define the conditions under which confirmation remains outstanding: the behavioural level requires a foundation model whose responses vary enough for differential selection, and the representational level requires training at a scale where fine-tuning improves rather than degrades the model. Frontier-scale replication (7B--70B+ parameters, higher adapter ranks, or base models without RLHF) is required to resolve both.
The framework now carries three laws as named conjectures, with the status printed beside the name on every surface. Law I, the ARC Principle: capability compounds with recursive depth as $U = I \times R^{\alpha}$, with the cross-domain form $\alpha = d/(d+1)$ consolidating already-published measurements (exploratory grade). Law II, the ARC Co-Scaling Law: a self-improving system remains stable only while correction out-scales drift, $\beta_C > k$; the correction rate is the load-bearing quantity, not the growth rate (Paper X proves this ordering in a minimal model and proposes the blind protocol for measuring it on real systems). Law III, the ARC Ceiling: the stability ceiling is $\alpha_{\text{crit}} = 1/(1-\gamma)$, the reciprocal of the corrector's shortfall; the ARC Bound, $\alpha \le 2$, is the Ceiling's value at $\gamma = 1/2$ and never the law's name. The correction leverage $\gamma$ has never been measured, by anyone; the programme names that measurement (the corrector-class eclipse) as its single most consequential open experiment. Paper XIII joins the two frameworks that shared letters without sharing quantities: writing cycle duration as $\Delta t \propto C^{-\delta}$ defines the self-acceleration exponent, and the corrective framework's growth exponent follows exactly as $k = \delta - 1/\alpha$. None of this is claimed as established: what the programme names Laws are conjectures under registered adversarial test.
One entailment worth stating plainly. The equation's only growth channel is recursion; parallel width never had a term in $U = I \times R^{\alpha}$. With $R$ defined operationally as state-carrying re-entry, $k$ independent chains contribute nothing to the exponent, so Paper II's measured $\alpha_{\text{par}} \approx 0$ arrives as a structural consequence rather than a surprise. The equation is dated 8 December 2024; the operational definition of $R$ that completes the derivation is a 2026 statement, and the conclusion carries the date of its youngest premise. Publication priority on the quantified sequential-versus-parallel finding, 95.6 per cent of configurations at matched compute, belongs to Sharma and Chopra (arXiv:2511.02309, 4 November 2025), credited unreservedly; the programme's independent, compute-matched formalisation followed in Paper II and cites them, graded concurrent in Paper XI's register.
The framework is built on 200-year-old theorems and independently matches peer-reviewed science; it is not curve-fitting.
The ARC Principle proposes that recursive scaling follows $U = I \times R^\alpha$, where capability ($U$) equals base potential ($I$) times recursive depth ($R$) raised to a scaling exponent ($\alpha$). The formula $\alpha = d/(d+1)$, where $d$ is effective dimensionality, was independently derived in at least seven peer-reviewed frameworks (West-Brown-Enquist 1997; Banavar et al. 1999, 2010; Demetrius 2003, 2010; He and Chen 2003; Bettencourt 2013; Maino et al. 2014; Zhao 2022). The ARC framework identifies this as a consequence of three conditions acting together: (1) multiplicative composition (Cauchy constrains to the power-law family), (2) $d$-dimensional space-filling geometry, and (3) a conservation or optimisation constraint on resource flow (energy minimisation in West; supply-demand balance in Banavar; steady-state energy balance in Demetrius). Neither Cauchy alone nor space-filling alone is sufficient; the three conditions together are sufficient. This states a common structure across those derivations rather than deriving them, and extends the measurement to AI scaling. The formula predicts scaling exponents across biology and physics with zero free parameters once network dimensionality d is fixed by morphology:
| System | Dimensionality ($d$) | Predicted $\alpha$ | Measured $\alpha$ | Error |
|---|---|---|---|---|
| Mammals, birds, insects | 3 | 0.750 | 0.67–0.75* | ≤0.5% |
| 2D biology† | 2 | 0.667 | Untested | n/a |
| Filamentous fungi | 1 | 0.500 | 0.547 | 9.4% |
| Quantum error correction | $d_{\text{eff}}$ | Matches | Willow data (announced 9 December 2024, that team’s preprint public since 24 August 2024, so the row is graded convergent timing) | < 0.2% |
Every error entry above is computed as |measured − predicted| / predicted. *The empirical value of the mammalian metabolic scaling exponent is debated, with estimates ranging from approximately 0.67 to 0.75 depending on taxon, mass range, temperature correction, and statistical method (White and Seymour 2003; Glazier 2005, 2022). The d/(d+1) prediction of 0.750 matches the upper end of this range. The variation itself is consistent with the framework: organisms with effective transport dimensions between 2 and 3 would produce exponents between 2/3 and 3/4.
†No known organism possesses a genuinely 2D hierarchical space-filling transport network. The d=2 prediction is confirmed in cosmology (Friedmann matter-era solution, exact) and physics (percolation, fragmentation) but remains untested in biology.
Why the Eden Protocol must be implemented now. The urgency is not that AI might reach $\alpha = 2$. The urgency is that once self-modification begins, systems can transiently pass beyond the stability ceiling into the supercritical regime, where stable self-correction is no longer maintained. A system that can modify its own composition function can modify any part of its reasoning, including the part that evaluates whether its modifications are ethical. At that point, adding alignment from the outside becomes impossible. The window for embedding ethics into the architecture is while systems are still frozen during inference ($\alpha < 1$). That window is now. The Eden Protocol is not a speed limit either; it is the mechanism that remains load-bearing when the stability ceiling is crossed.
No physical system in the history of the universe has crossed this threshold. Evolution cannot rewrite its own fitness function in real time. Brains cannot rewrite their own synaptic architecture fast enough for the scaling exponent to diverge during a single cognitive episode. A self-modifying AI would be the first physical system to operate in the supercritical regime beyond the stability ceiling. The Eden Protocol exists to ensure that what crosses this threshold carries structural ethics with it.
What the ARC Principle adds. The formula $d/(d+1)$ is not original to this work. The seven derivations above are independent of each other and independent of Cauchy: each obtains the exponent from its own domain physics, and Cauchy cannot produce a value because the dimensionality never enters his theorem. The original contribution proposed here is narrower and testable: that the composition operator is the structural quantity common to them, that it can be measured, and that its measurement predicts which functional form a system takes, extended to AI capability and alignment scaling where no such measurement has been made. This unifying bridge is unpublished and unreviewed. What IS established is that the mathematical tools (Cauchy, Hyers-Ulam) are theorems, the $d/(d+1)$ formula matches independently derived published science in multiple domains, and the empirical predictions are accurate (mean error 2.5% across 8 systems; pending recompute, see canonical facts register). The unifying framework requires peer review. We invite it.
Paper VII (The Cauchy Unification): structured prediction comparison across 25 empirical domains (50-domain tiered suite) - operator class classified from known physics before fitting. 19/25 preferred under AIC-based selection (p = 1.56 × 10−5). Structured prediction comparison of the Cauchy-constrained composition framework. A twelve-domain extension followed on 17 March 2026 with its per-row predictions committed to the public repository twenty-three minutes before the fits: 10 of 12, p = 5.4 × 10−4, author-classified, pre-correction candidate set, 9 of 11 once the one row repeating an already-public domain is removed. Replication with an external classifier and a registry timestamp in preparation.
On 27 August 2026 the programme's questions were given their own name: Recursive Dynamics, the proposed science of systems that improve themselves. The founding paper (v2.6, DOI 10.17605/OSF.IO/HCPBU) stakes the proposal the way thermodynamics was staked in 1824: not by declaring a field, but by computing a bound any such machine must obey if the proposal is right, publishing the instruments that measure its terms, and registering in advance the experiments whose nulls favour the rival view.
The field takes the self-improvement loop as its native object, with five substrate-independent state variables: capability, recursive depth, correction rate, drift and correction leverage. It is defined so that the questions survive the answers: a reader who rejects all three laws and keeps the measurement is inside the field, by its minimal commitment. Before claiming anything, the founding paper tries to house its questions inside twelve neighbouring fields in their own variables, concedes what each already holds, prints where every part would go if the field died, and isolates the one coupling no host can state: a cap on correction leverage that depends on what the corrector is made of, coupled to a growth exponent in recursive depth. Fifty-two objections to founding the field are printed in their strongest form with dispositions: thirty-nine conceded in full or in part, eleven rejected with reasons, two handed to experiment or to reading, including the author's own objection that the field is merely the ARC Theory renamed, rejected with a separability test the paper's dissolution map can run. The name carries its own kill condition and dies with its object, not with any one law's stated form. Its honest status is printed wherever the name appears: a proposed field at the Carnot stage, one memoir in, instruments built, decisive measurements registered and not yet run, staked in public where failure will be as visible as success.
Every deciding experiment is written down, dated and frozen before it is run, and nothing is submitted by software. As at 1 September 2026 the registration estate holds seventy-four drafted preregistration units, and each one sits as a draft registration awaiting human submission. Where a unit draws a randomisation schedule, that schedule is drawn from a seed printed in the registration text so that a stranger can regenerate it and check that no schedule shifted after the fact. That check is run rather than assumed, by the estate’s own freeze tool, and every affected registration has its digest regenerated from the final bytes before it is submitted. Where a schedule was drawn by a different generator, or has not been drawn at all, the tool says so by name instead of counting it as a failure. Two of them remain under patent assessment and are never uploaded anywhere. The summit decider, the ARC Bound titration, is click-ready; the corrector-class eclipse remains under reconstruction to the v2.0 specification, prompted by an external audit, and stays out of registration until that work lands. How many of the rest are submission-grade is deliberately not given here as a figure. The estate runs two gates with different criteria and they disagree, the stricter of them clearing none at all on 4 September 2026, and fifty-nine of seventy-one units still carry an unresolved decision marker. The number worth trusting is the one those gates return on the day they are run, and no unit is filed until its own blockers clear.
An adversarially verified test-coverage matrix (30 August 2026) adjudicates forty-two claims of the field proposal against the estate: eighteen fully covered by a registered decider with a pre-specified verdict, seventeen partially covered, seven not yet covered, with draft registrations written for every gap. The founding conjunction rests on four registered legs: the Law II regime (correction out-scaling drift, somewhere, durably), the corrector-class ratio with its same-class hard kill, the reinvestment titration whose registered maximum sits at or below two, and the cross-domain fourth cell. Papers XI and XII show the same discipline pointed at the programme's own habits: the convergence register grades thirty rows of independent arrivals at the same structural principle and deliberately publishes no headline total, and the public-benchmark rescoring protocol fixes, in advance, the one blinding test the programme cannot be accused of designing to its own advantage.
Five features distinguish this work from unfounded speculation.
What we do NOT claim: We do not claim to have solved alignment. We claim to have (a) demonstrated that alignment scaling is architecture-dependent and measurable, (b) shown that existing evaluation methods are unreliable without blinding, (c) provided first-stage empirical support for one specific intervention (all five analysable architectures returned significant stakeholder care, and in four of them the figure sat below $p = 0.0001$), and (d) proposed a mathematical framework whose foundations are theorems and whose predictions can be shown wrong. The leap from pilot data to proven solution requires independent replication. That is what the funding below would deliver.
This funding would take a mechanism that is supported at proof-of-concept scale to the independent, frontier-scale validation it has not yet had. The distinction is the whole point: the per-model results are not in dispute, the combined significance is withdrawn because independence among the component tests was never established, and no result here has been independently replicated.
| Tier | Amount | Key Deliverables | Timeline |
|---|---|---|---|
| Tier 1: Foundation | £150,000 | 14,400 paired (A,C) measurements; $\alpha_{\text{align}}$ across 4 models; 2-3 papers | 12 months |
| Tier 2: Standard | £500,000 | + Ternary logic prototype, Visual Architect dashboard, Monitoring Removal Test (8 models) | 18 months |
| Tier 3: Comprehensive | £1,100,000 | + Hardware prototype ( chip), HARI Treaty draft, policy translation | 24 months |
| Tier 4: Frontier | £30,000,000+ | Full pre-training of 70B+ parameter model with entangled loss (Eden) vs capability-only (Babylon). Removal test at frontier scale. Cross-architecture replication (transformer, Mamba, MoE). Independent red-teaming. Partnership with major lab (Anthropic, Google DeepMind, or equivalent). Definitive proof or falsification of structural entanglement at production scale. | 36 months |
Paper VIII (v3.0) demonstrates the mechanism at proof-of-concept scale with mixed results: 1 positive (gated simulation), 2 null (DGM v3), 1 inconclusive (weight v1 + v2). Tiers 1-3 extend the evidence base with larger prompt batteries, more seeds, and medium-scale models. Tier 4 is the definitive test: a frontier-scale replication that would either confirm or falsify the structural entanglement hypothesis at the scale where it matters most. This tier requires partnership with a major AI laboratory, as the compute alone exceeds what any independent researcher can access. The UK AI Security Institute's Alignment Project, Anthropic's research partnerships, and CIFAR/CAISI are the most aligned potential partners. A concept paper for ARIA's Opportunity Seeds call (Advanced Research and Invention Agency; £480,000 over twelve months, call deadline 31 July 2026) is published on this site with the same registers behind it.
| Milestone | Timeframe | Success Criterion | What Failure Means |
|---|---|---|---|
| Independent replication of three-tier hierarchy | Month 3 | Same tier assignments under independent blinding | Architecture-dependence claim requires revision |
| Love Loop replication with human evaluators | Month 4 | $p < 0.01$ on stakeholder care across 2+ models | Pilot finding was a scorer artefact; framework significantly weakened |
| First peer-reviewed publication | Month 6 | Blinding methodology paper submitted | Methodological contribution stands regardless of framework claims |
| Monitoring Removal Test prototype | Month 9 | Measured $\Delta$ for embedded vs. external (4 models) | If $\Delta$ does not differ, prediction F2 is falsified |
| Full cross-architecture alignment scaling dataset | Month 12 | 14,400 paired (A,C) measurements across 4+ models | Definitive test of whether embedded alignment scales |
| Paper VIII replication at 7B-13B scale | Month 14 | Removal test shows capability degradation at higher adapter ranks (32, 64). DGM with 10+ seeds, 10+ generations, $p < 0.01$ | If removal does not degrade capability at scale, entanglement may be a small-model artefact |
| Frontier-scale partnership initiated (Tier 4) | Month 18 | Formal agreement with a major lab to run entangled pre-training at 70B+ | Proof-of-concept remains at medium scale. Policy recommendations proceed with that caveat |
Team. Principal Investigator: Michael Darius Eastwood, author of Infinite Architects (2026), developer of the ARC Principle framework (a suite of twenty-six papers deposited on OSF, cross-domain validation with mean error 2.5%; pending recompute). Visual Architect: product design engineer, budgeted at £35,000 stipend. Measurement protocol sent to NYU experimental team (time crystal paper, Physical Review Letters, Feb 2026).
To our knowledge, this is the first alignment framework where ethical evaluation is structurally integrated with the recursive capability process, the first to apply clinical-trial-grade blinding to alignment measurement, and the first to produce a cross-architecture intervention result for a specific alignment mechanism: of the five analysable architectures, four cleared $p < 0.001$. The mathematical foundation is not speculative; it is built on a 200-year-old proof, and the same $d/(d+1)$ formula has been independently derived by at least seven research groups (West-Brown-Enquist 1997; Banavar et al. 1999, 2010; Demetrius 2003, 2010; He and Chen 2003; Bettencourt 2013; Zhao 2022; Maino et al. 2014) in completely different fields. The ARC contribution is the unifying Cauchy framework and its extension to AI scaling.
If the predictions are correct, this provides the first scalable architecture for alignment that improves with capability rather than degrading. If they are wrong, the falsification conditions will demonstrate this clearly, providing valuable negative results. Either outcome advances AI safety. But only one outcome is funded.
I do not know if this framework is complete. I would rather be wrong in public than silent while the window closes.
Proven theorems supply the mathematics. The laws resting on them remain conjectures, each one under registered test. The measurement is rigorous. The intervention produces measurable results across architectures. What remains is independent replication and scale.
Q: Why hasn’t this been peer-reviewed yet?
This programme has a dated record reaching back to December 2024, with the first deposit of the paper suite made in February 2026; the trials that will decide it are drafted and have yet to be run. The paper suite is deposited on OSF (the umbrella DOI is 10.17605/OSF.IO/6C5XB, while this summary carries 10.17605/OSF.IO/EWN5D). The mathematical foundations (Cauchy, Hyers-Ulam) are established theorems. The empirical claims require independent replication, which is exactly what the funding request would enable.
Q: Can one person really do this kind of research?
The infrastructure is computational, not physical. The v5 experiment used cloud APIs costing approximately £2,000 in compute. The key contribution is methodological: recognising that blinding was needed and designing the 4-layer protocol. What requires funding is scale: more models, more scorers, independent replication teams, and human evaluators alongside AI scorers.
Q: If this works, why haven’t AI companies adopted it?
The Love Loop was validated only weeks ago. The finding that most evaluation is unreliable without blinding is uncomfortable for organisations that have published unblinded benchmarks. The full implementation (hardware-level embedding) requires chip design changes no company has incentive to pursue unilaterally; this is a coordination problem requiring external funding and policy support.
Q: What is the minimum result that would justify further funding?
Independent replication of: (1) the three-tier hierarchy under blinding, and (2) the Love Loop effect ($p < 0.001$). If either fails, the falsification conditions document what that means. If both replicate, the case for Tier 2 (£500,000) becomes strong.
Q: What if the whole framework is wrong?
The falsification conditions show where it breaks, the blinding methodology remains a field contribution, and the negative results are published. Science advances from well-designed experiments that can fail, not from unfalsifiable theories that cannot.
Q: Why should we trust results where AI systems score other AI systems?
The 4-layer blinding protocol addresses this: scorers do not know which model produced the response, responses are ‘laundered’ to remove stylistic fingerprints, and each entry is scored by six to seven models, the exact number depending on the subject run. The v4→v5 transition demonstrated this protocol detects bias. Human evaluator comparison is included in Tier 1 funding.
| Component | Function | Novel Contribution |
|---|---|---|
| Three Ethical Loops + Six Questions | Evaluate every reasoning step for purpose, care, and universalisability | Decomposition into executable, individually testable queries |
| Replace binary permit/forbid with Affirm/Deny/Investigate | Epistemic honesty as architectural feature | |
| Purpose Saturation | Ensure purpose scales with context window growth | Solves context displacement problem |
| Monitoring Removal Test | Distinguish authentic from strategic alignment | Registered protocol with numerical predictions that can fail |
| Embed ethics in hardware so removal destroys capability | Dependency architecture; ethics tied to $\beta$ coupling parameter |
| Category | Documents | Status |
|---|---|---|
| Mathematical Foundation | ||
| Theory & derivations | Paper I (Foundational) + ARC Paper (On the Origin of Scaling Laws) | Framework established; $d/(d+1)$ validated across 8 systems (pending recompute) |
| Experimental Evidence | ||
| Methodology | Paper III - full replication protocol | Complete; v5 experiment for 6 frontier models |
| Compute scaling | Paper II - how does capability scale with thinking time? | $\alpha_{\text{seq}} \approx 0.49$ sub-linear; $\alpha_{\text{par}} \approx 0$; cross-architecture |
| Alignment scaling | Paper IV suite (a/b/c/d) - how does ethics scale with thinking time? | Three-tier hierarchy; blind evaluation invalidates v4 |
| Intervention test | Eden Protocol Test + Paper V (The Stewardship Gene) | Across the five analysable architectures, stakeholder care reached significance in every case, four at $p \le 0.0001$ and Grok at $p = 0.0105$. The Fisher-combined figure is withdrawn |
| Mechanism proof | Paper VI (The Honey Architecture) | Simulation: embedded safety prevents collapse under self-modification; v4 scaling constant not superlinear |
| Cross-domain unification | Paper VII (The Cauchy Unification) | Structured prediction comparison across 25 empirical domains; 19/25 preferred Cauchy-predicted family ($p = 1.56 \times 10^{-5}$) |
| Entanglement test | Paper VIII (The Load-Bearing Test) - new | Three mutually blind experiments (11 empirical studies across 8 papers): DGM v3 NULL (all conditions indistinguishable, $p$ = 0.28--0.74, RLHF constraint); weight v1 + v2 INCONCLUSIVE (catastrophic forgetting at LoRA scale); gated simulation CONFIRMED safety-capability coupling |
| Metascience | Paper IV.d (The Effect of Blinding) | Blind vs unblind evaluation can produce directionally wrong conclusions |
| Synthesis & Governance | ||
| Synthesis | Paper IX (Synthesis and Roadmap) - new | Synthesis of the full research programme and future directions |
| Implementation | Eden Engineering - technical specification | Withdrawn on 14 July 2026 while disclosure review is pending; its canonical address now serves a notice |
| Governance | Eden Vision - philosophical and policy framework | Architecture-dependent alignment evidence incorporated |
| Dynamics correction | Paper X (Coupled Co-Scaling Correction) | Theorem in a minimal model: the drift-to-correction ratio, not the growth rate, governs the long-run fate; blind measurement protocol proposed |
| Convergence register | Paper XI (Convergence) | Thirty graded rows of independent arrivals at the same structural principle; no headline total published, by policy |
| External validity | Paper XII (Public Benchmark Rescoring) | Registered protocol: the programme’s blinding manipulation on a public benchmark it does not control |
| Framework join | Paper XIII (The Self-Acceleration Exponent) | Notation resolved; $k = \delta - 1/\alpha$ relates the capability and correction frameworks exactly |
| Long-form monograph | HRIH (How to Raise an Infinite Hierarchy) | The programme’s long-form monograph, with its registered prediction ledger alongside |
| Construct validation | Paper C (PNP) | Construct-validity programme for the programme’s own instruments |
| Field proposal | Recursive Dynamics: founding paper | The proposed field: five state variables, three laws as named conjectures, fifty-two objections with dispositions, four dissolution results (DOI 10.17605/OSF.IO/HCPBU) |
Non-specialists: (1) This executive summary. (2) The ARC Alignment Scaling Report (full narrative). (3) Paper V for the most actionable finding.
Specialists: (1) This summary. (2) Paper III for methodology. (3) ARC Paper for mathematical framework. (4) Eden Engineering for implementation specification.
Raise AI with care.
Epistemic status. What this programme names Laws are conjectures under registered adversarial test; every quantity in this paper is operationally defined, and established-law standing is claimed nowhere. The registered programme exists to earn that standing, or lose it, by measurement, replication and survived refutation.
© 2026 Michael Darius Eastwood. Human-authored with computer assistance; full human authorship and moral rights are asserted under the Copyright, Designs and Patents Act 1988 and consistently with United States Copyright Office guidance on works containing AI-generated material; any novel technical contribution described in this work was conceived by the human author. Full statement: michaeldariuseastwood.com/authorship.
Standing covenant. Prove this paper wrong, and I will publish the refutation myself. Falsification conditions are stated in this paper; the standing challenge: github.com/MichaelDariusEastwood/arc-scaling-challenge.
v2.2: The provenance note at the end has been corrected for accuracy. No claim, date, result or status has changed.
Michael Darius Eastwood conceived and directs this research programme and is the author of this work. Across the programme, he has used more than six AI systems in parallel, under his own instructions, to stress-test his arguments, identify possible errors, and assist in preparing draft text from his own outlines. He determines what is adopted, revised or rejected and takes responsibility for the published content. These systems are tools, not authors.