For funders and assessors
This page exists because assessors have limited time and unaffiliated applicants carry a justified burden of proof. Below: every headline claim mapped to its checkable artefact, then the standard objections answered in the order they usually arrive.
| Claim | Check it | Time |
|---|---|---|
| Unblinded alignment evaluation reversed its own direction on 2 of 6 models tested | Paper IV.d, results table §4 | 10 min |
| β > k stability criterion, six theorems, proved within the model | Paper X + independent test suite in the repo (test_theorems_independent.py, no shared code with the implementation) | 15 min |
| Fail-closed ternary gate: 48/48 tests, composition proven | Clone the repo, run pytest in the prototype | 5 min |
| One published retraction (α ≈ 2.24 → ≈ 0.49) | Paper II, correction stated in the abstract | 3 min |
| Two of three Paper VIII experiments returned nulls, reported as nulls | Paper VIII, abstract and §7 | 5 min |
| Priority: manuscript timestamped 8 December 2024, before Willow (9 Dec) and Anthropic's alignment-faking paper (18 Dec) | Raw .eml with Google server headers, SHA-256 f0d1f38f…, on OSF | 5 min |
| Business track record £40,403 → £624,000, zero external investment | Independently verified through a statutory Official Receiver review | stated basis |
| Every claim sits on an explicit evidential rung | Per-paper CLAIMS ledgers in the repo: proved-within-model / verified / synthetic / pilot / open | 5 min |
Correct, and stated first everywhere. Nothing rests on credentials: every claim resolves to an artefact you can check above. The corpus is offered as the thesis; the public record is the viva; the examiner is anyone with a terminal.
The strongest counter-evidence is the failure record, published unprompted: a retraction of the headline exponent, two nulls of three experiments, a withdrawn metric, and an apparent drift effect reclassified as a probable scorer-bias artefact the moment a compliant protocol erased it. Discipline maintained with no supervisor, reviewer or funder to answer to is the only kind an assessor can actually rely on.
Cranks do not publish kill-conditions, run adversarial red teams against their own keystone paper (24 objections raised, 21 confirmed and fixed), credit precursors by name (Ashby, Conant-Ashby, the scalable-oversight literature), or retract. The claim ladder places every statement on an explicit evidential rung and nothing is cited a rung above its place.
Autistic and ADHD, diagnosed in adulthood. The standard route selects at every stage for one linear sequence; this cognition moves across domains instead: music, then a bootstrapped company, then self-taught High Court litigation, then this corpus. The same trait that failed the filter produced the work. Several funders now say credentials are not their criterion; this body of work is what taking that sentence literally looks like.
Mitigations are structural: everything public and reproducible (the programme survives its author), independent from-scratch test suites, funded plans include an engineer to distribute tacit knowledge and an advisory board with successor authority, and replication is budgeted as an adversarial collaboration briefed to refute, not confirm.
Cross-family blinded scoring requires routing evaluation data through a competitor's model. No frontier laboratory is structurally free to do that at scale. An unaffiliated measurer is. That is the uncomfortable structural fact at the centre of this programme.
michael@michaeldariuseastwood.com · London, United Kingdom. Application documents are available on request in accessible formats; reasonable-adjustment needs (Equality Act 2010) are documented and can be evidenced.