For funders and assessors

You should be sceptical of an unaffiliated applicant. Here is how to check everything without trusting anything.

This page exists because assessors have limited time and unaffiliated applicants carry a justified burden of proof. Below: every headline claim mapped to its checkable artefact, then the standard objections answered in the order they usually arrive.

Claim-to-evidence map

ClaimCheck itTime
Unblinded alignment evaluation reversed its own direction on 2 of 6 models testedPaper IV.d, results table §410 min
β > k stability criterion, six theorems, proved within the modelPaper X + independent test suite in the repo (test_theorems_independent.py, no shared code with the implementation)15 min
Fail-closed ternary gate: 48/48 tests, composition provenClone the repo, run pytest in the prototype5 min
One published retraction (α ≈ 2.24 → ≈ 0.49)Paper II, correction stated in the abstract3 min
Two of three Paper VIII experiments returned nulls, reported as nullsPaper VIII, abstract and §75 min
Priority: manuscript timestamped 8 December 2024, before Willow (9 Dec) and Anthropic's alignment-faking paper (18 Dec)Raw .eml with Google server headers, SHA-256 f0d1f38f…, on OSF5 min
Business track record £40,403 → £624,000, zero external investmentIndependently verified through a statutory Official Receiver reviewstated basis
Every claim sits on an explicit evidential rungPer-paper CLAIMS ledgers in the repo: proved-within-model / verified / synthetic / pilot / open5 min

The objections, answered in order

"No PhD, no institution."

Correct, and stated first everywhere. Nothing rests on credentials: every claim resolves to an artefact you can check above. The corpus is offered as the thesis; the public record is the viva; the examiner is anyone with a terminal.

"Solo researchers overclaim."

The strongest counter-evidence is the failure record, published unprompted: a retraction of the headline exponent, two nulls of three experiments, a withdrawn metric, and an apparent drift effect reclassified as a probable scorer-bias artefact the moment a compliant protocol erased it. Discipline maintained with no supervisor, reviewer or funder to answer to is the only kind an assessor can actually rely on.

"Is this crank physics?"

Cranks do not publish kill-conditions, run adversarial red teams against their own keystone paper (24 objections raised, 21 confirmed and fixed), credit precursors by name (Ashby, Conant-Ashby, the scalable-oversight literature), or retract. The claim ladder places every statement on an explicit evidential rung and nothing is cited a rung above its place.

"Why isn't this person inside an institution?"

Autistic and ADHD, diagnosed in adulthood. The standard route selects at every stage for one linear sequence; this cognition moves across domains instead: music, then a bootstrapped company, then self-taught High Court litigation, then this corpus. The same trait that failed the filter produced the work. Several funders now say credentials are not their criterion; this body of work is what taking that sentence literally looks like.

"Single-person risk."

Mitigations are structural: everything public and reproducible (the programme survives its author), independent from-scratch test suites, funded plans include an engineer to distribute tacit knowledge and an advisory board with successor authority, and replication is budgeted as an adversarial collaboration briefed to refute, not confirm.

"Why does the reliable measurement need an outsider?"

Cross-family blinded scoring requires routing evaluation data through a competitor's model. No frontier laboratory is structurally free to do that at scale. An unaffiliated measurer is. That is the uncomfortable structural fact at the centre of this programme.

Contact and formats

michael@michaeldariuseastwood.com · London, United Kingdom. Application documents are available on request in accessible formats; reasonable-adjustment needs (Equality Act 2010) are documented and can be evidenced.

← Home · Research · Evidence spine