the formal-limits pair
Two formal results and one dated record
A peer-reviewed paper and a formal preprint argue, on mathematical grounds, that alignment verification faces structural limits. A manuscript dated 8 December 2024 recorded the same direction before either existed. This page states what that establishes and what it does not, in the register’s own grading, and then states exactly what these results are worth to this record, which is more than the grading alone suggests.
Paper one · peer-reviewed
Hernández-Espinosa, Abrahão, Witkowski and Zenil, “Neurodivergent influenceability in agentic AI as a contingent solution to the AI alignment problem”, PNAS Nexus, volume 5, issue 4, published 14 April 2026 (received 23 July 2025, accepted 22 December 2025). DOI 10.1093/pnasnexus/pgag076.
Read it for what it actually says. The authors argue that perfect alignment is mathematically unattainable, invoking Gödel- and Turing-style undecidability, and they propose engineering deliberate cognitive diversity among AI agents, which they call artificial agentic neurodivergence, so that no single system dominates. The word neurodivergent in this paper describes machines, by design and by metaphor. It is not a paper about neurodivergent people, and this site does not cite it as one. An earlier version of this page did, and that reading is corrected here, on the record.
Paper two · preprint
Gumbau Mezquita, “The Unverifiability of Artificial General Intelligence (AGI) Alignment, Static and Dynamic: From Trakhtenbrot’s Wall to the Safety-Generality Tension”, arXiv 2606.28639, first version 26 June 2026, substantially expanded and retitled 6 July 2026. A preprint, not yet peer-reviewed. arXiv record.
It derives a soundness, completeness and tractability trilemma for alignment verification, argues that certification persists only for systems that have stopped evolving semantically, and formalises a supervisory regress: any supervisor able to audit a general system is itself a general system. Its first version credits the Hernández-Espinosa paper as its conceptual origin, which is why the register treats the pair as one dependent chain rather than two independent arrivals.
The dated record
On 8 December 2024, sixteen months before the peer-reviewed paper and nineteen before the preprint, a self-emailed manuscript package recorded the embedded-correction thesis and the control-failure passage, with a published hash and a Google-assigned date. The direction is the same: externally imposed verification degrades as capability scales, so correction belongs inside the system.
Here is the boundary, stated the way the falsification dashboard states it: these two results are consistent with the premise, they are one dependent chain, one is peer-reviewed and one is not, and neither tests this programme. Convergence in direction is not confirmation of a mechanism, and this page does not claim otherwise. The graded rows live in the register.
What these results give this record
Accuracy first cost this page its old headline. Here is what survives at full strength, because all of it is true.
First, the premise now has formal allies. My thesis has two halves: externally imposed control degrades as capability scales, and correction therefore belongs inside the system. Both of these papers are attacks on the first half’s alternative. A peer-reviewed paper argues perfect alignment is mathematically unattainable; a formal preprint argues that verifying alignment from outside is blocked in principle, and that any supervisor capable of auditing a general system is itself a general system, so oversight never terminates. Neither author has read my work. The ground under the external route is being formally attacked by people who have never heard of me, and my statement of that premise has a hash and a date sixteen months ahead of the first of them.
Second, the positive proposals are family. When the Hernández-Espinosa authors go looking for what protects an ecosystem of agents, they reach for internal properties: engineered diversity of reasoning styles, misalignment as a counterbalancing mechanism inside the population. My answer reaches for an internal property too: correction embedded in the recursive loop itself. These are not the same mechanism, and I do not blur them. But they are the same family of answer, safety as something a system is rather than something done to it, and the family was small and unfashionable when I joined it in December 2024.
Third, the falsification angle cuts my way. My first claim has carried a standing kill-condition from the start: demonstrate a purely external oversight mechanism that remains sufficient as capability scales, and the claim dies. The Gumbau preprint argues, on formal grounds, that no such demonstration can exist. If that argument survives peer review, the escape hatch I built for my critics closes in principle, and not by my hand. I keep the kill-condition open anyway, because that is the discipline; but it is worth saying plainly that the best current formal work suggests nobody will ever walk through it.
And the dates stand in the right order. Convergence in direction, from authors who do not cite me and never saw the manuscript, is the strongest form of independent arrival short of replication. The register still grades it as context rather than confirmation, because that is what it is. But two results I did not write have moved the ground toward the place this record was already standing. That is what the register calls context. It is also what momentum looks like.
The human note
I have ADHD and autism, recognised in adulthood, and the manuscript that anchors this record was produced by exactly the kind of cross-domain, pattern-seeking cognition those labels describe. When a peer-reviewed paper argues that diversity of reasoning styles is a protective property in intelligent systems, I notice the rhyme. But a rhyme is what it is: the paper is about machines, and the human version of the argument is mine to make as biography and as a stated framework, not as someone else’s peer review. That framework is Paper C, the Polymathic Neurodivergent Profile, and it stands or falls on its own terms.
Both author groups are welcome to this record. Every date, hash and boundary above is checkable without asking me, starting at the verify page; the output the record sits inside is at what zero funding built.