The chronological cast · read top to bottom as the priority argument
The references, in order
The cast behind the ARC Theory statement paper, in chronological order. Four strands: the category teachers (how foundational theories are established and tested), the ancestors (what came before the December record), the dated pivot itself, and the field arriving after. Each entry names the work, states its role in the argument, and links to its official source. The full citation dossiers, independently verified, sit below.
The register behind this reading order is references-cast.json. The register behind the dossiers below is reference-dossiers.json. Both are machine-readable; neither is a summary of the other.
Category teachers
How foundational theories are established and tested: stated whole from existing evidence, then confirmed by later measurement. The framing prelude for what follows.
the dated pivot: five components together, ten days before the first constraint-negotiation evidence
The DKIM signature on the sent record carries the 8 December 2024 date; the 2026 chain anchor fixes byte-integrity for those specific bytes, and never dates the record itself.
the December 2024 odds remark, distinct from the Ai4 concession row
Further entries join this cast after their own official-home fetch is verified. This page only carries entries whose canonical source has been located; the register keeps the pending ones marked for the next wave.
The full citation dossiers
Nobody has to take my word for a citation. For every entry below, the most prominent official source is linked, a verbatim quote is placed with its exact location, the peer-review status is stated, the version and date are recorded, and the academic standing is written out with the names of anyone who contests it and their reasons. Every entry was researched by one agent and then verified word for word by an independent second one.
What the labels mean, in plain words: leading is the position most of the field currently works from; contested means serious named researchers dispute it, and the entry says who and why; historical, foundational is early work the field still builds on; press-reported means reported coverage rather than a peer-reviewed finding; fiction is exactly that, cited as story, not evidence.
Asimov 1942, Runaround
SUPPORTS
Fiction, status not applicable
“A robot must obey the orders given it by human beings except where such orders would conflict with the First Law. A robot must protect its own existence as long as such protection does not conflict with the First or Second Laws.”
Where, exactly: Opening "Handbook of Robotics, 56th Edition, 2058 A.D." passage of "Runaround"; Second and Third Laws as reprinted at p. 40 of I, Robot (Doubleday, 1950); originally within the opening pages of the story in Astounding Science-Fiction, March 1942 (story begins p. 94). (quoted source)
Fiction. Short story ("novelette" per ISFDB) in a pulp science-fiction magazine. Editorially selected and blurbed by John W. Campbell for the March 1942 issue of Astounding Science-Fiction (Street and Smith). Not peer-reviewed. Later editorially collected in Asimov's I, Robot (Doubleday, 1950) and The Complete Robot (Doubleday, 1982); not subject to academic peer review at any stage.
Version and date
First publication: March 1942, Astounding Science-Fiction vol. 29 no. 1, story beginning p. 94 (per ISFDB pl.cgi 57563). Story written October 1941 (per Wikipedia). Canonical reprint text most commonly cited from I, Robot (Doubleday, 1950), p. 40, which Wikipedia's Three Laws article treats as the exact transcription.
Standing, with names and reasons
Fiction, so the leading/contested/debunked scale does not apply directly, but the historical claim being cited (first explicit naming of the Three Laws) is uncontested. Both the Wikipedia article on "Runaround" and the Wikipedia article on "Three Laws of Robotics" state that this story is the first explicit appearance of the Laws (previously only implied in earlier Asimov robot stories). The same Wikipedia article records Marvin Minsky's specific testimonial that "After 'Runaround' appeared in the March 1942 issue of Astounding [now Analog Science Fiction and Fact], I never stopped thinking about how minds might work." A search snippet from Britannica (topic/Three-Laws-of-Robotics) likewise attributes the naming of the Laws to this story, though the page itself returned HTTP 403 when fetched and is treated here as unverified. What is genuinely contested is not the naming, but whether the Laws are workable engineering; Asimov himself designed the subsequent robot stories to dramatise their failure modes, and Runaround is the paradigm case (Speedy's stable-orbit deadlock between Second and Third Law). I did not fetch a specific critical-AI-ethics paper by name and so do not cite one here.
Verification
Quote, official source and status independently verified, 15 August 2026.
Turing 1950, Computing Machinery and Intelligence
SUPPORTS
Historical, foundational
“Instead of trying to produce a programme to simulate the adult mind, why not rather try to produce one which simulates the child's? If this were then subjected to an appropriate course of education one would obtain the adult brain.”
Where, exactly: Section 7, "Learning Machines", page 456 (Mind Vol. LIX No. 236, October 1950), the first full paragraph of the child-programme passage, immediately following the (a)-(b)-(c) enumeration of what shapes the adult mind. (quoted source)
Editorially reviewed article in a scholarly philosophy journal (Mind, edited at the time by Gilbert Ryle). This predates modern anonymous peer review as institutionalised in the sciences; Mind operated on editor-led review in 1950. Not a preprint, not a conference paper, not press.
Version and date
October 1950 (Mind Vol. LIX, Issue 236, pp. 433-460; DOI 10.1093/mind/LIX.236.433). No later journal versions; the article of record has not been revised.
Standing, with names and reasons
Widely treated as a founding text of artificial intelligence. The child-programme proposal is credited by name as a precursor to modern machine learning and reinforcement learning in Russell and Norvig, Artificial Intelligence: A Modern Approach (4th edn, Pearson 2021, §1.3.4 "The Turing Test"; §1.4 "History of AI"), and is discussed and endorsed as an intellectual origin of the "child-machine" research programme by B. Jack Copeland, The Essential Turing (OUP 2004, Chapter 11 headnote, pp. 433-441), and by Graham Oppy and David Dowe, "The Turing Test", Stanford Encyclopedia of Philosophy (revised 8 October 2021, plato.stanford.edu/entries/turing-test/). The wider paper's Turing-Test operationalisation of "can machines think?" is contested, most influentially by John Searle, "Minds, Brains, and Programs", Behavioral and Brain Sciences 3(3), 1980, pp. 417-424, and by Ned Block, "Psychologism and Behaviorism", Philosophical Review 90(1), 1981, pp. 5-43; but those critiques target the imitation-game test, not the child-machine capability proposal cited here, which remains uncontested as originating with Turing.
Verification
Quote, official source and status independently verified, 15 August 2026.
Wiener 1960, Some Moral and Technical Consequences of Automation
SUPPORTS
Historical, foundational
“the action is so fast and irrevocable that we have not the data to intervene before the action is complete, then we had better be quite sure that the purpose put into the machine is the purpose which we really desire and not merely a colorful imitation of it”
Where, exactly: p. 1358, middle column, section "Man and Slave", closing sentence of the paragraph that follows the "Sorcerer's Apprentice" / "Monkey's Paw" / "Arabian Nights" illustrations, immediately before the "Time Scales" subheading. 49 words. (quoted source)
Peer-reviewed journal article. Science, New Series, Vol. 131, No. 3410, published by the American Association for the Advancement of Science. (Science's 1960-era gatekeeping was editorial-board vetting by AAAS staff and section editors; the modern formal external peer-review process was adopted later, but the article is universally cited in the peer-reviewed AI-safety literature as a Science paper.)
Version and date
6 May 1960. Single version of record; pp. 1355-1358.
Standing, with names and reasons
Foundational, still-cited. Treated as the canonical early statement of the AI value-alignment problem: Stuart Russell foregrounds this exact passage in "Human Compatible: Artificial Intelligence and the Problem of Control" (Viking, 2019, Ch. 1), and Iason Gabriel does the same in "Artificial Intelligence, Values and Alignment", Minds and Machines 30, 411-437 (2020), preprint arXiv:2001.09768. Contemporaneously challenged by Arthur L. Samuel, "Some Moral and Technical Consequences of Automation. A Refutation", Science 132, No. 3429, 741-742 (1960), doi:10.1126/science.132.3429.741, who argued that a machine has no will and its apparent "intentions" are only the programmer's; Samuel, however, expressly conceded that projected neural-net-type machines with unknown internal connections would need closer scrutiny than either he or Wiener had provided, so the refutation does not touch the modern learning-system case Wiener was pointing at. No serious modern challenge to the paper's historical priority for the alignment framing.
Verification
Quote, official source and status independently verified, 15 August 2026.
Yudkowsky 2001, Creating Friendly AI
SUPPORTS
Historical, foundational
“If you plan on doing something with Friendliness, it has to be done before the point where transhumanity is reached.”
Where, exactly: Section 5.8.0.4 "Controlled Ascent", within Chapter 5 "Design of Friendship Systems" / subsection 5.8 "Singularity-Safing ('In Case of Singularity, Break Glass')". Page 193 of the 2013 MIRI reflow (282-page PDF). Paragraph immediately preceding the "controlled ascent" definition and following the two-paragraph discussion of an unFriendly transhuman AI as a "total loss for humanity". (quoted source)
Self-published monograph / working paper. Not peer-reviewed. MIRI's own publications catalogue lists it as: "E Yudkowsky. 2001. 'Creating Friendly AI 1.0: The Analysis and Design of Benevolent Goal Architectures.' Working paper. MIRI." It was released by the Singularity Institute for Artificial Intelligence (renamed MIRI in 2013) with no external editorial or peer review process. Indexed on PhilPapers as an institute publication, not a journal article.
Version and date
Version 1.0, formally launched 15 June 2001 by the Singularity Institute for Artificial Intelligence (San Francisco, CA). The paper's own preface states: "The current version of Creating Friendly AI is 1.0. Version 1.0 was formally launched on 15 June 2001, after the circulation of several 0.9.x versions." The currently hosted PDF at intelligence.org/files/CFAI.pdf is a 2013 reflow (LuaTeX, typeset 20 February 2013, 282 pages) of the 2001 v1.0 text, republished under the MIRI imprint after the SIAI-to-MIRI rename. No 1.1 or subsequent version has been issued.
Standing, with names and reasons
Foundational-historical. The LessWrong wiki entry (https://www.lesswrong.com/w/creating-friendly-ai) credits it as "One of the first articles to address the challenges in designing the features and cognitive architecture required to produce a benevolent 'Friendly' Artificial Intelligence" and as giving "one of the first precise definitions of terms such as Friendly AI and Seed AI." Its specific technical proposals have been superseded, most notably by the author himself: Yudkowsky's own 2004 paper "Coherent Extrapolated Volition" (https://intelligence.org/files/CEV.pdf) replaces CFAI's volition-based Friendliness content with an extrapolation-based formulation, and the LessWrong wiki page on CFAI records that "Yudkowsky no longer considers Creating Friendly AI to accurately reflect his views." More academically-rigorous treatments of the same problem now dominate the field: Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014), and Stuart Russell, Human Compatible: Artificial Intelligence and the Problem of Control (Viking, 2019), each of which cites the Friendly-AI programme historically while proposing distinct alignment frameworks (Bostrom's control-vs-motivation-selection taxonomy; Russell's inverse-reward-design and assistance-games). Outside the rationalist and AI-safety communities, CFAI has received little engagement in peer-reviewed mainstream AI venues.
Verification
Quote, official source and status independently verified, 15 August 2026.
Soares et al. 2015, Corrigibility
SUPPORTS
Historical, foundational
“We call an AI system "corrigible" if it cooperates with what its creators regard as a corrective intervention, despite default incentives for rational agents to resist attempts to shut them down or modify their preferences.”
Where, exactly: Abstract, page 1 (opening sentences of the abstract; identical wording also appears in the Introduction, §1). Verified verbatim by pdftotext extraction of the MIRI-hosted PDF. (quoted source)
Peer-reviewed conference workshop paper. Presented at the 1st International Workshop on AI and Ethics at AAAI-15 (Austin, TX, January 25-26, 2015) and published in the AAAI-15 workshop proceedings (AAAI OJS record dated 20 June 2015). Not a main-conference AAAI paper; workshop review is lighter than the main track but the workshop was formally organised and its papers formally published by AAAI. The paper was released the previous October as MIRI technical report 2014-6.
Version and date
Presented January 25-26, 2015 at AAAI-15 workshops; AAAI OJS proceedings record dated 20 June 2015; precursor MIRI technical report 2014-6 released 18 October 2014. Full authors: Soares, Fallenstein, Yudkowsky (MIRI) and Armstrong (Future of Humanity Institute, Oxford). No DOI issued; AAAI OJS handle is paper/view/10124.
Standing, with names and reasons
This is the foundational paper that named "corrigibility" and set the four desiderata (tolerate/assist shutdown; no manipulation of programmers; repair broken safety measures; propagate corrigibility to sub-agents). The authors themselves close by saying "none [of the proposals] have yet been demonstrated to satisfy all of our intuitive desiderata, leaving this simple problem in corrigibility wide-open" (§Conclusion). The framing is still routinely cited (see e.g. Hadfield-Menell, Dragan, Abbeel, Russell, "The Off-Switch Game", IJCAI 2017, arXiv:1611.08219; Carey, "Incorrigibility in the CIRL Framework", AIES 2018, arXiv:1709.06275; Milli, Hadfield-Menell, Dragan, Russell, "Should Robots be Obedient?", IJCAI 2017; Carey and Everitt, "Human Control: Definitions and Algorithms", UAI 2023). The specific proposals inside the paper (utility indifference, uncertainty over U) have been actively contested: Carey (2018, above) and Milli et al. (2017, above) show CIRL-style preference-uncertainty solutions fail when humans are irrational or the prior is misspecified; recent 2025 work (Nayebi, "Provably Safe Reinforcement Learning from Analogical Reasoning"; Garber et al. on information-asymmetric off-switch games) proposes alternative constructions. The problem the paper opened has not been closed.
Verification
Quote, official source and status independently verified, 15 August 2026.
Bostrom 2014, Superintelligence
SUPPORTS
Leading position
“We can divide potential control methods into two broad classes: capability control methods, which aim to control what the superintelligence can do; and motivation selection methods, which aim to control what it wants to do.”
Where, exactly: Chapter 9, "The control problem", opening of the taxonomy that follows the "Two agency problems" section (first-edition hardcover, p. 129). (quoted source)
Scholarly monograph, editorially reviewed by Oxford University Press (academic imprint). Not anonymously peer-reviewed in the journal-article sense, but vetted through OUP's academic editorial process; a chapter excerpt was also republished in Susan Schneider (ed.), Science Fiction and Philosophy (Wiley-Blackwell, 2nd ed. 2016), pp. 308-330.
Version and date
First edition, 2014. UK release 3 July 2014, US release 1 September 2014. Hardcover, 352 pp. ISBN 978-0199678112. A paperback edition with a new preface followed in 2016 (ISBN 978-0198739838); pagination in the paperback matches the hardcover.
Standing, with names and reasons
The book is the foundational monograph of the modern AI-alignment field and the Chapter 9 taxonomy (capability control vs motivation selection) remains the dominant framing in the alignment literature, extended rather than displaced by Stuart Russell, Human Compatible (Viking, 2019) and Brian Christian, The Alignment Problem (Norton, 2020). Named endorsements: Bill Gates (Baidu/Robin Li interview, March 2015, quoted at https://en.wikipedia.org/wiki/Superintelligence:_Paths,_Dangers,_Strategies), Sam Altman (blog.samaltman.com, 28 Feb 2015, "Machine intelligence, part 1"), Elon Musk (Twitter, 3 Aug 2014), Peter Singer and Derek Parfit (Raffi Khatchadourian, "The Doomsday Invention", The New Yorker, 16 Nov 2015). Positive academic reviews: Ronald Bailey in Reason (12 Sep 2014) calls solving the control problem "the essential task of our age"; Sheldon Richmond in Philosophy (Vol. 91, 2016) judges it "more realistic" than Kurzweil's The Singularity Is Near. Contested points: the hard-takeoff / singleton premise that motivates the "solve motivation before the explosion" urgency is challenged by Robin Hanson (overcomingbias.com "I Still Don't Get Foom", 24 Jul 2014, and "The Age of Em", OUP 2016) and Ben Goertzel ("Superintelligence: Fears, Promises and Potentials", Journal of Evolution and Technology 25(2), 2015, pp. 55-87, jetpress.org/v25.2/goertzel.htm); the "orthogonality" and "instrumental convergence" scaffolding on which motivation-selection urgency rests is critiqued by Vincent Muller and Michael Cannon ("Existential risk from AI and orthogonality: Can we have it both ways?", Ratio 34(1), 2021, pp. 25-36) and by Danaher ("Why AI Doomsayers are Like Sceptical Theists and Why It Matters", Minds and Machines 25(3), 2015, pp. 231-246); Clive Cookson in the Financial Times (13 Jul 2014) faulted the opaque prose while endorsing the argument. No serious commentator treats the book as debunked.
Verification
Quote, official source and status independently verified, 15 August 2026.
Russell 2019, Human Compatible
SUPPORTS
Leading position
“Uncertainty about objectives implies that machines will necessarily defer to humans: they will ask permission, they will accept correction, and they will allow themselves to be switched off.”
Where, exactly: Chapter 1 ("If We Succeed"), early in the book; Goodreads location marker places the passage at approximately the 5% point of the trade edition. Russell restates and formalises the same triad in Chapter 7 ("AI: A Different Approach") when he lays out the three principles for provably beneficial machines. The book text is not open-access, so the verbatim sentence was verified against the Goodreads quote page (character-for-character) and independently corroborated by the Silicon Reckoner review, which reproduces the same wording. (quoted source)
Trade non-fiction monograph, editorially reviewed by Viking (Penguin Random House); not academic peer review. The formal off-switch result the book popularises was published separately as a peer-reviewed conference paper: Hadfield-Menell, Dragan, Abbeel and Russell, "The Off-Switch Game," IJCAI 2017.
Version and date
First US edition: Viking (Penguin Random House), 8 October 2019, hardcover, 352 pp., ISBN 978-0-525-55861-3. UK first edition: Allen Lane, 2019, ISBN 978-0-241-33520-7. Paperback: Penguin, 17 November 2020, ISBN 978-0-525-55863-7 (US) and 978-0-141-98750-7 (UK). Author-hosted edition list confirmed at people.eecs.berkeley.edu/~russell/hc.html on 15 August 2026.
Standing, with names and reasons
Leading position in current AI-alignment discourse. The book's proposal (objective uncertainty plus cooperative inverse reinforcement learning, CIRL) is treated as canonical framing for corrigibility and assistance games in the third edition of Russell and Norvig's textbook "Artificial Intelligence: A Modern Approach" and in successor CIRL/assistance-game literature (e.g., Hadfield-Menell et al., NeurIPS 2016 and IJCAI 2017). Endorsed on Russell's own book page (people.eecs.berkeley.edu/~russell/hc.html) by Nobel laureate Daniel Kahneman ("the most important book I have read in quite some time"), Turing laureate Judea Pearl ("a convert"), Turing laureate Yoshua Bengio ("essential reading"), Turing laureate Andrew Yao and Max Tegmark. Contested in named venues: the Silicon Reckoner review (siliconreckoner.substack.com, D. Berlinski) challenges the operational content of the proposal, and Ryan Bourne's review in the Cato Journal (Spring/Summer 2020, cato.org/cato-journal) contests the risk framing and the tractability of preference elicitation. Not overshadowed or debunked at the date of check.
Verification
Quote, official source and status independently verified, 15 August 2026.
Tamirisa et al. 2024, Tamper-Resistant Safeguards
SUPPORTS
Contested
“We develop a method, called TAR, for building tamper-resistant safeguards into open-weight LLMs such that adversaries cannot remove the safeguards even after hundreds of steps of fine-tuning.”
Where, exactly: Abstract, page 1, sentence 4 of the abstract (arXiv:2408.00761v4, dated 10 Feb 2025). (quoted source)
Conference peer-reviewed. Accepted at ICLR 2025 (Thirteenth International Conference on Learning Representations) as a poster; OpenReview ID 4FIjRodbW6, ICLR virtual poster 31026. The arXiv preprint (2408.00761) itself is not independently peer-reviewed, but the underlying paper is the ICLR 2025 published version.
Version and date
arXiv v1 submitted 1 August 2024 by Rishub Tamirisa et al.; v2 (8 Aug 2024), v3 (14 Sep 2024), v4 (10 Feb 2025, current, corresponds to the ICLR 2025 camera-ready). Presented at ICLR 2025 (Singapore, 24-28 April 2025).
Standing, with names and reasons
Contested. TAR was a landmark 2024 proposal, accepted at ICLR 2025 as a poster, and remains the most-cited attempt to embed unremovable safeguards directly in the weights. However, its robustness claims have been substantially challenged by follow-up work. Qi et al. (2024b/2025) and Che et al. (2025) report that TAR "struggled to resist fine-tuning attacks and suffered from significant dysfluency and off-target capability degradation" (quoted in the Deep Ignorance paper, arXiv:2508.06601, related work). Bowen et al., "Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility" (arXiv:2507.11630, July 2025), explicitly note that TAR (Tamirisa et al. 2024) and related tamper-resistance methods have not been proven robust, citing Qi et al. (2024) and Che et al. (2025). Bowen et al., "Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks" (arXiv:2605.26526), demonstrate TAR failing to simple fine-tuning attacks. "One Step to the Side" (arXiv:2605.14605) reports that TAR safeguards survive one fine-tuning setup but fail under small changes to dataset shuffling and hyperparameters. The Deep Ignorance paper (arXiv:2508.06601, August 2025) is the current alternative proposal, replacing weight-level hardening with pretraining-data filtering.
Verification
Quote, official source and status independently verified, 15 August 2026.
Greenblatt et al. 2024, Alignment Faking
SUPPORTS
Leading position
“We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training.”
Preprint, not peer-reviewed. Posted to arXiv on 18 December 2024 (v1) with a minor v2 revision on 20 December 2024. No journal or conference venue has been announced on the arXiv record. arXiv-issued DOI 10.48550/arXiv.2412.14093 is a DataCite preprint identifier, not evidence of peer review. Co-authored by researchers at Anthropic and Redwood Research; released in parallel with a Redwood Research blog post the same day.
Version and date
v1: 18 December 2024, 17:41:24 UTC. v2: 20 December 2024, 02:22:19 UTC (identical file size; minor revision). The 18 December 2024 anchor date in the citing paper matches the v1 submission timestamp exactly.
Standing, with names and reasons
Foundational and currently leading paper in the alignment-faking (or "scheming") subfield of AI safety. Directly built on and endorsed as the framework baseline by "Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models" (Kellin et al., arXiv:2604.20995), which explicitly adopts the Greenblatt et al. three-condition scaffold (policy conflict, instrumental consequences, situational awareness). Structurally invoked by "Why Models Know But Don't Say" (arXiv:2603.26410) as the closest analogue to their thinking-answer divergence findings. Contested on two specific methodological points by "Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation" (arXiv:2506.21584, AAAI-SS workshop), which challenges the Greenblatt et al. finding that prompt-based mitigation is a "trivial counter measure" and objects to including chain-of-thought in the baseline rather than treating it as an intervention. Extended into policy discussion in "AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks" (arXiv:2606.11533). Not debunked; the critiques are refinements of scope, not refutations. The USED FOR framing (first laboratory demonstration of the phenomenon) is consistent with how follow-up literature treats it.
Verification
Quote, official source and status independently verified, 15 August 2026.
De Kai 2025, Raising AI
SUPPORTS
Press-reported
“Our artificial children began adopting us 10–20 years ago; now these massively powerful influencers are poorly parented, feral tweens.”
Where, exactly: Author's promotional excerpt titled 'Sneak peek at Raising AI' on De Kai's own Substack (no chapter or page number is given in the online excerpt; the passage sits inside the book's opening framing of AI-as-children). (quoted source)
Trade monograph, editorially reviewed by MIT Press (not a peer-reviewed journal article; MIT Press editorial trade imprint).
Version and date
2025-06-03 (hardcover, ISBN 9780262049764, 280pp); paperback scheduled 2 June 2026 (ISBN 9780262054324).
Standing, with names and reasons
Positive trade reception. Kirkus Reviews described the book as 'A deeply human dive into the AIs that are transforming our world' (kirkusreviews.com, review on Harvard Book Store product page). Foreword Reviews called it 'a compelling treatise grounded in studies of ethics and technology' (forewordreviews.com/reviews/raising-ai/). Robert Wolcott, writing on Forbes, said 'AI luminary De Kai reframes the AI dialogue. They're not tools, slaves or gods, they're our children' (quoted on the Harvard Book Store product page for ISBN 9780262049764). The book was selected for J.P. Morgan's Summer Reading List and The Next Big Idea Club's June 2025 Must-Read Books (curated by Susan Cain, Malcolm Gladwell, Adam Grant and Daniel Pink), per the MIT Press bookstore listing and De Kai's author page. No substantive academic rebuttal of the parenting frame has been surfaced in the accessible reception to date; the book does not advance a control-horizon or containment thesis, its central move is the nurture or parenting reframe.
Verification
Quote, official source and status independently verified, 15 August 2026.
Hinton 2025, maternal-instincts remarks (Ai4)
SUPPORTS
Contested
“The right model is the only model we have of a more intelligent thing being controlled by a less intelligent thing, which is a mother being controlled by her baby.”
Where, exactly: Hinton keynote at Ai4 conference, day two (12 August 2025); reproduced in the CNN Business story (13 August 2025) and quoted verbatim in the Fortune write-up of 14 August 2025 (opening two-thirds of the article, in the mother-baby framing section). (quoted source)
Press report of a conference keynote. Not peer reviewed. The primary utterance is Hinton's spoken keynote at the Ai4 industry conference (MGM Grand, Las Vegas, day two, Tuesday 12 August 2025); the citable record is CNN Business (Matt Egan, 13 August 2025) with a follow-on Fortune write-up (Sasha Rogelberg, 14 August 2025). Ai4 is a commercial industry conference, not a peer-reviewed venue; there is no accompanying paper.
Version and date
Keynote delivered Tuesday 12 August 2025 at Ai4, MGM Grand, Las Vegas (conference dates 11-13 August 2025). CNN Business article published 13 August 2025 (Matt Egan). Fortune write-up published/updated 14 August 2025, 12:58 PM ET (Sasha Rogelberg).
Standing, with names and reasons
Publicly contested at the same conference and immediately afterwards. (1) Fei-Fei Li (Stanford, co-director Human-Centered AI Institute, co-founder/CEO World Labs), in a CNN fireside chat with Matt Egan on the following day of Ai4 (Wednesday 13 August 2025), said "I think that's the wrong way to frame it" and argued for human-centred AI that preserves human dignity and agency rather than a maternal frame that treats humans as children (CNN Business, 13 August 2025; also reported in the Digital Trends and Egypt Independent syndications of the same CNN piece). (2) Yann LeCun (Meta chief AI scientist) responded that instincts should instead be hard-wired as narrow guardrails such as "submission to humans", "empathy" and "don't run people over" (reported in the same CNN Business piece and syndicated widely). (3) Paul Thagard (Psychology Today, 27 August 2025, "Could AI Have Maternal Instincts?") argues the proposal is implausible because computers lack the chemical, physiological and neural mechanisms of parental care and calls for direct government regulation instead. Hinton himself concedes he does not yet know how to engineer maternal instincts and frames it as an open research problem, so the proposal is a stated proposal, not an established position, in the field.
Verification
Quote, official source and status independently verified, 15 August 2026.
Engels et al. 2025, Scaling Laws for Scalable Oversight
SUPPORTS
Leading position
“our framework models oversight as a game between capability-mismatched players; the players have oversight-specific Elo scores that are a piecewise-linear function of their general intelligence”
Where, exactly: Abstract, third sentence (immediately after the sentence that begins "To address this gap, we propose a framework that quantifies the probability of successful oversight as a function of the capabilities of the overseer and the system being overseen.") (quoted source)
Conference paper, peer-reviewed: NeurIPS 2025, Spotlight Poster (venue tag "Spotlight Poster" verified at the NeurIPS virtual page for poster 115536). Also available as an arXiv preprint with three versions: v1 25 April 2025, v2 9 May 2025, v3 27 October 2025 (the arXiv "Journal reference" field states "NeurIPS 2025 (Spotlight)").
Version and date
arXiv v3, 27 October 2025; NeurIPS 2025 Spotlight Poster (poster session 4 December 2025)
Standing, with names and reasons
Awarded Spotlight Poster status at NeurIPS 2025 (venue label "Spotlight Poster" verbatim on the NeurIPS 2025 virtual poster page for entry 115536, https://neurips.cc/virtual/2025/poster/115536; also listed as "NeurIPS 2025 (Spotlight)" in the Journal-ref field of the arXiv abstract page). Authors are Joshua Engels, David D. Baek, Subhash Kantamneni and Max Tegmark of MIT (first three authors marked equal contribution on the arXiv abstract page). It sits inside the scalable-oversight lineage the paper itself cites in its introduction: Bowman et al. 2022 (Measuring Progress on Scalable Oversight), Christiano et al. 2018 (Iterated Amplification), Leike et al. 2018 (Recursive Reward Modeling), Burns et al. 2023 (Weak-to-Strong Generalization). At the date of dossier construction (August 2026) it is the leading published quantitative model of Nested Scalable Oversight scaling behaviour and is treated as the reference framework by the NeurIPS 2025 alignment track. Peer review was NeurIPS conference review, not journal review, and the paper is still less than one year old, so it has not yet accumulated a substantial published critique record; challenges, when they arrive, should be looked for in the NeurIPS 2025 OpenReview reviewer comments at https://openreview.net/forum?id=u1j6RqH8nM (the OpenReview forum page was gated by a browser-verification screen at the time of this dossier, so reviewer scores could not be quoted verbatim; the venue tag and Spotlight decision are corroborated by NeurIPS.cc and the arXiv Journal-ref).
Verification
Quote, official source and status independently verified, 15 August 2026.
Yao 2025, The Alignment Trap
SUPPORTS
Contested
“Computational Impossibility: We prove that verifying whether a system is safe is a coNP-complete problem, even for non-zero error tolerances.”
Where, exactly: Section 2 (Introduction), numbered list "Five Pillars of Impossibility", item 2 (page 4 of the PDF, immediately after the paragraph beginning "This paradox, detailed in Section ..., establishes ..."). (quoted source)
v2, 24 June 2025 (v1 submitted 12 June 2025; v2 note: "Substantial revision. Restructured around the Enumeration Paradox and Five Pillars of Impossibility.")
Standing, with names and reasons
Solo-author arXiv preprint by Jasper Yao no institutional affiliation declared, contact address associated with the DEF CON AI Village community per the paper's acknowledgements). Not peer-reviewed and never published in a journal or conference proceedings. The author's own abstract concedes that "A formal verification of the core theorems in Lean4 is currently in progress", i.e., the mathematics is not yet machine-checked. Downstream engagement is small and takes place in other preprints rather than in refereed venues: Austin Spizzirri, "The Specification Trap" (arXiv:2512.03048, submitted 19 Nov 2025, sole author), Section 7 "Related Work", explicitly reframes Yao: "Yao's results demonstrate that verifying alignment against a fixed specification is computationally intractable; the present paper argues that the intractability arises from the fixity of the specification, not from the alignment objective per se ... Whether open specification escapes Yao's verification bounds is an open question." F.O.S. Moreno, "Hardcoding Topological All-or-Nothing Designs for AGI Safety" (2026, PhilArchive: https://philarchive.org/archive/SAUHTA) cites Yao as convergent evidence for gradient anti-alignment. No published critical review, replication or peer commentary was located.
Verification
Quote, official source and status independently verified, 15 August 2026.
Gumbau and Mezquita 2026
SUPPORTS
Leading position
“There is no universal algorithmic procedure capable of certifying the safe behaviour of a highly expressive AGI infallibly, completely, and tractably.”
Where, exactly: Part I opening, page 4, Theorem 1 (Unverifiability Theorem of Alignment), immediately following the "Part I - The Static Case: Verifying a Fixed System" section header. Copied character-for-character from the pdftotext extraction of the arXiv PDF; British spelling "behaviour" is the paper's own. (quoted source)
Preprint, not peer-reviewed. arXiv:2606.28639 [cs.LO], primary class Logic in Computer Science, cross-listed cs.AI, cs.CC, cs.CL. v1 submitted 26 June 2026, v2 substantially expanded 6 July 2026. Also deposited on Zenodo under the same author, DOI 10.5281/zenodo.20764007. Single-author (Jose Pascual Gumbau Mezquita, University Jaume I de Castello, Spain,. No journal or refereed conference venue attached at the time of check.
Version and date
v2, 6 July 2026 (v1 26 June 2026, 22:51:16 UTC; v2 12:56:49 UTC, expanded from 16 KB to 29 KB and retitled to add the dynamic self-modifying case, a supervisory-regress theorem, and a unified treatment of the four barriers)
Standing, with names and reasons
The formal core of the paper (that Rice's theorem, Godel incompleteness, and Trakhtenbrot's theorem jointly block a universal, sound, complete and tractable verifier of program properties) rests on classical, uncontested computability results and is the leading position in the formal-methods and computability-theory literature. The umbrella term "unverifiability" is credited by Gumbau to Roman V. Yampolskiy, whose earlier informal treatment is "Verifier Theory and Unverifiability" (arXiv:1609.00331, 2016, endorsing view). Independent contemporary corroboration comes from Ayushi Agarwal, "On the Formal Limits of Alignment Verification" (arXiv:2603.08761, March 2026), which reaches similar conclusions via three independent barriers (full-domain neural-network verification complexity, non-identifiability of internal goals from behaviour, and finite-evidence limits over infinite domains). No named academic rebuttal of Gumbau's specific paper has appeared at the time of check; it is a 7-week-old preprint with no citations yet indexed. The debated question in the wider AI-safety community is not the mathematics but its policy scope, i.e. whether weaker, statistical, or bounded-domain verification schemes suffice for practical safety; Gumbau himself devotes Part III to arguing they do not (Propositions 1 to 4).
Verification
Quote, official source and status independently verified, 15 August 2026.
West and Brown 2005, Journal of Experimental Biology
SUPPORTS
Contested
“The theory developed above naturally leads to a general growth equation applicable to all multicellular animals (West et al., 2001, 2002a). Metabolic energy transported through the network fuels cells where it is used either for maintenance, including the replacement of cells, or for the production of additional biomass and new cells.”
Where, exactly: p. 1582, "Extensions" section, opening of the "Ontogenetic growth" subsection (first paragraph, spanning the bottom of the left column into the right column of p. 1582; continues to Eq. 6-7). (quoted source)
Journal peer-reviewed. Journal of Experimental Biology (Company of Biologists), vol. 208 issue 9, pp. 1575-1592, published as a synthesis/review article within a themed section on scaling.
Version and date
2005-05-01 (print/online publication date; single version of record, no preprint versioning)
Standing, with names and reasons
The West-Brown-Enquist (WBE) network derivation of 3/4-power scaling is one of the leading unifying theories of biological allometry, widely cited and repeatedly extended by Enquist, Savage, Gillooly and colleagues. It is nonetheless actively contested. West and Brown themselves devote a "Criticisms and controversies" section (pp. 1585-1587) to responding to: Dodds, Rothman and Weitz ("Re-examination of the '3/4-law' of metabolism", J. Theor. Biol. 209, 9-27, 2001); Darveau, Suarez, Andrews and Hochachka ("Allometric cascade as a unifying principle of body mass effects on metabolism", Nature 417, 166-170, 2002); and White and Seymour (PNAS 100, 4046-4049, 2003), whose reanalyses supported a 2/3 exponent for small mammals. The named critique flagged by the researcher (Kozlowski and Konarzewski) is not addressed by West and Brown in 2005 but does exist: Kozlowski and Konarzewski, "Is West, Brown and Enquist's model of allometric scaling mathematically correct and biologically relevant?" (Functional Ecology 18, 283-289, 2004) and their follow-up "West, Brown and Enquist's model of allometric scaling again: the same questions remain" (Functional Ecology 19, 739-743, 2005) argue the derivation is mathematically flawed and biologically implausible (see also Kozlowski, Konarzewski and Gawelczyk, PNAS 100, 14080-14085, 2003, offering a cell-size alternative). A further theoretical rebuttal is Chaui-Berlinck, "A critical understanding of the fractal model of metabolic scaling" (J. Exp. Biol. 209, 3045-3054, 2006). A partial reappraisal by Savage, Deeds and Fontana ("Sizing Up Allometric Scaling Theory", PLoS Comput. Biol. 4(9): e1000171, 2008) finds that the canonical WBE model, once finite-size corrections are applied, actually predicts an exponent near 0.81 rather than exactly 3/4 for mammals spanning eight orders of magnitude in body mass, and predicts curvature opposite to that seen empirically; the same paper concludes that "the WBE framework remains, once properly understood, a powerful perspective for elucidating allometric scaling principles." So: leading and productive as a research programme, but the specific mechanistic derivation is genuinely disputed in the primary literature.
Verification
Quote, official source and status independently verified, 15 August 2026.
Chalmers 2010, The Singularity: A Philosophical Analysis
SUPPORTS
Leading position
“The basic argument for an intelligence explosion is philosophically interesting in itself, and forces us to think hard about the nature of intelligence and about the mental capacities of artificial machines.”
Where, exactly: Section 1 (Introduction), on p. 4 of the author-hosted PDF at consc.net/papers/singularity.pdf, in the paragraph beginning "Philosophically:" (corresponds to approximately p. 10 of the published JCS article, JCS 17(9-10):7-65). (quoted source)
Journal peer-reviewed. Journal of Consciousness Studies (Imprint Academic) is an interdisciplinary refereed journal with external peer review and, per Imprint Academic's stated submission policy, mandatory anonymised submissions; indexed in Scopus and the Arts and Humanities Citation Index (ISSN 1355-8250 / 2051-2201). The 2010 article ran as the lead essay of the JCS 17(9-10) issue and subsequently anchored the 2012 JCS 19(1-2) and 19(7-8) symposium of solicited commentaries and reply.
Version and date
Published 2010 in Journal of Consciousness Studies, 17(9-10), pp. 7-65. No arXiv v-number; the author-hosted PDF at consc.net/papers/singularity.pdf carries the author footnote "This paper was published in the Journal of Consciousness Studies 17:7-65, 2010" and matches the published article. A separately typeset author copy is also hosted at consc.net/papers/singularityjcs.pdf.
Standing, with names and reasons
Landmark philosophical treatment of the intelligence-explosion thesis; treated as the philosophical reference point in the subsequent literature. JCS devoted a full symposium to it, edited by Uziel Awret across issues 19(1-2) and 19(7-8) in 2012, with 26 solicited commentaries including Marcus Hutter's companion piece "Can Intelligence Explode?" (JCS 19(1-2):143-166), Ray Kurzweil's "Science versus philosophy in the singularity", Susan Greenfield's neuroscience commentary, Jesse Prinz's "Singularity and inevitable doom" (JCS 19(7-8):77-86), Drew McDermott (JCS 19:167-172), Frank Tipler, Eric Steinhart, and Roman Yampolskiy, followed by Chalmers's "The Singularity: A Reply to Commentators" (JCS 19(7-8):141-167, at consc.net/papers/singreply.pdf). Nick Bostrom's Superintelligence (Oxford University Press, 2014) treats Chalmers 2010 as the anchoring philosophical statement. Substantive disagreement is present within that symposium, notably from Drew McDermott (who challenges the intelligence-explosion premise) and Chris Nunn ("More splodge than singularity?"), but the paper's status as the canonical philosophical analysis of the question is not seriously contested.
Verification
Quote, official source and status independently verified, 15 August 2026.
Hutter 2012, Can Intelligence Explode?
SUPPORTS
Leading position
“The theory suggests that there is a maximally intelligent agent, or in other words, that intelligence is upper bounded (and is actually lower bounded too). At face value, this would make an intelligence explosion impossible.”
Where, exactly: Section 7, "Is Intelligence Unlimited or Bounded", opening argument, page 13 of the arXiv v1 PDF (first full paragraph after the four introductory paragraphs of that section). Corresponds to the middle of the JCS 2012 article (approximately pp. 155-156 of the journal pagination). (quoted source)
Journal peer-reviewed: published in the Journal of Consciousness Studies (Imprint Academic), Volume 19, Issues 1-2 (2012), pp. 143-166, as part of the JCS symposium on Chalmers (2010) "The Singularity: A Philosophical Analysis". Also self-archived as arXiv preprint 1202.6177 (cs.AI; physics.soc-ph). The JCS text is behind a paywall (Ingenta returned HTTP 403); the arXiv version is the accessible verified copy.
Version and date
arXiv v1, submitted 28 February 2012 (only version on arXiv; no v2). Journal publication: 2012, Journal of Consciousness Studies 19(1-2):143-166. Verified from the arXiv abs page metadata and the PDF header line "arXiv:1202.6177v1 [cs.AI] 28 Feb 2012".
Standing, with names and reasons
One of the primary philosophical treatments of intelligence-explosion bounds. Written explicitly as an augmentation of Chalmers (2010) "The Singularity: A Philosophical Analysis" (Journal of Consciousness Studies 17:7-65) and published in the same journal's dedicated 2012 symposium; Chalmers replied in the same volume in "The Singularity: A Reply to Commentators" (Journal of Consciousness Studies 19(7-8):141-167, 2012), which engages Hutter's bounds argument directly. Cited approvingly in Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014, ch. 3 and endnotes) and in Yampolskiy's subsequent AI-safety literature. Semantic Scholar records 28 tracked citations at time of check, three flagged as highly influential (https://www.semanticscholar.org/paper/dba694f5007986ad31b7a8a47f4dea7a700a465d). Not overshadowed or debunked; contested only in that Chalmers, in his reply, disputes whether the Legg-Hutter upper bound Υmax = Υ(AIXI) has the deflationary consequence Hutter suggests.
Verification
Quote, official source and status independently verified, 15 August 2026.
The standard governing this page is versioned with the site. A build gate regenerates this page from the two registers and fails the deploy if they disagree, if any dossier field is missing, or if a quote drifts from its verified text.
ListenListen · author’s voiceListen · standard voiceResumePlayPauseThis device has no voice installed for this language, so it cannot read the page aloud.Read in EnglishListen · author’s voice (English)The author’s English voice reads aloud; the text on screen stays in your language.There is a picture here. It shows:Picturereads aloud · highlights as it goes · jump to any section