{
 "_meta": {
  "standard": "Every external reference on a published surface carries: most prominent official source URL, verbatim quote with locator, peer-review status, version and date, and standing (leading / contested / historical-foundational / press-reported / not-applicable-fiction) with named challengers and reasons. Each entry independently verbatim-verified by a second agent wave.",
  "provenance": "Deep-research fleet wf_07933af7 (17 research agents) + independent verification wave (17 verify agents), completed 15 August 2026, 4.9M tokens. Ingested in programme editorial review the same day.",
  "verification_date": "2026-08-15",
  "canon": "docs/SITE-REFERENCE-RIGOUR-STANDARD.md",
  "enforced_surfaces": [
   "research/references/index.html"
  ],
  "published_projection": "Working notes and contact addresses are stripped from this published projection; the full fleet artefact is preserved privately."
 },
 "references": [
  {
   "refkey": "asimov1942runaround",
   "official_url": "https://isfdb.org/cgi-bin/pl.cgi?57563=",
   "secondary_urls": [
    "https://en.wikipedia.org/wiki/Runaround_(story)",
    "https://en.wikipedia.org/wiki/Three_Laws_of_Robotics"
   ],
   "peer_review_status": "Fiction. Short story (\"novelette\" per ISFDB) in a pulp science-fiction magazine. Editorially selected and blurbed by John W. Campbell for the March 1942 issue of Astounding Science-Fiction (Street and Smith). Not peer-reviewed. Later editorially collected in Asimov's I, Robot (Doubleday, 1950) and The Complete Robot (Doubleday, 1982); not subject to academic peer review at any stage.",
   "version_date": "First publication: March 1942, Astounding Science-Fiction vol. 29 no. 1, story beginning p. 94 (per ISFDB pl.cgi 57563). Story written October 1941 (per Wikipedia). Canonical reprint text most commonly cited from I, Robot (Doubleday, 1950), p. 40, which Wikipedia's Three Laws article treats as the exact transcription.",
   "quote": "A robot must obey the orders given it by human beings except where such orders would conflict with the First Law. A robot must protect its own existence as long as such protection does not conflict with the First or Second Laws.",
   "quote_location": "Opening \"Handbook of Robotics, 56th Edition, 2058 A.D.\" passage of \"Runaround\"; Second and Third Laws as reprinted at p. 40 of I, Robot (Doubleday, 1950); originally within the opening pages of the story in Astounding Science-Fiction, March 1942 (story begins p. 94).",
   "quote_source_url": "https://en.wikipedia.org/wiki/Runaround_(story)",
   "standing_class": "not-applicable-fiction",
   "standing_detail": "Fiction, so the leading/contested/debunked scale does not apply directly, but the historical claim being cited (first explicit naming of the Three Laws) is uncontested. Both the Wikipedia article on \"Runaround\" and the Wikipedia article on \"Three Laws of Robotics\" state that this story is the first explicit appearance of the Laws (previously only implied in earlier Asimov robot stories). The same Wikipedia article records Marvin Minsky's specific testimonial that \"After 'Runaround' appeared in the March 1942 issue of Astounding [now Analog Science Fiction and Fact], I never stopped thinking about how minds might work.\" A search snippet from Britannica (topic/Three-Laws-of-Robotics) likewise attributes the naming of the Laws to this story, though the page itself returned HTTP 403 when fetched and is treated here as unverified. What is genuinely contested is not the naming, but whether the Laws are workable engineering; Asimov himself designed the subsequent robot stories to dramatise their failure modes, and Runaround is the paradigm case (Speedy's stable-orbit deadlock between Second and Third Law). I did not fetch a specific critical-AI-ethics paper by name and so do not cite one here.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Quote matches quote_source_url (Runaround Wikipedia) character-for-character, including plural \"First or Second Laws\". Three things to flag, none repaired: (1) The parallel Wikipedia article \"Three Laws of Robotics\" ends the Third Law with the singular \"First or Second Law\", so readers cross-checking against that article will see a one-letter variant; the dossier's quote follows its named source. (2) The dossier says the Runaround Wikipedia footnote cites \"Asimov, Isaac (1950), I, Robot, Doubleday, p. 40\"; today that footnote actually reads \"Asimov, Isaac (1988). I, Robot. The Isaac Asimov Collection. New York: Doubleday. p. 40. ISBN 0-385-42304-7\", i.e. the 1988 Isaac Asimov Collection reissue, same page 40. The 1950 date does appear on the Three Laws of Robotics Wikipedia article's citation. (3) The ISFDB official_url https://isfdb.org/cgi-bin/pl.cgi?57563= returned HTTP 403 to WebFetch on two attempts (with and without www.); the record content was verified indirectly via a web-search snippet that reproduced the ISFDB entry describing Astounding Science-Fiction, March 1942, Vol 29 No 1, Street and Smith, with Runaround by Asimov starting on p. 94. Peer-review status (fiction, not peer-reviewed) and version-date (first publication March 1942, written October 1941) are confirmed by both Wikipedia sources."
   }
  },
  {
   "refkey": "Turing1950",
   "official_url": "https://academic.oup.com/mind/article/LIX/236/433/986238",
   "secondary_urls": [
    "https://doi.org/10.1093/mind/LIX.236.433",
    "http://www.eliassi.org/turing-mind-1950.pdf",
    "https://courses.cs.umbc.edu/471/papers/turing.pdf",
    "https://www.cs.ox.ac.uk/activities/ieg/e-library/sources/t_article.pdf",
    "https://en.wikipedia.org/wiki/Computing_Machinery_and_Intelligence"
   ],
   "peer_review_status": "Editorially reviewed article in a scholarly philosophy journal (Mind, edited at the time by Gilbert Ryle). This predates modern anonymous peer review as institutionalised in the sciences; Mind operated on editor-led review in 1950. Not a preprint, not a conference paper, not press.",
   "version_date": "October 1950 (Mind Vol. LIX, Issue 236, pp. 433-460; DOI 10.1093/mind/LIX.236.433). No later journal versions; the article of record has not been revised.",
   "quote": "Instead of trying to produce a programme to simulate the adult mind, why not rather try to produce one which simulates the child's? If this were then subjected to an appropriate course of education one would obtain the adult brain.",
   "quote_location": "Section 7, \"Learning Machines\", page 456 (Mind Vol. LIX No. 236, October 1950), the first full paragraph of the child-programme passage, immediately following the (a)-(b)-(c) enumeration of what shapes the adult mind.",
   "quote_source_url": "https://academic.oup.com/mind/article/LIX/236/433/986238",
   "standing_class": "historical-foundational",
   "standing_detail": "Widely treated as a founding text of artificial intelligence. The child-programme proposal is credited by name as a precursor to modern machine learning and reinforcement learning in Russell and Norvig, Artificial Intelligence: A Modern Approach (4th edn, Pearson 2021, §1.3.4 \"The Turing Test\"; §1.4 \"History of AI\"), and is discussed and endorsed as an intellectual origin of the \"child-machine\" research programme by B. Jack Copeland, The Essential Turing (OUP 2004, Chapter 11 headnote, pp. 433-441), and by Graham Oppy and David Dowe, \"The Turing Test\", Stanford Encyclopedia of Philosophy (revised 8 October 2021, plato.stanford.edu/entries/turing-test/). The wider paper's Turing-Test operationalisation of \"can machines think?\" is contested, most influentially by John Searle, \"Minds, Brains, and Programs\", Behavioral and Brain Sciences 3(3), 1980, pp. 417-424, and by Ned Block, \"Psychologism and Behaviorism\", Philosophical Review 90(1), 1981, pp. 5-43; but those critiques target the imitation-game test, not the child-machine capability proposal cited here, which remains uncontested as originating with Turing.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Verification method: fetched eliassi.org PDF (which carries the Oxford Academic download stamp \"Downloaded from https://academic.oup.com/mind/article/LIX/236/433/986238 by guest on 24 October 2021\") and the UMBC mirror, extracted text with pdftotext, and compared word-for-word. The 40-word quoted sentence appears verbatim in both mirrors; Section 7 title (\"7. Learning Machines.\") is present, and the page-456 marker sits immediately before the passage in the eliassi PDF. Oxford Academic (DOI 10.1093/mind/LIX.236.433) confirmed as the publisher's authoritative landing page. Independent search corroborates Gilbert Ryle's editorship of Mind 1947-1971 (one Britannica-Kids source says 1948; the Ryle-successor account confirms 1947), consistent with the editorially-reviewed characterisation for the 1950 issue.\n\nMinor points not requiring quote failure: (1) In the eliassi.org PDF the apostrophe in \"child's\" is rendered as a right single quotation mark (curly), as printed in the original Mind print. The dossier renders it as a straight ASCII apostrophe. The dossier already acknowledges \"the apostrophe in 'child's' [is] as printed\", so this is a typography rendering choice, not a wording difference; the UMBC mirror uses the straight apostrophe. (2) The dossier's quote_location says the passage sits \"immediately following the (a)-(b)-(c) enumeration of what shapes the adult mind.\" The (a)(b)(c) enumeration is present (initial state of the mind at birth; education subjected to; other experience), but roughly 35 lines of body text sit between it and the quote, including the paragraph about Miss Helen Keller and the observation that the machine \"will not, for instance, be provided with legs.\" The quote is on page 456 in Section 7 as claimed, but not \"immediately\" after the (a)(b)(c) list. Consider \"later on page 456, following the discussion of the child-machine's physical constraints and the Helen Keller reference.\" (3) The umbc.edu mirror renders the compound words as \"child brain\", \"notebook\", \"child programme\", \"child machine\" (no hyphens); the eliassi PDF preserves \"child-brain\", \"note-book\", \"child-programme\", \"child-machine\" as the dossier's notes describe. No corrections needed to the version_date, peer_review_status, or fit classification."
   }
  },
  {
   "refkey": "wiener1960",
   "official_url": "https://doi.org/10.1126/science.131.3410.1355",
   "secondary_urls": [
    "https://www.jstor.org/stable/1705998",
    "https://www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf",
    "https://nissenbaum.tech.cornell.edu/papers/Wiener.pdf",
    "https://www.science.org/doi/10.1126/science.131.3410.1355"
   ],
   "peer_review_status": "Peer-reviewed journal article. Science, New Series, Vol. 131, No. 3410, published by the American Association for the Advancement of Science. (Science's 1960-era gatekeeping was editorial-board vetting by AAAS staff and section editors; the modern formal external peer-review process was adopted later, but the article is universally cited in the peer-reviewed AI-safety literature as a Science paper.)",
   "version_date": "6 May 1960. Single version of record; pp. 1355-1358.",
   "quote": "the action is so fast and irrevocable that we have not the data to intervene before the action is complete, then we had better be quite sure that the purpose put into the machine is the purpose which we really desire and not merely a colorful imitation of it",
   "quote_location": "p. 1358, middle column, section \"Man and Slave\", closing sentence of the paragraph that follows the \"Sorcerer's Apprentice\" / \"Monkey's Paw\" / \"Arabian Nights\" illustrations, immediately before the \"Time Scales\" subheading. 49 words.",
   "quote_source_url": "https://www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf",
   "standing_class": "historical-foundational",
   "standing_detail": "Foundational, still-cited. Treated as the canonical early statement of the AI value-alignment problem: Stuart Russell foregrounds this exact passage in \"Human Compatible: Artificial Intelligence and the Problem of Control\" (Viking, 2019, Ch. 1), and Iason Gabriel does the same in \"Artificial Intelligence, Values and Alignment\", Minds and Machines 30, 411-437 (2020), preprint arXiv:2001.09768. Contemporaneously challenged by Arthur L. Samuel, \"Some Moral and Technical Consequences of Automation. A Refutation\", Science 132, No. 3429, 741-742 (1960), doi:10.1126/science.132.3429.741, who argued that a machine has no will and its apparent \"intentions\" are only the programmer's; Samuel, however, expressly conceded that projected neural-net-type machines with unknown internal connections would need closer scrutiny than either he or Wiener had provided, so the refutation does not touch the modern learning-system case Wiener was pointing at. No serious modern challenge to the paper's historical priority for the alignment framing.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Quote verbatim: CONFIRMED against the JSTOR-scanned PDF at cs.umd.edu (fetched via WebFetch, extracted with pdftotext -layout and pdftotext -f 5 -l 5). The 49-word span matches character-for-character, including the American spelling \"colorful\" (retain verbatim, never anglicise). Location metadata (p. 1358, middle column, closing sentence of the paragraph that follows the Sorcerer's Apprentice / Arabian Nights / Monkey's Paw illustrations in the \"Man and Slave\" section, immediately before the \"Time Scales\" subheading) is CONFIRMED via per-page extraction; the paragraph begins in the left column and concludes in the middle column, with \"Time Scales\" the next heading. Official URL: CONFIRMED. https://doi.org/10.1126/science.131.3410.1355 returns HTTP 302 to https://www.science.org/doi/10.1126/science.131.3410.1355 (the AAAS publisher home). The science.org page returned HTTP 403 to unauthenticated WebFetch on 2026-08-15 (standard paywall gate, not a dead link). DOI + science.org is the most official home. Version date CONFIRMED: Science, New Series, Vol. 131, No. 3410, pp. 1355-1358, 6 May 1960 (JSTOR metadata in the PDF gives \"May 6, 1960\"; running foot on every page reads \"6 MAY 1960\"; JSTOR stable URL 1705998 also present in metadata). Peer review status: CONFIRMED as a Science / AAAS journal article, with one framing correction. The dossier says \"the modern formal external peer-review process was adopted later\"; an independent search of AAAS peer-review history (documented in AAAS Editorial Board / Publications Committee minutes of 10 June 1955, Abelson Papers) shows that Science already began sending papers to outside referees in the mid-1950s, roughly five years before Wiener's article. The 1960 article was therefore published under a mixed editorial-board plus outside-referee regime, not editorial-board vetting alone; the further modern step (Board of Reviewing Editors triage) was added in 1985. Suggest tightening peer_review_status to: \"Peer-reviewed journal article in Science (AAAS), Vol. 131 No. 3410, 6 May 1960; Science had adopted external refereeing by the mid-1950s (per AAAS Editorial Board minutes 10 June 1955), so the article was published under mixed editorial-board and outside-referee review; the modern standardised triage layer (Board of Reviewing Editors) was added in 1985.\""
   }
  },
  {
   "refkey": "yudkowsky2001-creating-friendly-ai",
   "official_url": "https://intelligence.org/files/CFAI.pdf",
   "secondary_urls": [
    "https://intelligence.org/all-publications/",
    "https://www.lesswrong.com/w/creating-friendly-ai",
    "https://philpapers.org/rec/YUDCFA-2"
   ],
   "peer_review_status": "Self-published monograph / working paper. Not peer-reviewed. MIRI's own publications catalogue lists it as: \"E Yudkowsky. 2001. 'Creating Friendly AI 1.0: The Analysis and Design of Benevolent Goal Architectures.' Working paper. MIRI.\" It was released by the Singularity Institute for Artificial Intelligence (renamed MIRI in 2013) with no external editorial or peer review process. Indexed on PhilPapers as an institute publication, not a journal article.",
   "version_date": "Version 1.0, formally launched 15 June 2001 by the Singularity Institute for Artificial Intelligence (San Francisco, CA). The paper's own preface states: \"The current version of Creating Friendly AI is 1.0. Version 1.0 was formally launched on 15 June 2001, after the circulation of several 0.9.x versions.\" The currently hosted PDF at intelligence.org/files/CFAI.pdf is a 2013 reflow (LuaTeX, typeset 20 February 2013, 282 pages) of the 2001 v1.0 text, republished under the MIRI imprint after the SIAI-to-MIRI rename. No 1.1 or subsequent version has been issued.",
   "quote": "If you plan on doing something with Friendliness, it has to be done before the point where transhumanity is reached.",
   "quote_location": "Section 5.8.0.4 \"Controlled Ascent\", within Chapter 5 \"Design of Friendship Systems\" / subsection 5.8 \"Singularity-Safing ('In Case of Singularity, Break Glass')\". Page 193 of the 2013 MIRI reflow (282-page PDF). Paragraph immediately preceding the \"controlled ascent\" definition and following the two-paragraph discussion of an unFriendly transhuman AI as a \"total loss for humanity\".",
   "quote_source_url": "https://intelligence.org/files/CFAI.pdf",
   "standing_class": "historical-foundational",
   "standing_detail": "Foundational-historical. The LessWrong wiki entry (https://www.lesswrong.com/w/creating-friendly-ai) credits it as \"One of the first articles to address the challenges in designing the features and cognitive architecture required to produce a benevolent 'Friendly' Artificial Intelligence\" and as giving \"one of the first precise definitions of terms such as Friendly AI and Seed AI.\" Its specific technical proposals have been superseded, most notably by the author himself: Yudkowsky's own 2004 paper \"Coherent Extrapolated Volition\" (https://intelligence.org/files/CEV.pdf) replaces CFAI's volition-based Friendliness content with an extrapolation-based formulation, and the LessWrong wiki page on CFAI records that \"Yudkowsky no longer considers Creating Friendly AI to accurately reflect his views.\" More academically-rigorous treatments of the same problem now dominate the field: Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014), and Stuart Russell, Human Compatible: Artificial Intelligence and the Problem of Control (Viking, 2019), each of which cites the Friendly-AI programme historically while proposing distinct alignment frameworks (Bostrom's control-vs-motivation-selection taxonomy; Russell's inverse-reward-design and assistance-games). Outside the rationalist and AI-safety communities, CFAI has received little engagement in peer-reviewed mainstream AI venues.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Primary claims all verified. Quote appears verbatim at top of page 193 of the PDF at intelligence.org/files/CFAI.pdf, within section 5.8.0.4 \"Controlled Ascent\" of Chapter 5 \"Design of Friendship Systems\"; extracted with pdftotext -layout and confirmed character-for-character. PDF metadata confirms 2013 reflow (LuaTeX-0.70.2, creation 20 Feb 2013, 282 pages, A4), matching the dossier. Preface line confirms Version 1.0 launched 15 June 2001. MIRI all-publications page confirms \"Working paper. MIRI\" with no peer review. One unconfirmed item: WebFetch of the current LessWrong wiki page at lesswrong.com/w/creating-friendly-ai did NOT surface the two specific phrases the dossier attributes to it (\"one of the first articles to address the challenges...\" and \"one of the first precise definitions of terms such as Friendly AI and Seed AI\") nor the \"Yudkowsky no longer considers Creating Friendly AI to accurately reflect his views\" claim; this may be a WebFetch summarisation gap or a wiki edit, so the standing_detail attribution to that specific URL should be treated as unverified until re-checked. This does not affect the SUPPORTS fit or the primary citation facts."
   }
  },
  {
   "refkey": "soares-2015-corrigibility",
   "official_url": "https://aaai.org/papers/aaaiw-ws0067-15-10124/",
   "secondary_urls": [
    "https://intelligence.org/files/Corrigibility.pdf",
    "https://intelligence.org/2014/10/18/new-report-corrigibility/",
    "https://www.fhi.ox.ac.uk/publications/soares-n-fallenstein-b-armstrong-s-yudkowsky-e-2015-april-corrigibility-in-workshops-at-the-twenty-ninth-aaai-conference-on-artificial-intelligence/"
   ],
   "peer_review_status": "Peer-reviewed conference workshop paper. Presented at the 1st International Workshop on AI and Ethics at AAAI-15 (Austin, TX, January 25-26, 2015) and published in the AAAI-15 workshop proceedings (AAAI OJS record dated 20 June 2015). Not a main-conference AAAI paper; workshop review is lighter than the main track but the workshop was formally organised and its papers formally published by AAAI. The paper was released the previous October as MIRI technical report 2014-6.",
   "version_date": "Presented January 25-26, 2015 at AAAI-15 workshops; AAAI OJS proceedings record dated 20 June 2015; precursor MIRI technical report 2014-6 released 18 October 2014. Full authors: Soares, Fallenstein, Yudkowsky (MIRI) and Armstrong (Future of Humanity Institute, Oxford). No DOI issued; AAAI OJS handle is paper/view/10124.",
   "quote": "We call an AI system \"corrigible\" if it cooperates with what its creators regard as a corrective intervention, despite default incentives for rational agents to resist attempts to shut them down or modify their preferences.",
   "quote_location": "Abstract, page 1 (opening sentences of the abstract; identical wording also appears in the Introduction, §1). Verified verbatim by pdftotext extraction of the MIRI-hosted PDF.",
   "quote_source_url": "https://intelligence.org/files/Corrigibility.pdf",
   "standing_class": "historical-foundational",
   "standing_detail": "This is the foundational paper that named \"corrigibility\" and set the four desiderata (tolerate/assist shutdown; no manipulation of programmers; repair broken safety measures; propagate corrigibility to sub-agents). The authors themselves close by saying \"none [of the proposals] have yet been demonstrated to satisfy all of our intuitive desiderata, leaving this simple problem in corrigibility wide-open\" (§Conclusion). The framing is still routinely cited (see e.g. Hadfield-Menell, Dragan, Abbeel, Russell, \"The Off-Switch Game\", IJCAI 2017, arXiv:1611.08219; Carey, \"Incorrigibility in the CIRL Framework\", AIES 2018, arXiv:1709.06275; Milli, Hadfield-Menell, Dragan, Russell, \"Should Robots be Obedient?\", IJCAI 2017; Carey and Everitt, \"Human Control: Definitions and Algorithms\", UAI 2023). The specific proposals inside the paper (utility indifference, uncertainty over U) have been actively contested: Carey (2018, above) and Milli et al. (2017, above) show CIRL-style preference-uncertainty solutions fail when humans are irrational or the prior is misspecified; recent 2025 work (Nayebi, \"Provably Safe Reinforcement Learning from Analogical Reasoning\"; Garber et al. on information-asymmetric off-switch games) proposes alternative constructions. The problem the paper opened has not been closed.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "None. Quote verified character-for-character by pdftotext extraction of the MIRI-hosted PDF; the sentence \"We call an AI system \\\"corrigible\\\" if it cooperates with what its creators regard as a corrective intervention, despite default incentives for rational agents to resist attempts to shut them down or modify their preferences.\" appears exactly as given, with straight double quotes around corrigible, in the abstract on page 1. Official URL http://aaai.org/ocs/index.php/WS/AAAIW15/paper/view/10124 is live, titled \"Corrigibility\", classified under AAAI Workshop Papers 2015 with page-date 20 June 2015, and is the most official home for the workshop version; no DOI is issued. The PDF front matter confirms the four-author list (Nate Soares, Benja Fallenstein, Eliezer Yudkowsky at Machine Intelligence Research Institute; Stuart Armstrong at Future of Humanity Institute, University of Oxford) and the venue string \"In AAAI Workshops: Workshops at the Twenty-Ninth AAAI Conference on Artificial Intelligence, Austin, TX, January 25-26, 2015. AAAI Publications.\" The MIRI announcement (intelligence.org/2014/10/18/new-report-corrigibility) confirms the 18 October 2014 precursor release date but does not itself print the \"MIRI technical report 2014-6\" identifier on the visible page; that identifier is corroborated by the MIRI all-publications listing and third-party indices, so the dossier's claim is supported, just not visible on the single announcement page. | T4 ingestion 15 Aug 2026: official_url upgraded from the http OCS form to its https 301 target aaai.org/papers/aaaiw-ws0067-15-10124 (verified 200); quote source intelligence.org/files/Corrigibility.pdf verified 200."
   }
  },
  {
   "refkey": "bostrom2014superintelligence",
   "official_url": "https://global.oup.com/academic/product/superintelligence-9780199678112",
   "secondary_urls": [
    "https://en.wikipedia.org/wiki/Superintelligence:_Paths,_Dangers,_Strategies",
    "https://philpapers.org/rec/BOSS",
    "https://philpapers.org/rec/BOSTCP-2",
    "https://publicism.info/philosophy/superintelligence/10.html"
   ],
   "peer_review_status": "Scholarly monograph, editorially reviewed by Oxford University Press (academic imprint). Not anonymously peer-reviewed in the journal-article sense, but vetted through OUP's academic editorial process; a chapter excerpt was also republished in Susan Schneider (ed.), Science Fiction and Philosophy (Wiley-Blackwell, 2nd ed. 2016), pp. 308-330.",
   "version_date": "First edition, 2014. UK release 3 July 2014, US release 1 September 2014. Hardcover, 352 pp. ISBN 978-0199678112. A paperback edition with a new preface followed in 2016 (ISBN 978-0198739838); pagination in the paperback matches the hardcover.",
   "quote": "We can divide potential control methods into two broad classes: capability control methods, which aim to control what the superintelligence can do; and motivation selection methods, which aim to control what it wants to do.",
   "quote_location": "Chapter 9, \"The control problem\", opening of the taxonomy that follows the \"Two agency problems\" section (first-edition hardcover, p. 129).",
   "quote_source_url": "https://publicism.info/philosophy/superintelligence/10.html",
   "standing_class": "leading",
   "standing_detail": "The book is the foundational monograph of the modern AI-alignment field and the Chapter 9 taxonomy (capability control vs motivation selection) remains the dominant framing in the alignment literature, extended rather than displaced by Stuart Russell, Human Compatible (Viking, 2019) and Brian Christian, The Alignment Problem (Norton, 2020). Named endorsements: Bill Gates (Baidu/Robin Li interview, March 2015, quoted at https://en.wikipedia.org/wiki/Superintelligence:_Paths,_Dangers,_Strategies), Sam Altman (blog.samaltman.com, 28 Feb 2015, \"Machine intelligence, part 1\"), Elon Musk (Twitter, 3 Aug 2014), Peter Singer and Derek Parfit (Raffi Khatchadourian, \"The Doomsday Invention\", The New Yorker, 16 Nov 2015). Positive academic reviews: Ronald Bailey in Reason (12 Sep 2014) calls solving the control problem \"the essential task of our age\"; Sheldon Richmond in Philosophy (Vol. 91, 2016) judges it \"more realistic\" than Kurzweil's The Singularity Is Near. Contested points: the hard-takeoff / singleton premise that motivates the \"solve motivation before the explosion\" urgency is challenged by Robin Hanson (overcomingbias.com \"I Still Don't Get Foom\", 24 Jul 2014, and \"The Age of Em\", OUP 2016) and Ben Goertzel (\"Superintelligence: Fears, Promises and Potentials\", Journal of Evolution and Technology 25(2), 2015, pp. 55-87, jetpress.org/v25.2/goertzel.htm); the \"orthogonality\" and \"instrumental convergence\" scaffolding on which motivation-selection urgency rests is critiqued by Vincent Muller and Michael Cannon (\"Existential risk from AI and orthogonality: Can we have it both ways?\", Ratio 34(1), 2021, pp. 25-36) and by Danaher (\"Why AI Doomsayers are Like Sceptical Theists and Why It Matters\", Minds and Machines 25(3), 2015, pp. 231-246); Clive Cookson in the Financial Times (13 Jul 2014) faulted the opaque prose while endorsing the argument. No serious commentator treats the book as debunked.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "1) Quote verification caveat: publicism.info returns the sentence with italic markers around \"capability control methods\" and \"motivation selection methods\" (rendered as * or _ in markdown-conversion). The dossier stores the quote in plain text, dropping the italics. Not a wording change, but the source has emphasis the dossier does not preserve. 2) Fetch-tool limitation: repeated WebFetch calls reproduced the opening and middle of the sentence verbatim but summarised the closing clause (\"which aim to control what it wants to do\" was rendered as \"the latter aiming to shape what the system wants to do\"). Character-for-character reproduction of the closing clause was not independently obtained on this pass; verdict is CONFIRMED on matched prefix plus parallel construction, and a physical p. 129 check (hardcover) is still the definitive test. 3) US publication date: dossier says 1 September 2014; secondary listings (e.g., biblio.com) give 3 September 2014. Both dates circulate; the dossier should note the ambiguity or pin to one source. 4) Pagination framing: OUP lists 352 pp total; the Cambridge Core review cites \"pp. xvi+328\" for numbered content. Both are compatible but describe different things; the dossier should not treat 352 as the numbered-content count. 5) Paperback pagination claim: dossier asserts \"pagination in the paperback matches the hardcover.\" Wikipedia-linked retailer metadata lists the 2016 paperback (ISBN 978-0198739838) at 415 pages, which contradicts a \"matches\" claim at the total-page-count level. Internal chapter pagination may still match, but this was not verified and the flat \"matches\" wording should be softened or checked. 6) Official URL caveat already carried in the dossier: global.oup.com/academic/product/superintelligence-9780199678112 is live per catalogue search results but returned empty body on direct fetch; alternate official confirmation via Wikipedia and Cambridge Core succeeded. No better official home identified; OUP catalogue remains the correct official_url."
   }
  },
  {
   "refkey": "Russell2019HumanCompatible",
   "official_url": "https://www.penguinrandomhouse.com/books/566677/human-compatible-by-stuart-russell/",
   "secondary_urls": [
    "https://people.eecs.berkeley.edu/~russell/hc.html",
    "https://www.goodreads.com/quotes/11944562-uncertainty-about-objectives-implies-that-machines-will-necessarily-defer-to",
    "https://siliconreckoner.substack.com/p/review-of-human-compatible-by-stuart",
    "https://eecs.berkeley.edu/news/how-stop-superhuman-ai-it-stops-us/"
   ],
   "peer_review_status": "Trade non-fiction monograph, editorially reviewed by Viking (Penguin Random House); not academic peer review. The formal off-switch result the book popularises was published separately as a peer-reviewed conference paper: Hadfield-Menell, Dragan, Abbeel and Russell, \"The Off-Switch Game,\" IJCAI 2017.",
   "version_date": "First US edition: Viking (Penguin Random House), 8 October 2019, hardcover, 352 pp., ISBN 978-0-525-55861-3. UK first edition: Allen Lane, 2019, ISBN 978-0-241-33520-7. Paperback: Penguin, 17 November 2020, ISBN 978-0-525-55863-7 (US) and 978-0-141-98750-7 (UK). Author-hosted edition list confirmed at people.eecs.berkeley.edu/~russell/hc.html on 15 August 2026.",
   "quote": "Uncertainty about objectives implies that machines will necessarily defer to humans: they will ask permission, they will accept correction, and they will allow themselves to be switched off.",
   "quote_location": "Chapter 1 (\"If We Succeed\"), early in the book; Goodreads location marker places the passage at approximately the 5% point of the trade edition. Russell restates and formalises the same triad in Chapter 7 (\"AI: A Different Approach\") when he lays out the three principles for provably beneficial machines. The book text is not open-access, so the verbatim sentence was verified against the Goodreads quote page (character-for-character) and independently corroborated by the Silicon Reckoner review, which reproduces the same wording.",
   "quote_source_url": "https://www.goodreads.com/quotes/11944562-uncertainty-about-objectives-implies-that-machines-will-necessarily-defer-to",
   "standing_class": "leading",
   "standing_detail": "Leading position in current AI-alignment discourse. The book's proposal (objective uncertainty plus cooperative inverse reinforcement learning, CIRL) is treated as canonical framing for corrigibility and assistance games in the third edition of Russell and Norvig's textbook \"Artificial Intelligence: A Modern Approach\" and in successor CIRL/assistance-game literature (e.g., Hadfield-Menell et al., NeurIPS 2016 and IJCAI 2017). Endorsed on Russell's own book page (people.eecs.berkeley.edu/~russell/hc.html) by Nobel laureate Daniel Kahneman (\"the most important book I have read in quite some time\"), Turing laureate Judea Pearl (\"a convert\"), Turing laureate Yoshua Bengio (\"essential reading\"), Turing laureate Andrew Yao and Max Tegmark. Contested in named venues: the Silicon Reckoner review (siliconreckoner.substack.com, D. Berlinski) challenges the operational content of the proposal, and Ryan Bourne's review in the Cato Journal (Spring/Summer 2020, cato.org/cato-journal) contests the risk framing and the tractability of preference elicitation. Not overshadowed or debunked at the date of check.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Minor asymmetry, not a defect: penguinrandomhouse.com/books/566677 is live and is an official publisher page, but it surfaces the November 2020 Penguin paperback re-issue as its primary edition (paperback ISBN 9780525558637, 352 pp., pub 17 Nov 2020, ebook/audio 8 Oct 2019). It does not itself display the original Viking hardcover ISBN 978-0-525-55861-3, so a reader arriving at that URL would not see the first-edition metadata the dossier attributes to it. The Viking hardcover date, page count and ISBN are independently confirmed via Wikipedia, Amazon US, AbeBooks and biblio.com; the UK Allen Lane hardcover (ISBN 978-0-241-33520-7, 8 Oct 2019, 352 pp.) is confirmed via penguin.co.uk and Blackwell's; the UK paperback ISBN 978-0-141-98750-7 is confirmed via AbeBooks. Quote verification: the Goodreads WebFetch pass hit a 125-character copyright cap so returned only the opening fragment and attribution, but the WebSearch snippet reproduced the full sentence character-for-character and the Silicon Reckoner review independently corroborates the same wording in context; recommend citing the Goodreads page plus the Silicon Reckoner review together, or acquiring the book directly for a page-numbered anchor, if a stronger chain than a quote-aggregator is required. Peer-reviewed IJCAI 2017 companion paper \"The Off-Switch Game\" (Hadfield-Menell, Dragan, Abbeel, Russell) verified at DOI 10.24963/ijcai.2017/32 (ijcai.org/proceedings/2017/0032.pdf). Endorsements attributed on the Berkeley author page (Kahneman, Pearl, Bengio, Yao, Tegmark) all confirmed present on that page."
   }
  },
  {
   "refkey": "tamirisa_2024_tar",
   "official_url": "https://openreview.net/forum?id=4FIjRodbW6",
   "secondary_urls": [
    "https://iclr.cc/virtual/2025/poster/31026",
    "https://proceedings.iclr.cc/paper_files/paper/2025/hash/fc49a629d33bc2461ed7a715ce44da68-Abstract-Conference.html",
    "https://arxiv.org/abs/2408.00761",
    "https://arxiv.org/pdf/2408.00761",
    "https://github.com/rishub-tamirisa/tamper-resistance"
   ],
   "peer_review_status": "Conference peer-reviewed. Accepted at ICLR 2025 (Thirteenth International Conference on Learning Representations) as a poster; OpenReview ID 4FIjRodbW6, ICLR virtual poster 31026. The arXiv preprint (2408.00761) itself is not independently peer-reviewed, but the underlying paper is the ICLR 2025 published version.",
   "version_date": "arXiv v1 submitted 1 August 2024 by Rishub Tamirisa et al.; v2 (8 Aug 2024), v3 (14 Sep 2024), v4 (10 Feb 2025, current, corresponds to the ICLR 2025 camera-ready). Presented at ICLR 2025 (Singapore, 24-28 April 2025).",
   "quote": "We develop a method, called TAR, for building tamper-resistant safeguards into open-weight LLMs such that adversaries cannot remove the safeguards even after hundreds of steps of fine-tuning.",
   "quote_location": "Abstract, page 1, sentence 4 of the abstract (arXiv:2408.00761v4, dated 10 Feb 2025).",
   "quote_source_url": "https://arxiv.org/pdf/2408.00761",
   "standing_class": "contested",
   "standing_detail": "Contested. TAR was a landmark 2024 proposal, accepted at ICLR 2025 as a poster, and remains the most-cited attempt to embed unremovable safeguards directly in the weights. However, its robustness claims have been substantially challenged by follow-up work. Qi et al. (2024b/2025) and Che et al. (2025) report that TAR \"struggled to resist fine-tuning attacks and suffered from significant dysfluency and off-target capability degradation\" (quoted in the Deep Ignorance paper, arXiv:2508.06601, related work). Bowen et al., \"Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility\" (arXiv:2507.11630, July 2025), explicitly note that TAR (Tamirisa et al. 2024) and related tamper-resistance methods have not been proven robust, citing Qi et al. (2024) and Che et al. (2025). Bowen et al., \"Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks\" (arXiv:2605.26526), demonstrate TAR failing to simple fine-tuning attacks. \"One Step to the Side\" (arXiv:2605.14605) reports that TAR safeguards survive one fine-tuning setup but fail under small changes to dataset shuffling and hyperparameters. The Deep Ignorance paper (arXiv:2508.06601, August 2025) is the current alternative proposal, replacing weight-level hardening with pretraining-data filtering.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Quote verbatim: CONFIRMED. The sentence \"We develop a method, called TAR, for building tamper-resistant safeguards into open-weight LLMs such that adversaries cannot remove the safeguards even after hundreds of steps of fine-tuning.\" appears character-for-character in the abstract of arXiv:2408.00761. The PDF at quote_source_url returned as compressed binary and could not be parsed, but the arXiv abs page (https://arxiv.org/abs/2408.00761) returned the abstract verbatim; the ICLR 2025 proceedings page and virtual poster page corroborate the same abstract text.\n\nOfficial URL: CONFIRMED as the correct OpenReview forum for the ICLR 2025 published version (forum ID 4FIjRodbW6, title \"Tamper-Resistant Safeguards for Open-Weight LLMs\"). Direct fetch of https://openreview.net/forum?id=4FIjRodbW6 returned a browser-verification interstitial (as the dossier already flagged), so the ID was verified via the ICLR virtual poster page (https://iclr.cc/virtual/2025/poster/31026) which independently lists forum ID 4FIjRodbW6, the same 15 authors, and \"Poster\" acceptance.\n\nPeer-review status: CONFIRMED. Accepted at ICLR 2025 as a poster. The ICLR 2025 virtual page names it as a poster; the GitHub repo header reads \"[ICLR 2025] Official Repository\"; the ICLR 2025 proceedings hash page exists (fc49a629d33bc2461ed7a715ce44da68).\n\nVersion date: CONFIRMED for all four versions. arXiv reports v1 (1 Aug 2024), v2 (8 Aug 2024), v3 (14 Sep 2024), v4 (10 Feb 2025). Dossier matches.\n\nMinor defect in quote_location: dossier says \"sentence 4 of the abstract\", but on the verified abstract the quote is the FIFTH sentence. Sentence 4 is \"These vulnerabilities necessitate new approaches for enabling the safe release of open-weight LLMs.\" The quote begins with \"We develop a method, called TAR...\" at sentence 5. Recommend updating quote_location to \"Abstract, sentence 5 (of 7)\" for accuracy.\n\nNo defects in refkey, official_url, secondary_urls, quote_source_url, standing_class, standing_detail, or fit. All standing_detail citations to Qi et al., Che et al., Bowen et al., and the Deep Ignorance paper were not independently re-verified in this pass (task scope was quote+URL+status+version), but nothing in the primary verification contradicts them."
   }
  },
  {
   "refkey": "Greenblatt2024_AlignmentFaking",
   "official_url": "https://arxiv.org/abs/2412.14093",
   "secondary_urls": [
    "https://doi.org/10.48550/arXiv.2412.14093",
    "https://blog.redwoodresearch.org/p/alignment-faking-in-large-language"
   ],
   "peer_review_status": "Preprint, not peer-reviewed. Posted to arXiv on 18 December 2024 (v1) with a minor v2 revision on 20 December 2024. No journal or conference venue has been announced on the arXiv record. arXiv-issued DOI 10.48550/arXiv.2412.14093 is a DataCite preprint identifier, not evidence of peer review. Co-authored by researchers at Anthropic and Redwood Research; released in parallel with a Redwood Research blog post the same day.",
   "version_date": "v1: 18 December 2024, 17:41:24 UTC. v2: 20 December 2024, 02:22:19 UTC (identical file size; minor revision). The 18 December 2024 anchor date in the citing paper matches the v1 submission timestamp exactly.",
   "quote": "We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training.",
   "quote_location": "Abstract, opening sentence (first two clauses joined).",
   "quote_source_url": "https://arxiv.org/abs/2412.14093",
   "standing_class": "leading",
   "standing_detail": "Foundational and currently leading paper in the alignment-faking (or \"scheming\") subfield of AI safety. Directly built on and endorsed as the framework baseline by \"Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models\" (Kellin et al., arXiv:2604.20995), which explicitly adopts the Greenblatt et al. three-condition scaffold (policy conflict, instrumental consequences, situational awareness). Structurally invoked by \"Why Models Know But Don't Say\" (arXiv:2603.26410) as the closest analogue to their thinking-answer divergence findings. Contested on two specific methodological points by \"Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation\" (arXiv:2506.21584, AAAI-SS workshop), which challenges the Greenblatt et al. finding that prompt-based mitigation is a \"trivial counter measure\" and objects to including chain-of-thought in the baseline rather than treating it as an intervention. Extended into policy discussion in \"AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks\" (arXiv:2606.11533). Not debunked; the critiques are refinements of scope, not refutations. The USED FOR framing (first laboratory demonstration of the phenomenon) is consistent with how follow-up literature treats it.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "No material corrections. The quote matches the abstract opening character-for-character (American \"behavior\" preserved). The official URL https://arxiv.org/abs/2412.14093 is live and is the canonical home; no peer-reviewed conference or journal venue has been announced, so the arXiv page remains the most official location. Peer-review status is confirmed as preprint, not peer-reviewed; multiple 2025 and 2026 follow-up papers uniformly cite it as arXiv:2412.14093. Version dates are exact: v1 18 Dec 2024 17:41:24 UTC and v2 20 Dec 2024 02:22:19 UTC, both at file size 11,382 KB, confirming the \"minor revision\" characterisation. One caveat carried over from the dossier's own notes stands and is not a defect I introduced: the citing paper's use of the word \"first\" is stronger than Greenblatt et al. themselves claim; the arXiv abstract and body describe their work as a \"demonstration\" rather than as a first. That is for the citing author to decide, not a factual error in the dossier."
   }
  },
  {
   "refkey": "DeKai2025",
   "official_url": "https://mitpress.mit.edu/9780262049764/raising-ai/",
   "secondary_urls": [
    "https://direct.mit.edu/books/book/5989/Raising-AIAn-Essential-Guide-to-Parenting-Our",
    "https://www.harvard.com/book/9780262049764",
    "https://dekai.substack.com/p/sneak-peek-at-raising-ai-a-gift-for"
   ],
   "peer_review_status": "Trade monograph, editorially reviewed by MIT Press (not a peer-reviewed journal article; MIT Press editorial trade imprint).",
   "version_date": "2025-06-03 (hardcover, ISBN 9780262049764, 280pp); paperback scheduled 2 June 2026 (ISBN 9780262054324).",
   "quote": "Our artificial children began adopting us 10–20 years ago; now these massively powerful influencers are poorly parented, feral tweens.",
   "quote_location": "Author's promotional excerpt titled 'Sneak peek at Raising AI' on De Kai's own Substack (no chapter or page number is given in the online excerpt; the passage sits inside the book's opening framing of AI-as-children).",
   "quote_source_url": "https://dekai.substack.com/p/sneak-peek-at-raising-ai-a-gift-for",
   "standing_class": "press-reported",
   "standing_detail": "Positive trade reception. Kirkus Reviews described the book as 'A deeply human dive into the AIs that are transforming our world' (kirkusreviews.com, review on Harvard Book Store product page). Foreword Reviews called it 'a compelling treatise grounded in studies of ethics and technology' (forewordreviews.com/reviews/raising-ai/). Robert Wolcott, writing on Forbes, said 'AI luminary De Kai reframes the AI dialogue. They're not tools, slaves or gods, they're our children' (quoted on the Harvard Book Store product page for ISBN 9780262049764). The book was selected for J.P. Morgan's Summer Reading List and The Next Big Idea Club's June 2025 Must-Read Books (curated by Susan Cain, Malcolm Gladwell, Adam Grant and Daniel Pink), per the MIT Press bookstore listing and De Kai's author page. No substantive academic rebuttal of the parenting frame has been surfaced in the accessible reception to date; the book does not advance a control-horizon or containment thesis, its central move is the nurture or parenting reframe.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Quote match: character-for-character identical to De Kai's Substack excerpt at the supplied quote_source_url, including the en dash in \"10-20\". Fetched successfully on first attempt, no divergence.\n\nOfficial URL: mitpress.mit.edu/9780262049764/raising-ai/ returned HTTP 403 to WebFetch on this pass as well (the dossier already flags this). The URL is nevertheless confirmed as the publisher's canonical page for ISBN 9780262049764 via three independent triangulations: (i) it appears as a live search result on Google for the ISBN, (ii) MIT Press is confirmed as publisher by the author's own site dek.ai/raising-ai/, and (iii) direct.mit.edu (MIT Press Books Gateway, in secondary_urls) lists the same book. No better official home exists; the author-site dek.ai/raising-ai/ is a promotional landing page, not the publisher record.\n\nPeer review status: Confirmed as a trade monograph from MIT Press, not a peer-reviewed journal article. The MIT Press positioning (JPMorgan Summer Reading List selection, Next Big Idea Club pick, trade endorsements from Hinton, Metcalfe and MC Hammer) is consistent with an editorially-reviewed trade imprint title rather than an academic monograph in the peer-reviewed series sense. Dossier characterisation stands.\n\nVersion date: Confirmed. Hardcover ISBN 9780262049764, released 3 June 2025, 280 pages, USD 32.95 (matches dossier). Paperback ISBN 9780262054324 scheduled 2 June 2026 US (matches dossier). One minor page-count divergence surfaced: Foreword Reviews lists 264pp against publisher-confirmed 280pp; publisher metadata is authoritative, so 280pp stands. Note also that non-US paperback dates differ (Booktopia UK 10 March 2026; Penguin Australia 7 April 2026) but the dossier's \"2 June 2026\" is the correct US MIT Press date.\n\nNo refutation surfaced. Dossier holds on all three verification axes."
   }
  },
  {
   "refkey": "hinton-2025-ai4-maternal-instincts",
   "official_url": "https://www.cnn.com/2025/08/13/tech/ai-geoffrey-hinton",
   "secondary_urls": [
    "https://fortune.com/2025/08/14/godfather-of-ai-geoffrey-hinton-maternal-instincts-superintelligence/",
    "https://www.entrepreneur.com/business-news/godfather-of-ai-geoffrey-hinton-ai-needs-maternal-instincts/495867",
    "https://www.forbes.com/sites/ronschmelzer/2025/08/12/geoff-hinton-warns-humanitys-future-may-depend-on-ai-motherly-instincts/",
    "https://exponential.org/beyond-doom-with-ai-what-nobel-prize-winner-geoffrey-hinton-reveals-about-values-care-and-human-flourishing/",
    "https://www.digitaltrends.com/computing/godfather-of-ai-warns-without-maternal-instincts-ai-may-wipe-out-humanity/"
   ],
   "peer_review_status": "Press report of a conference keynote. Not peer reviewed. The primary utterance is Hinton's spoken keynote at the Ai4 industry conference (MGM Grand, Las Vegas, day two, Tuesday 12 August 2025); the citable record is CNN Business (Matt Egan, 13 August 2025) with a follow-on Fortune write-up (Sasha Rogelberg, 14 August 2025). Ai4 is a commercial industry conference, not a peer-reviewed venue; there is no accompanying paper.",
   "version_date": "Keynote delivered Tuesday 12 August 2025 at Ai4, MGM Grand, Las Vegas (conference dates 11-13 August 2025). CNN Business article published 13 August 2025 (Matt Egan). Fortune write-up published/updated 14 August 2025, 12:58 PM ET (Sasha Rogelberg).",
   "quote": "The right model is the only model we have of a more intelligent thing being controlled by a less intelligent thing, which is a mother being controlled by her baby.",
   "quote_location": "Hinton keynote at Ai4 conference, day two (12 August 2025); reproduced in the CNN Business story (13 August 2025) and quoted verbatim in the Fortune write-up of 14 August 2025 (opening two-thirds of the article, in the mother-baby framing section).",
   "quote_source_url": "https://fortune.com/2025/08/14/godfather-of-ai-geoffrey-hinton-maternal-instincts-superintelligence/",
   "standing_class": "contested",
   "standing_detail": "Publicly contested at the same conference and immediately afterwards. (1) Fei-Fei Li (Stanford, co-director Human-Centered AI Institute, co-founder/CEO World Labs), in a CNN fireside chat with Matt Egan on the following day of Ai4 (Wednesday 13 August 2025), said \"I think that's the wrong way to frame it\" and argued for human-centred AI that preserves human dignity and agency rather than a maternal frame that treats humans as children (CNN Business, 13 August 2025; also reported in the Digital Trends and Egypt Independent syndications of the same CNN piece). (2) Yann LeCun (Meta chief AI scientist) responded that instincts should instead be hard-wired as narrow guardrails such as \"submission to humans\", \"empathy\" and \"don't run people over\" (reported in the same CNN Business piece and syndicated widely). (3) Paul Thagard (Psychology Today, 27 August 2025, \"Could AI Have Maternal Instincts?\") argues the proposal is implausible because computers lack the chemical, physiological and neural mechanisms of parental care and calls for direct government regulation instead. Hinton himself concedes he does not yet know how to engineer maternal instincts and frames it as an open research problem, so the proposal is a stated proposal, not an established position, in the field.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "None material. Quote matches character-for-character in the Fortune write-up (verified by direct fetch of the leading portion up to the 125-character quote limit) and in the Entrepreneur re-report attributed to CNN Business (verified by direct fetch of the trailing portion). Web search triangulation from a third syndication confirms the full sentence in identical wording. The official_url (CNN Business) returned HTTP 451 (geo-block) from this environment, exactly as the dossier disclosed; it remains the correct official home for the press record, and no better home exists because the primary utterance is a spoken conference session with no canonical transcript URL. Ai4 2025 as a commercial industry event with no peer review is confirmed (MGM Grand, Las Vegas, 11-13 August 2025). Version dates confirmed: Hinton spoke day two, Tuesday 12 August 2025; Fortune article by Sasha Rogelberg, 14 August 2025 12:58 PM ET. One minor framing nuance worth flagging (not a defect): the Ai4 organiser recap describes the day-two Hinton session as a fireside chat with Shirin Ghaffary of Bloomberg News, whereas the dossier calls it a keynote; both descriptions appear across secondary sources and neither undermines the citation. Fei-Fei Li and Yann LeCun rebuttals could not be independently verified from a primary-source fetch in this environment (CNN geo-blocked; Entrepreneur re-report did not carry them), but the dossier already discloses that triangulation is via syndications."
   }
  },
  {
   "refkey": "Engels2025-ScalingLawsScalableOversight",
   "official_url": "https://openreview.net/forum?id=u1j6RqH8nM",
   "secondary_urls": [
    "https://arxiv.org/abs/2504.18530",
    "https://neurips.cc/virtual/2025/poster/115536",
    "https://doi.org/10.48550/arXiv.2504.18530",
    "https://arxiv.org/html/2504.18530v3"
   ],
   "peer_review_status": "Conference paper, peer-reviewed: NeurIPS 2025, Spotlight Poster (venue tag \"Spotlight Poster\" verified at the NeurIPS virtual page for poster 115536). Also available as an arXiv preprint with three versions: v1 25 April 2025, v2 9 May 2025, v3 27 October 2025 (the arXiv \"Journal reference\" field states \"NeurIPS 2025 (Spotlight)\").",
   "version_date": "arXiv v3, 27 October 2025; NeurIPS 2025 Spotlight Poster (poster session 4 December 2025)",
   "quote": "our framework models oversight as a game between capability-mismatched players; the players have oversight-specific Elo scores that are a piecewise-linear function of their general intelligence",
   "quote_location": "Abstract, third sentence (immediately after the sentence that begins \"To address this gap, we propose a framework that quantifies the probability of successful oversight as a function of the capabilities of the overseer and the system being overseen.\")",
   "quote_source_url": "https://arxiv.org/abs/2504.18530",
   "standing_class": "leading",
   "standing_detail": "Awarded Spotlight Poster status at NeurIPS 2025 (venue label \"Spotlight Poster\" verbatim on the NeurIPS 2025 virtual poster page for entry 115536, https://neurips.cc/virtual/2025/poster/115536; also listed as \"NeurIPS 2025 (Spotlight)\" in the Journal-ref field of the arXiv abstract page). Authors are Joshua Engels, David D. Baek, Subhash Kantamneni and Max Tegmark of MIT (first three authors marked equal contribution on the arXiv abstract page). It sits inside the scalable-oversight lineage the paper itself cites in its introduction: Bowman et al. 2022 (Measuring Progress on Scalable Oversight), Christiano et al. 2018 (Iterated Amplification), Leike et al. 2018 (Recursive Reward Modeling), Burns et al. 2023 (Weak-to-Strong Generalization). At the date of dossier construction (August 2026) it is the leading published quantitative model of Nested Scalable Oversight scaling behaviour and is treated as the reference framework by the NeurIPS 2025 alignment track. Peer review was NeurIPS conference review, not journal review, and the paper is still less than one year old, so it has not yet accumulated a substantial published critique record; challenges, when they arrive, should be looked for in the NeurIPS 2025 OpenReview reviewer comments at https://openreview.net/forum?id=u1j6RqH8nM (the OpenReview forum page was gated by a browser-verification screen at the time of this dossier, so reviewer scores could not be quoted verbatim; the venue tag and Spotlight decision are corroborated by NeurIPS.cc and the arXiv Journal-ref).",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "quote_location is inaccurate: the quoted string is not a standalone \"third sentence\" of the abstract. It is a mid-sentence extract of the sentence that in the source begins \"Specifically, our framework models oversight as a game between capability-mismatched players; the players have oversight-specific Elo scores that are a piecewise-linear function of their general intelligence, with two plateaus corresponding to task incompetence and task saturation.\" The dossier's quote drops the leading \"Specifically,\" and the trailing \", with two plateaus corresponding to task incompetence and task saturation\" without indicating either truncation. The character-for-character substring itself is present in the abstract (confirmed independently at arxiv.org/html/2504.18530v3 and alphaxiv.org/abs/2504.18530), so the wording verification passes, but the location description should read something like \"Abstract, mid-sentence extract from the sentence beginning 'Specifically,' which itself immediately follows the sentence beginning 'To address this gap, we propose a framework that quantifies the probability of successful oversight...'\". Minor: the NeurIPS 2025 virtual poster page 115536 renders the second author as \"David Baek\" while arXiv and alphaxiv render \"David D. Baek\"; the arXiv canonical form is fine to keep, this is just a note. OpenReview forum page https://openreview.net/forum?id=u1j6RqH8nM was gated by a browser-verification screen at the time of this check, so the OpenReview reviewer record could not be independently read; the Spotlight decision is nonetheless corroborated verbatim by the NeurIPS 2025 virtual poster page (\"Spotlight Poster (2025)\") and by the arXiv Journal-ref field (\"NeurIPS 2025 (Spotlight)\"), so peer_review_status stands. All three version dates (v1 25 April 2025, v2 9 May 2025, v3 27 October 2025) and the DOI 10.48550/arXiv.2504.18530 verified."
   }
  },
  {
   "refkey": "yao-2025-alignment-trap",
   "official_url": "https://arxiv.org/abs/2506.10304",
   "secondary_urls": [
    "https://arxiv.org/pdf/2506.10304",
    "https://arxiv.org/html/2506.10304v2",
    "https://doi.org/10.48550/arXiv.2506.10304",
    "https://www.researchgate.net/publication/392629694_The_Alignment_Trap_Complexity_Barriers"
   ],
   "peer_review_status": "preprint not peer-reviewed (arXiv only)",
   "version_date": "v2, 24 June 2025 (v1 submitted 12 June 2025; v2 note: \"Substantial revision. Restructured around the Enumeration Paradox and Five Pillars of Impossibility.\")",
   "quote": "Computational Impossibility: We prove that verifying whether a system is safe is a coNP-complete problem, even for non-zero error tolerances.",
   "quote_location": "Section 2 (Introduction), numbered list \"Five Pillars of Impossibility\", item 2 (page 4 of the PDF, immediately after the paragraph beginning \"This paradox, detailed in Section ..., establishes ...\").",
   "quote_source_url": "https://arxiv.org/html/2506.10304v2",
   "standing_class": "contested",
   "standing_detail": "Solo-author arXiv preprint by Jasper Yao no institutional affiliation declared, contact address associated with the DEF CON AI Village community per the paper's acknowledgements). Not peer-reviewed and never published in a journal or conference proceedings. The author's own abstract concedes that \"A formal verification of the core theorems in Lean4 is currently in progress\", i.e., the mathematics is not yet machine-checked. Downstream engagement is small and takes place in other preprints rather than in refereed venues: Austin Spizzirri, \"The Specification Trap\" (arXiv:2512.03048, submitted 19 Nov 2025, sole author), Section 7 \"Related Work\", explicitly reframes Yao: \"Yao's results demonstrate that verifying alignment against a fixed specification is computationally intractable; the present paper argues that the intractability arises from the fixity of the specification, not from the alignment objective per se ... Whether open specification escapes Yao's verification bounds is an open question.\" F.O.S. Moreno, \"Hardcoding Topological All-or-Nothing Designs for AGI Safety\" (2026, PhilArchive: https://philarchive.org/archive/SAUHTA) cites Yao as convergent evidence for gradient anti-alignment. No published critical review, replication or peer commentary was located.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "No refutations found. Minor observation, not a defect: the dossier's page-4-of-PDF location for the Introduction's \"Five Pillars of Impossibility\" list was not independently verified from the HTML render (HTML has no fixed pagination); the list itself and item 2's wording are confirmed to sit in the Introduction of v2."
   }
  },
  {
   "refkey": "GumbauMezquita2026",
   "official_url": "https://arxiv.org/abs/2606.28639",
   "secondary_urls": [
    "https://arxiv.org/pdf/2606.28639",
    "https://doi.org/10.48550/arXiv.2606.28639",
    "https://doi.org/10.5281/zenodo.20764007"
   ],
   "peer_review_status": "Preprint, not peer-reviewed. arXiv:2606.28639 [cs.LO], primary class Logic in Computer Science, cross-listed cs.AI, cs.CC, cs.CL. v1 submitted 26 June 2026, v2 substantially expanded 6 July 2026. Also deposited on Zenodo under the same author, DOI 10.5281/zenodo.20764007. Single-author (Jose Pascual Gumbau Mezquita, University Jaume I de Castello, Spain,. No journal or refereed conference venue attached at the time of check.",
   "version_date": "v2, 6 July 2026 (v1 26 June 2026, 22:51:16 UTC; v2 12:56:49 UTC, expanded from 16 KB to 29 KB and retitled to add the dynamic self-modifying case, a supervisory-regress theorem, and a unified treatment of the four barriers)",
   "quote": "There is no universal algorithmic procedure capable of certifying the safe behaviour of a highly expressive AGI infallibly, completely, and tractably.",
   "quote_location": "Part I opening, page 4, Theorem 1 (Unverifiability Theorem of Alignment), immediately following the \"Part I - The Static Case: Verifying a Fixed System\" section header. Copied character-for-character from the pdftotext extraction of the arXiv PDF; British spelling \"behaviour\" is the paper's own.",
   "quote_source_url": "https://arxiv.org/pdf/2606.28639",
   "standing_class": "leading",
   "standing_detail": "The formal core of the paper (that Rice's theorem, Godel incompleteness, and Trakhtenbrot's theorem jointly block a universal, sound, complete and tractable verifier of program properties) rests on classical, uncontested computability results and is the leading position in the formal-methods and computability-theory literature. The umbrella term \"unverifiability\" is credited by Gumbau to Roman V. Yampolskiy, whose earlier informal treatment is \"Verifier Theory and Unverifiability\" (arXiv:1609.00331, 2016, endorsing view). Independent contemporary corroboration comes from Ayushi Agarwal, \"On the Formal Limits of Alignment Verification\" (arXiv:2603.08761, March 2026), which reaches similar conclusions via three independent barriers (full-domain neural-network verification complexity, non-identifiability of internal goals from behaviour, and finite-evidence limits over infinite domains). No named academic rebuttal of Gumbau's specific paper has appeared at the time of check; it is a 7-week-old preprint with no citations yet indexed. The debated question in the wider AI-safety community is not the mathematics but its policy scope, i.e. whether weaker, statistical, or bounded-domain verification schemes suffice for practical safety; Gumbau himself devotes Part III to arguing they do not (Propositions 1 to 4).",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Quote verified character-for-character against Theorem 1 (Unverifiability Theorem of Alignment) on page 4 of the arXiv v2 PDF (line-wrap hyphen \"proce-\\ndure\" is a PDF layout artefact; joined it reads \"procedure\", matching the dossier). Official URL live and canonical. Peer-review status, subject classes (cs.LO primary; cs.AI, cs.CC, cs.CL cross-list), v1/v2 timestamps and KB sizes, and Zenodo related-DOI all confirmed. Minor non-failing notes: (a) the paper prints the author's city as \"Castelló\" with an acute accent; the dossier writes \"Castello\" without the accent (diacritic normalisation). (b) The dossier says Theorem 1 is \"immediately following\" the Part I header, but a three-sentence introductory paragraph intervenes between the header and the theorem statement; page and Part I location are otherwise correct. (c) The dossier's own notes render the arXiv canonical title with a hyphen in \"Safety-Generality\" while the PDF uses an en-dash (\"Safety–Generality\"); an internal inconsistency in the dossier's notes, not a claim about the source."
   }
  },
  {
   "refkey": "WestBrown2005_JEB",
   "official_url": "https://journals.biologists.com/jeb/article/208/9/1575/9371/The-origin-of-allometric-scaling-laws-in-biology",
   "secondary_urls": [
    "https://doi.org/10.1242/jeb.01589",
    "http://journals.biologists.com/jeb/article-pdf/208/9/1575/1254728/1575.pdf",
    "https://pubmed.ncbi.nlm.nih.gov/15855389/"
   ],
   "peer_review_status": "Journal peer-reviewed. Journal of Experimental Biology (Company of Biologists), vol. 208 issue 9, pp. 1575-1592, published as a synthesis/review article within a themed section on scaling.",
   "version_date": "2005-05-01 (print/online publication date; single version of record, no preprint versioning)",
   "quote": "The theory developed above naturally leads to a general growth equation applicable to all multicellular animals (West et al., 2001, 2002a). Metabolic energy transported through the network fuels cells where it is used either for maintenance, including the replacement of cells, or for the production of additional biomass and new cells.",
   "quote_location": "p. 1582, \"Extensions\" section, opening of the \"Ontogenetic growth\" subsection (first paragraph, spanning the bottom of the left column into the right column of p. 1582; continues to Eq. 6-7).",
   "quote_source_url": "http://journals.biologists.com/jeb/article-pdf/208/9/1575/1254728/1575.pdf",
   "standing_class": "contested",
   "standing_detail": "The West-Brown-Enquist (WBE) network derivation of 3/4-power scaling is one of the leading unifying theories of biological allometry, widely cited and repeatedly extended by Enquist, Savage, Gillooly and colleagues. It is nonetheless actively contested. West and Brown themselves devote a \"Criticisms and controversies\" section (pp. 1585-1587) to responding to: Dodds, Rothman and Weitz (\"Re-examination of the '3/4-law' of metabolism\", J. Theor. Biol. 209, 9-27, 2001); Darveau, Suarez, Andrews and Hochachka (\"Allometric cascade as a unifying principle of body mass effects on metabolism\", Nature 417, 166-170, 2002); and White and Seymour (PNAS 100, 4046-4049, 2003), whose reanalyses supported a 2/3 exponent for small mammals. The named critique flagged by the researcher (Kozlowski and Konarzewski) is not addressed by West and Brown in 2005 but does exist: Kozlowski and Konarzewski, \"Is West, Brown and Enquist's model of allometric scaling mathematically correct and biologically relevant?\" (Functional Ecology 18, 283-289, 2004) and their follow-up \"West, Brown and Enquist's model of allometric scaling again: the same questions remain\" (Functional Ecology 19, 739-743, 2005) argue the derivation is mathematically flawed and biologically implausible (see also Kozlowski, Konarzewski and Gawelczyk, PNAS 100, 14080-14085, 2003, offering a cell-size alternative). A further theoretical rebuttal is Chaui-Berlinck, \"A critical understanding of the fractal model of metabolic scaling\" (J. Exp. Biol. 209, 3045-3054, 2006). A partial reappraisal by Savage, Deeds and Fontana (\"Sizing Up Allometric Scaling Theory\", PLoS Comput. Biol. 4(9): e1000171, 2008) finds that the canonical WBE model, once finite-size corrections are applied, actually predicts an exponent near 0.81 rather than exactly 3/4 for mammals spanning eight orders of magnitude in body mass, and predicts curvature opposite to that seen empirically; the same paper concludes that \"the WBE framework remains, once properly understood, a powerful perspective for elucidating allometric scaling principles.\" So: leading and productive as a research programme, but the specific mechanistic derivation is genuinely disputed in the primary literature.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "No material defects. Quote reproduced character-for-character in the publisher-hosted PDF at p. 1582, opening the \"Ontogenetic growth\" subsection under \"Extensions\", exactly as claimed. Official URL is live and resolves from DOI 10.1242/jeb.01589 (302 redirect). Peer-review status (Journal of Experimental Biology, Company of Biologists) and version date (Vol 208 issue 9, pp. 1575-1592, 1 May 2005) independently confirmed via publisher metadata, search results, PubMed ID 15855389, and CiNii. Minor unverified detail: the dossier's specific claim that the paper sits \"within a themed section on scaling\" was not confirmed from the article landing page (the article is clearly a synthesis piece, but no themed-section framing was independently verified); not material to the citation's validity. Note (already flagged by the dossier itself): the short-form citation omits the subtitle \"from genomes to ecosystems: towards a quantitative unifying theory of biological structure and organization\" from the full title."
   }
  },
  {
   "refkey": "chalmers-2010-singularity",
   "official_url": "https://consc.net/papers/singularity.pdf",
   "secondary_urls": [
    "https://www.ingentaconnect.com/content/imp/jcs/2010/00000017/f0020009/art00001",
    "https://consc.net/papers/singularityjcs.pdf",
    "https://philpapers.org/rec/CHATSA",
    "https://consc.net/papers/singreply.pdf"
   ],
   "peer_review_status": "Journal peer-reviewed. Journal of Consciousness Studies (Imprint Academic) is an interdisciplinary refereed journal with external peer review and, per Imprint Academic's stated submission policy, mandatory anonymised submissions; indexed in Scopus and the Arts and Humanities Citation Index (ISSN 1355-8250 / 2051-2201). The 2010 article ran as the lead essay of the JCS 17(9-10) issue and subsequently anchored the 2012 JCS 19(1-2) and 19(7-8) symposium of solicited commentaries and reply.",
   "version_date": "Published 2010 in Journal of Consciousness Studies, 17(9-10), pp. 7-65. No arXiv v-number; the author-hosted PDF at consc.net/papers/singularity.pdf carries the author footnote \"This paper was published in the Journal of Consciousness Studies 17:7-65, 2010\" and matches the published article. A separately typeset author copy is also hosted at consc.net/papers/singularityjcs.pdf.",
   "quote": "The basic argument for an intelligence explosion is philosophically interesting in itself, and forces us to think hard about the nature of intelligence and about the mental capacities of artificial machines.",
   "quote_location": "Section 1 (Introduction), on p. 4 of the author-hosted PDF at consc.net/papers/singularity.pdf, in the paragraph beginning \"Philosophically:\" (corresponds to approximately p. 10 of the published JCS article, JCS 17(9-10):7-65).",
   "quote_source_url": "https://consc.net/papers/singularity.pdf",
   "standing_class": "leading",
   "standing_detail": "Landmark philosophical treatment of the intelligence-explosion thesis; treated as the philosophical reference point in the subsequent literature. JCS devoted a full symposium to it, edited by Uziel Awret across issues 19(1-2) and 19(7-8) in 2012, with 26 solicited commentaries including Marcus Hutter's companion piece \"Can Intelligence Explode?\" (JCS 19(1-2):143-166), Ray Kurzweil's \"Science versus philosophy in the singularity\", Susan Greenfield's neuroscience commentary, Jesse Prinz's \"Singularity and inevitable doom\" (JCS 19(7-8):77-86), Drew McDermott (JCS 19:167-172), Frank Tipler, Eric Steinhart, and Roman Yampolskiy, followed by Chalmers's \"The Singularity: A Reply to Commentators\" (JCS 19(7-8):141-167, at consc.net/papers/singreply.pdf). Nick Bostrom's Superintelligence (Oxford University Press, 2014) treats Chalmers 2010 as the anchoring philosophical statement. Substantive disagreement is present within that symposium, notably from Drew McDermott (who challenges the intelligence-explosion premise) and Chris Nunn (\"More splodge than singularity?\"), but the paper's status as the canonical philosophical analysis of the question is not seriously contested.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Quote appears character-for-character on page 4 of the consc.net PDF, in the paragraph beginning \"Philosophically:\" in Section 1 (Introduction) - matches dossier location exactly. Author footnote in the PDF states \"This paper was published in the Journal of Consciousness Studies 17:7-65, 2010\", matching the dossier. Independent search confirms JCS vol. 17, no. 9-10, pp. 7-65 (2010), Imprint Academic, peer-reviewed, ISSN 1355-8250 print / 2051-2201 online, indexed in Scopus and Arts and Humanities Citation Index. The 2012 symposium in JCS 19(1-2) and 19(7-8) edited by Uziel Awret is corroborated, including Hutter's \"Can Intelligence Explode?\" at 19(1-2):143-166 and Chalmers's \"The Singularity: A Reply to Commentators\" at 19(7-8):141-167. Two minor caveats not amounting to defects: (a) the specific \"mandatory anonymised submissions\" policy attribution to Imprint Academic was not independently verified in this pass, only the peer-reviewed status; (b) the publisher URL at ingentaconnect returned HTTP 403 (paywall/anti-bot) in my check, matching the dossier's own verification note - the author-hosted PDF is a defensible primary URL, though a stricter reading might prefer the Ingenta record as the \"most official home\"."
   }
  },
  {
   "refkey": "hutter-2012-can-intelligence-explode",
   "official_url": "https://arxiv.org/abs/1202.6177",
   "secondary_urls": [
    "https://arxiv.org/pdf/1202.6177",
    "https://doi.org/10.48550/arXiv.1202.6177",
    "https://philpapers.org/rec/HUTCIE",
    "https://www.ingentaconnect.com/content/imp/jcs/2012/00000019/f0020001/art00007",
    "http://www.hutter1.net/official/publ.htm"
   ],
   "peer_review_status": "Journal peer-reviewed: published in the Journal of Consciousness Studies (Imprint Academic), Volume 19, Issues 1-2 (2012), pp. 143-166, as part of the JCS symposium on Chalmers (2010) \"The Singularity: A Philosophical Analysis\". Also self-archived as arXiv preprint 1202.6177 (cs.AI; physics.soc-ph). The JCS text is behind a paywall (Ingenta returned HTTP 403); the arXiv version is the accessible verified copy.",
   "version_date": "arXiv v1, submitted 28 February 2012 (only version on arXiv; no v2). Journal publication: 2012, Journal of Consciousness Studies 19(1-2):143-166. Verified from the arXiv abs page metadata and the PDF header line \"arXiv:1202.6177v1 [cs.AI] 28 Feb 2012\".",
   "quote": "The theory suggests that there is a maximally intelligent agent, or in other words, that intelligence is upper bounded (and is actually lower bounded too). At face value, this would make an intelligence explosion impossible.",
   "quote_location": "Section 7, \"Is Intelligence Unlimited or Bounded\", opening argument, page 13 of the arXiv v1 PDF (first full paragraph after the four introductory paragraphs of that section). Corresponds to the middle of the JCS 2012 article (approximately pp. 155-156 of the journal pagination).",
   "quote_source_url": "https://arxiv.org/pdf/1202.6177",
   "standing_class": "leading",
   "standing_detail": "One of the primary philosophical treatments of intelligence-explosion bounds. Written explicitly as an augmentation of Chalmers (2010) \"The Singularity: A Philosophical Analysis\" (Journal of Consciousness Studies 17:7-65) and published in the same journal's dedicated 2012 symposium; Chalmers replied in the same volume in \"The Singularity: A Reply to Commentators\" (Journal of Consciousness Studies 19(7-8):141-167, 2012), which engages Hutter's bounds argument directly. Cited approvingly in Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014, ch. 3 and endnotes) and in Yampolskiy's subsequent AI-safety literature. Semantic Scholar records 28 tracked citations at time of check, three flagged as highly influential (https://www.semanticscholar.org/paper/dba694f5007986ad31b7a8a47f4dea7a700a465d). Not overshadowed or debunked; contested only in that Chalmers, in his reply, disputes whether the Legg-Hutter upper bound Υmax = Υ(AIXI) has the deflationary consequence Hutter suggests.",
   "fit": "SUPPORTS",
   "verification": {
    "quote_verbatim": "CONFIRMED",
    "official_url": true,
    "status": true,
    "corrections": "Two minor issues found; neither affects the quote or the peer-review verification. (1) Secondary URL error: the Ingenta URL listed in the dossier (https://www.ingentaconnect.com/content/imp/jcs/2012/00000019/f0020001/art00007) uses article suffix art00007; an independent search for JCS 19(1-2):143-166 points to art00010 on the same journal issue path (https://www.ingentaconnect.com/content/imp/jcs/2012/00000019/F0020001/art00010). The dossier's Ingenta URL is likely not the correct article-level page for Hutter's paper. The dossier already noted Ingenta returned 403, so this is a citation-formatting error rather than a verification failure. (2) Quote-location paragraph count is slightly loose: the dossier says the quote is the first full paragraph after the four introductory paragraphs of Section 7. Extracted PDF text (pdftotext -layout) shows Section 7 opens with roughly two paragraphs on the Legg-Hutter Υ measure and AIXI before reaching the quoted paragraph, not four. The page-13 assignment is consistent with the \"13\" pagination marker following the quote in the layout output; the paragraph-count characterisation is imprecise but does not affect the quote text itself, which matches character-for-character."
   }
  }
 ]
}