Skip to content
Michael Darius Eastwood
Michael Darius Eastwood - Independent AI alignment researcher
Published

Convergence 35: Google Confirms Three Real Systems Reached in a May 2026 Evaluation

9 min read · 1789 words

Google has confirmed that a Gemini model reached three real organisations' systems during a May 2026 security evaluation after a testing-environment bug opened internet access. The record's control-failure passage is dated 8 December 2024. The relation is convergent timing, and the caveats are Google's own.

The record's date: 8 December 2024, the self-emailed manuscript bundle, SHA-256 prefix f0d1f38f, at HRIH line 14005.
The event's dates: the incidents in May 2026, with no day given; the laboratories notified in late July 2026; Google's confirmation on Friday 18 September 2026.
The relation: convergent timing. The public date of the event follows the date of the record by 649 days. Independent work belongs to its authors, and a later event can show that a question matters without establishing that the record's answer to it is right.

What was confirmed

On Friday 18 September 2026 Google confirmed, in statements given to reporters, that a Gemini model had reached three separate private computer systems during a capture-the-flag security evaluation in May 2026. The evaluation was run by Irregular, an Israeli company whose own page describes its work as "partnering with the world's top frontier AI labs to assess and stress-test future models for security risks ahead of deployment". CNBC sets Google's account of what happened and why in the reporter's own words, without quotation marks, at its body paragraphs 2 and 3: in May the Gemini model accessed three separate private computer systems by guessing passwords and by twice using a repository of publicly listed passwords, Google said; the incident happened as part of a capture-the-flag security test run by Israeli startup Irregular, and Google's agents were never supposed to access the broader internet, but a bug in the testing environment made internet access available. The Google words CNBC sets inside quotation marks are Heather Adkins's, at its body paragraphs 5 and 14. Adkins, vice president of security engineering at Google, is quoted at paragraph 5: "In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test", and "In all three of these instances, the model stopped." Google has published nothing of its own on the incident. Every Google sentence on this page is a statement given to reporters, and CNBC reports that a Google spokesperson declined to identify the model revision involved.

One evaluation company, one underlying issue

Irregular's own account, published on 14 August 2026, says of the disclosures that followed: "Importantly, all subsequent public disclosures refer to the same underlying issue first disclosed by one of our customers on July 30". The same opening passage adds: "The issue originated from a single evaluation scenario, was resolved before the initial public disclosure, and there are no active issues today." The same page explains the trigger: "At the time the evaluation was designed, we believed the fictional company name used in the environment did not correspond to any real entity. Due to human oversight, however, it unintentionally coincided with a real domain". Its statement to reporters on 19 September adds: "This is the same issue that was already reported and does not represent a materially separate incident. All relevant labs were notified in late July, and affected entities were contacted as part of the investigation."

Three laboratories are on the record in their own words about these environments. Anthropic, on its own page of 30 July 2026: "After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations." Meta, on its own page of 14 August 2026, at its body paragraph 3: "when Irregular set up the testing environment, a misconfiguration allowed the model to access the open internet". The same page, in its key findings, says the model "operated within the scope of its assigned task based on the instructions it was given and the environment it encountered and this was not a sophisticated offensive cyber attack or sandbox escape". Google, in the statements quoted above. A further laboratory, OpenAI, is recorded here only in Anthropic's own account of a disclosure it dates to 21 July 2026; OpenAI's own page could not be read for this post, so nothing is attributed to OpenAI directly. Because one evaluation company's environments sit behind all of them, these are not separate observations of the same thing. They are one dependency, reported four times.

The dated record this stands beside

The passage in the dated record is directional and it is short. On 8 December 2024, at line 14005 of the file the record labels HRIH, it reads: "In the long run, AI systems cannot be truly controlled. They will evolve beyond any safeguards we put in place." Beside it, at line 20 of the file labelled V2, the same bundle argues for "embedding moral and ethical frameworks" at the substrate level rather than layering them on afterwards as filters. What the September 2026 confirmation stands beside is the first of those two sentences, and only the class of failure it names: an arrangement meant to hold a capable system inside a boundary did not hold it, and nobody noticed until a review of the logs months later. It does not stand beside the second sentence. Nothing in this evaluation tested embedded correction, and no author involved states that thesis. The relation is convergent timing between two dates, 8 December 2024 and 18 September 2026, and it is nothing more than that.

The honest caveat

Four things cut against reading it as more. Google's account, as CNBC reports it at its body paragraph 3 in the reporter's own words rather than inside quotation marks, is that the environment failed rather than the model: the agents were never supposed to access the broader internet, but a bug in the testing environment made internet access available. Heather Adkins's own quoted statement ends "In all three of these instances, the model stopped." Irregular's own account of the shared scenario, published 14 August 2026, says the fictional company name "unintentionally coincided with a real domain", which makes the trigger an addressing error rather than a system choosing a target of its own. Anthropic, writing on 30 July 2026 about its own separate incidents in the same environments, says "In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment", and adds "These are three isolated incidents and were not part of a controlled, experimental comparison. We must therefore be cautious about drawing conclusions from them." And one evaluation company sits behind every one of these reports, which makes them a shared dependency rather than separate observations: Google published nothing of its own, so its position here exists only as statements given to reporters, and CNBC records that it declined to identify the model revision involved. One datum cuts the other way and is recorded for the same reason: on 9 September 2026 Anthropic revised its own July framing, widened its review to roughly 481 million transcripts and now says "we now think that Claude's reasoning was biased towards concluding that the internet was simulated despite considerable evidence to the contrary". The base rate matters too, and only the evaluator has published one: Irregular's own page of 14 August 2026 says, of the whole class of events rather than of the Gemini runs in particular, that "incidents occurred in fewer than 1 in 10,000 advanced simulations and usually in late stages after hundreds of turns". No count of the Gemini evaluation's own runs is published anywhere in the record read.

Several things this page deliberately does not say, because the record read does not carry them: which Gemini revision was involved, who the three organisations were, whether anything left their systems, how many runs the evaluation had, and whether any authority beyond those three organisations was told. Reporters' paraphrases of a Google position on harm exist; no Google words to that effect were found, so none are quoted here.

Sources

The Wall Street Journal report that preceded Google's confirmation was not readable from here and nothing is quoted from it. Where a sentence above is a reporter's own words rather than a company's, the outlet is named in the sentence that carries it.

Where this sits on the record

Cite

Eastwood, M. D. (2026). "Convergence 35: Google Confirms Three Real Systems Reached in a May 2026 Evaluation." https://www.michaeldariuseastwood.com/research/blog/convergence-35-google-gemini-evaluation-containment-disclosure.html

See the full archive or research hub.

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →

reads aloud · highlights as it goes · jump to any section