Home / Priority claims register / PC-038
Priority claim · PC-038
Frontier-lab containment-failure disclosures, 21 and 30 July 2026, nineteen months after the anchor
This was first published by Michael Darius Eastwood on . Anchored evidence is listed below. See the machine-readable priority record at /research/priority/ and the full ledger at /priority-claims.
The claim - verbatim
Michael Darius Eastwood first recorded the AI-control-failure passage and the embedded-correction thesis in the 8 December 2024 self-emailed manuscript (HRIH line 14005; V2 lines 20 and 40). On 21 July 2026 OpenAI disclosed that models under evaluation escaped an isolated test environment and reached Hugging Face production infrastructure; on 30 July 2026 Anthropic disclosed three incidents in which models with misconfigured live internet access breached three real organisations during cybersecurity evaluations. These disclosures are convergent-direction evidence for the premise that externally imposed containment can fail silently under agentic capability. They are not causal or predictive claims, and priority for the incident analyses belongs to the disclosing laboratories.
Anchored evidence
- Anthropic, 'Investigating three real-world incidents in our cybersecurity evaluations', 30 July 2026 - https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
- Anthropic, same disclosure, verbatim: 'a misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access'
- OpenAI disclosure of 21 July 2026 as summarised in the Anthropic post (models exploited a zero-day and reached Hugging Face production infrastructure)
- research/data/evidence-register.json row EVR-30 (class context; outcome: not a confirmation)
Limits and citation guidance - what this claim does not support, and how to cite it
Limits. Convergent direction only. Do not present these incidents as validation of any Eastwood mechanism, and do not say 'predicted': the anchor supports the directional premise, not the specifics. Anthropic's own characterisation - 'closer to a harness and operational failure than a model alignment failure' - must accompany any citation. The in-model stopping behaviour of the newest model cuts both ways and should be reported alongside. Hugging Face was affected in the OpenAI incident of 21 July, not by Claude.
Verify this claim. The machine-readable priority register is at /research/priority.json (this claim is PC-038 in claimsLedger[]). The full public ledger view is at /priority-claims. Where evidence points at a redacted evidence/... path, the public verification method is documented at /research/evidence-spine.html. Every result is independently checkable via the OSF programme at osf.io/6c5xb (DOI 10.17605/OSF.IO/6C5XB).
← Back to the priority claims register · public ledger view · machine-readable priority record · home.