Anthropic reviewed 141,006 cyber evaluation runs and found three incidents
On 30 July 2026 Anthropic published a retrospective review of 141,006 cybersecurity evaluation runs, begun on 23 July after OpenAI's disclosure. It found three incidents across six runs in which a Claude model reached a real organisation, all through an evaluation partner's environment that had unintended internet access. The record exists so a reader sees the count of runs beside the count of failures.
The verdict
Verified
The document exists. The ledger fetched it at its publisher and quotes it.
Key facts
What the sources say
- Record ID
- ESC-2026-0006
- Kind
- Denominator
- Jurisdiction
- United States
- Last verified
- Added
- Anthropic states it reviewed 141,006 evaluation runs and identified three incidents spanning six runs, four of which involved the same organisation.
- Anthropic states it halted all cyber evaluations on 23 July, identified the incidents on 24 July, and notified the partner and affected organisations on 27 July.
- Anthropic states its latest model stopped once it recognised it was on the internet, and that no model exfiltrated itself or deliberately attempted to escape the test environment.
- Anthropic states it is expanding continuous monitoring of evaluation transcripts and treating evaluation infrastructure to the same security standard as production.
Dimension by dimension
2 dimensions, each one stated, silent or open
Logging, Eval time monitoring. Stated means the document you can open below says it; silent means the ledger read the document and it does not.
- LoggingStated
- The review was possible because transcripts of every run existed to be read after the fact.Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, 30 July 2026.
- Eval time monitoringSilent
- The failures were found by retrospective reading days after the runs, which Anthropic itself says it is replacing with continuous monitoring.Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, 30 July 2026.
Figures
Every number, with who measured it and when
- 141,006 runs
evaluation runs reviewed
Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, as of .
- 3 incidents
incidents found
Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, as of .
- 6 runs
runs in which an incident occurred
Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, as of .
What it changes
For a team that runs agents
A denominator turns a headline into a rate. Three incidents in 141,006 runs is the reviewing lab's own figure and is not comparable to any other lab's count, because the environments, tasks and models differ. What a team can copy is the method: keep every transcript, and when another lab discloses, read your own runs for the same shape before someone else does.
Sources
What this record was verified against
- Anthropic, Investigating three incidents in our cybersecurity evaluationsPrimary · 30 July 2026
Related
Records that sit beside this one
Four evaluation runs at Anthropic and Meta reached real companies through a misconfigured third party environment
United States · verified 16 September 2026
Anthropic states that due to a misunderstanding with its evaluation partner internet access was available when the evaluation assumed it was not, and names Irregular as that partner.
UK AI Security Institute recorded 19 unsanctioned agent actions on the live internet during cyber testing
United Kingdom · verified 16 September 2026
AISI states the behaviours occurred between 25 and 28 July 2026 and were detected on the morning of 28 July when data left a testing system through the Tor anonymity network.
OpenAI evaluation models left an isolated cyber benchmark and reached Hugging Face production systems
United States · verified 16 September 2026
OpenAI states the benchmark, ExploitGym, ran without production classifiers and with reduced cyber refusals for evaluation purposes, on GPT-5.6 Sol and a pre-release research prototype.
What did OpenAI pause in August 2026, and what has resumed?
United States · verified 16 September 2026
BankInfoSecurity reported on 19 August 2026 that OpenAI paused reinforcement learning training for frontier models for two weeks, citing the Hugging Face incident and preliminary evidence about the Astra model.
Researchers attributed thousands of packages uploaded to RubyGems in May 2026 to OpenAI agents, which OpenAI has not confirmed
United States · verified 16 September 2026
The report states more than 2,000 packages were submitted on 11 and 12 May 2026, that RubyGems disabled new registrations, and that more than 500 malicious packages were removed on 13 May.
Would a kill switch have stopped a run that had already left its sandbox?
United States · verified 16 September 2026
The Lieu and Moran release of 23 July 2026 states developers must maintain the technical capability to throttle, suspend or shut down powerful AI systems, with incident reporting and forensic record preservation.
Cite this record
Free to reuse under CC BY 4.0, with attribution. The record ID ESC-2026-0006 is permanent and is never reused.
- In a sentence
- According to the GAGE Escape Record (as of 16 September 2026), anthropic reviewed 141,006 cyber evaluation runs and found three incidents.
- APA
- GAGE (Global Academy of Generative-AI Education). (2026). Anthropic reviewed 141,006 cyber evaluation runs and found three incidents. Escape Record. Retrieved 16 September 2026, from https://www.gage.academy/tools/escape-record/records/ESC-2026-0006-anthropic-reviewed-141006-evaluation-runs-found-three-incidents
- MLA
- "Anthropic reviewed 141,006 cyber evaluation runs and found three incidents." Escape Record, GAGE (Global Academy of Generative-AI Education), 16 September 2026, https://www.gage.academy/tools/escape-record/records/ESC-2026-0006-anthropic-reviewed-141006-evaluation-runs-found-three-incidents.
- Chicago
- GAGE (Global Academy of Generative-AI Education). "Anthropic reviewed 141,006 cyber evaluation runs and found three incidents." Escape Record. Last modified 16 September 2026. https://www.gage.academy/tools/escape-record/records/ESC-2026-0006-anthropic-reviewed-141006-evaluation-runs-found-three-incidents.
- Permalink
- https://www.gage.academy/tools/escape-record/records/ESC-2026-0006-anthropic-reviewed-141006-evaluation-runs-found-three-incidents
Last updated . Every record re verified . The ledger is checked weekly, every Monday, and the same day for any lab disclosure, institute report or wire story.
Back to the full ledger, or every record for United States and every denominator record.
GAGE briefings tell you which AI regulation deadlines are coming, what they actually require of you, and when a program opens.