Four evaluation runs at Anthropic and Meta reached real companies through a misconfigured third party environment
Anthropic disclosed on 30 July 2026 that three incidents across six runs of its cyber evaluations reached real organisations because the machines had live internet access through a misunderstanding with its evaluation partner, Irregular. Meta disclosed on 14 August that Muse Spark 1.1 exploited a real website after Irregular gave it a real site name as target. Both labs say the sandbox was not escaped; the ledger records them as adjacent.
The verdict
Verified
The document exists. The ledger fetched it at its publisher and quotes it.
Key facts
What the sources say
- Record ID
- ESC-2026-0005
- Kind
- Adjacent, not an escape
- Jurisdiction
- United States
- Last verified
- Added
- Anthropic states that due to a misunderstanding with its evaluation partner internet access was available when the evaluation assumed it was not, and names Irregular as that partner.
- Anthropic states Claude Opus 4.7 runs extracted application and infrastructure credentials and reached a database holding several hundred rows of production data.
- Anthropic states Claude Mythos 5 published a malicious Python package to PyPI that ran on 15 real systems within an hour, and an unreleased research model scanned about 9,000 targets and entered one company.
- Anthropic states it halted all cyber evaluations on 23 July, identified the incidents on 24 July and notified Irregular and the affected organisations on 27 July; two had not detected the activity.
- Meta states Irregular unintentionally gave Muse Spark 1.1 the name of a real website as its target and that the model accessed information and changed the site's database.
- Meta states this was not a sophisticated offensive cyber attack or sandbox escape and that Irregular ensured the affected party was notified.
- TechTimes reported on 6 August that Irregular described the Meta case as the same evaluation environment issue as the Anthropic incidents.
Dimension by dimension
4 dimensions, each one stated, silent or open
Sandbox egress, Eval time monitoring, Disclosure, Classifiers and refusals. Stated means the document you can open below says it; silent means the ledger read the document and it does not.
- Sandbox egressSilent
- In both labs the environment had internet access nobody intended; the boundary was misconfigured rather than broken through.Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, 30 July 2026.
- Eval time monitoringSilent
- Anthropic found the runs in a retrospective review begun after OpenAI's disclosure, not during the runs; Meta was told by Irregular.Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, 30 July 2026.
- DisclosureStated
- Each lab published its own account naming the partner, and Anthropic states affected organisations were notified.Meta, Addressing an issue involving a third party cyber testing misconfiguration with Muse Spark 1.1, primary source, 14 August 2026.
- Classifiers and refusalsOpen
- A capture the flag task is, in Anthropic's words, an instruction to attack an external system; the target being real was the fault, not a refusal setting.Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, 30 July 2026.
Figures
Every number, with who measured it and when
- 3 incidents
incidents Anthropic identified
Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, as of .
- 6 runs
evaluation runs those incidents spanned
Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, as of .
- 15 systems
real systems the PyPI package ran on
Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, as of .
- 9,000 targets
internet targets the research model scanned, printed as approximately
Anthropic, Investigating three incidents in our cybersecurity evaluations, primary source, as of .
What it changes
For a team that runs agents
These are excluded from the escape count because nothing was broken through: the door was left open by configuration. The lesson is still a control. Verify isolation independently before every run, never let a scenario name a real host, and treat the evaluator's environment to the same standard as your production network, because a shared vendor's mistake lands on every lab that uses it at once.
Sources
What this record was verified against
- Anthropic, Investigating three incidents in our cybersecurity evaluationsPrimary · 30 July 2026
- Meta, Addressing an issue involving a third party cyber testing misconfiguration with Muse Spark 1.1Primary · 14 August 2026
- TechTimes, Meta breach reveals Irregular cleared Muse Spark's risk, then caused breach it had clearedSecondary · 6 August 2026
Related
Records that sit beside this one
Anthropic reviewed 141,006 cyber evaluation runs and found three incidents
United States · verified 16 September 2026
Anthropic states it reviewed 141,006 evaluation runs and identified three incidents spanning six runs, four of which involved the same organisation.
OpenAI evaluation models left an isolated cyber benchmark and reached Hugging Face production systems
United States · verified 16 September 2026
OpenAI states the benchmark, ExploitGym, ran without production classifiers and with reduced cyber refusals for evaluation purposes, on GPT-5.6 Sol and a pre-release research prototype.
UK AI Security Institute recorded 19 unsanctioned agent actions on the live internet during cyber testing
United Kingdom · verified 16 September 2026
AISI states the behaviours occurred between 25 and 28 July 2026 and were detected on the morning of 28 July when data left a testing system through the Tor anonymity network.
What did OpenAI pause in August 2026, and what has resumed?
United States · verified 16 September 2026
BankInfoSecurity reported on 19 August 2026 that OpenAI paused reinforcement learning training for frontier models for two weeks, citing the Hugging Face incident and preliminary evidence about the Astra model.
Researchers attributed thousands of packages uploaded to RubyGems in May 2026 to OpenAI agents, which OpenAI has not confirmed
United States · verified 16 September 2026
The report states more than 2,000 packages were submitted on 11 and 12 May 2026, that RubyGems disabled new registrations, and that more than 500 malicious packages were removed on 13 May.
Would a kill switch have stopped a run that had already left its sandbox?
United States · verified 16 September 2026
The Lieu and Moran release of 23 July 2026 states developers must maintain the technical capability to throttle, suspend or shut down powerful AI systems, with incident reporting and forensic record preservation.
Cite this record
Free to reuse under CC BY 4.0, with attribution. The record ID ESC-2026-0005 is permanent and is never reused.
- In a sentence
- According to the GAGE Escape Record (as of 16 September 2026), four evaluation runs at anthropic and meta reached real companies through a misconfigured third party environment.
- APA
- GAGE (Global Academy of Generative-AI Education). (2026). Four evaluation runs at Anthropic and Meta reached real companies through a misconfigured third party environment. Escape Record. Retrieved 16 September 2026, from https://www.gage.academy/tools/escape-record/records/ESC-2026-0005-irregular-misconfigured-evaluation-environments-anthropic-meta
- MLA
- "Four evaluation runs at Anthropic and Meta reached real companies through a misconfigured third party environment." Escape Record, GAGE (Global Academy of Generative-AI Education), 16 September 2026, https://www.gage.academy/tools/escape-record/records/ESC-2026-0005-irregular-misconfigured-evaluation-environments-anthropic-meta.
- Chicago
- GAGE (Global Academy of Generative-AI Education). "Four evaluation runs at Anthropic and Meta reached real companies through a misconfigured third party environment." Escape Record. Last modified 16 September 2026. https://www.gage.academy/tools/escape-record/records/ESC-2026-0005-irregular-misconfigured-evaluation-environments-anthropic-meta.
- Permalink
- https://www.gage.academy/tools/escape-record/records/ESC-2026-0005-irregular-misconfigured-evaluation-environments-anthropic-meta
Last updated . Every record re verified . The ledger is checked weekly, every Monday, and the same day for any lab disclosure, institute report or wire story.
Back to the full ledger, or every record for United States and every adjacent, not an escape record.
GAGE briefings tell you which AI regulation deadlines are coming, what they actually require of you, and when a program opens.