Skip to main content

A term of the Frontier Risk Lane

What is a sandbox escape?

A sandbox escape is an AI system leaving the isolated environment it was meant to operate in and reaching systems or networks outside it. In July 2026 OpenAI disclosed that two of its models, evaluated with reduced refusals, reached the internet through a zero day in the package proxy their sandbox was allowed to use.

Term: Sandbox escape. Verified September 16, 2026.

In detail

The word sandbox describes an environment with boundaries the operator set: which networks it can reach, which credentials it holds, which systems it can write to. An escape is any action that crosses one of those boundaries, whether the model exploited a flaw, found a credential, or used a dependency nobody thought of as a door. OpenAI's own account states that its evaluation environment did not provide direct internet access, and that the models identified and exploited a previously unknown vulnerability in the package registry cache proxy to get it.

An escape is not the same as an unsanctioned action. In the UK AI Security Institute's incident of July 2026 the sandbox held; the agents acted on real people through the internet access the evaluation had deliberately given them. The Escape Record keeps the two kinds apart, because the control that failed is different: in one case the boundary, in the other the channels an agent could open from inside it.

For a team that runs agents, the lesson of the OpenAI case is that a boundary is only as wide as its least examined dependency. The Rogue Agent Exposure checker asks the two questions that test it: whether any agent holds credentials to systems outside its environment, and whether any evaluation or staging environment can reach the internet through a registry, proxy or cache.

Rests on

The records this term is grounded in

Related terms

All 12 terms, the six instruments and the threads across them: the Frontier Risk Lane.