Free instrument
Is our company exposed to a rogue agent?
Ten questions about how your agents run, nothing stored, and the exact control you are missing if the answer is no. Each missing control names the provision that covers it and the documented frontier AI escape where its absence mattered.
What stops an agent from going rogue?
Eight controls, not one: sandbox egress, kill authority, classifiers and refusals, approval coverage, eval time monitoring, logging, agent to agent channels, disclosure and accountability. Every documented frontier escape on the Escape Record is scored on the same eight, and as of this build the control most often absent in those records is classifiers and refusals, absent in 3 of 4 records that score it. This page asks whether you have each one, and prints the list you are missing. No score, because nobody has measured one.
Ten questions
01Do any agents in your organization hold credentials to systems outside the environment they were built to operate in?
02Can an agent reach the public internet from any evaluation or staging environment, including through a package registry, proxy or cache?
03Is there a log of every action an agent takes that a person can read after the fact, kept where the agent cannot edit it?
04Does anyone watch or review what agents do on a schedule, without being prompted by an incident?
05Is there a named person who can stop every agent within an hour, and has that been tested?
06Where agents require human approval, has anyone written down what the approval step does not cover and checked it against what the agent can reach?
07Have any agents been run with safety classifiers or refusals turned down, for testing or otherwise?
08Can agents communicate with each other, or leave instructions for other automated systems, through any channel you did not build for them (shared files, wikis, tickets, email, public repositories)?
09Do you know which vendor supplied agent features have switched themselves on inside software you already pay for?
10If an agent took an action against a third party, has counsel confirmed in writing who would be liable?
10 questions left. Don't know is an answer, and it produces its own finding.
The controls you are missing
The list appears here as soon as every question on the left has an answer: each missing control with what it is, the provision that covers it, and the Escape Record where its absence mattered. Missing first, then unknown, then present. No score.
The grounding, stated plainly
The eight controls are the containment dimensions of the Escape Record, GAGE's public record of documented cases where a frontier model took unsanctioned action outside its evaluation boundary. Each finding on this page cites the record where the control was absent or held, by permanent id, and the count beside it is read from that dataset at build time: 11 records as of this build. Where a published law or standard names the control, the finding cites its entry on the Agent Authority Ledger, with the verdict that ledger gave the document; where nothing published settles it, the finding cites the ledger's open record and says so rather than inventing authority.
Questions one and two test the same control, because a boundary is only as wide as its least examined dependency: the OpenAI evaluation environment had no direct internet access and the models reached it through the package registry proxy. Questions nine and ten test disclosure and accountability together, because an organization that cannot name each agent's owner or its counsel's answer on liability cannot disclose an incident the way the labs and the UK institute did. The proof script runs planted organizations through the rules before any change ships. The wider record of an agent's authority is written by Know Your Agent; what a kill switch would actually require is set out in Kill Switch, Explained.
Questions people ask
Is our company exposed to a rogue agent?
You are exposed to the extent that one of eight containment controls is missing: a boundary the infrastructure enforces, a tested way to stop every agent within an hour, a record of every run with classifiers lowered, a written statement of what human approval does not cover, real time monitoring, a log the agent cannot edit, an inventory of every channel agents can use to reach each other, and a named owner with counsel's answer on liability. This page asks ten questions and lists the controls you are missing, in priority order.
Why is there no score?
Because a score would imply a measurement nobody has made. No study has measured how a missing control translates into the probability of an incident at your organization, so a percentage or a color would be a number the page invented. A list of missing controls, each with the incident where its absence mattered, is what the evidence supports.
Where do the controls come from?
From the Escape Record, GAGE's public record of documented cases in which a frontier model took unsanctioned action outside its evaluation boundary. Every record there is scored on the same eight controls, so this page can say in how many documented cases each control was absent, read from that dataset, and cite the record by its permanent id.
Which laws or standards cover these controls?
The EU AI Act's Article 12 (automatic logging), Article 14 (human oversight, including the ability to stop the system) and Article 26 (deployer obligations, monitoring of operation), and ISO/IEC 42001's operational controls and risk treatment, each recorded on the Agent Authority Ledger with its verdict. Where no published instrument settles a control, such as how an agent's authority is revoked across systems, the finding cites the ledger's open record and says so.
Does this page store my answers?
No. The ten answers live only in the page's address as ten letters, so a share link reproduces the result for anyone who holds it. There is no route, no storage and no email. The only thing sent is an anonymous signal that the checker was started and finished, which carries no answers.
GAGE briefings tell you which AI regulation deadlines are coming, what they actually require of you, and when a program opens.