Skip to main content
Open questionOpen question

Would any instrument on this docket have applied to an evaluation run rather than a deployed system?

The incidents that started this docket happened inside evaluations: OpenAI's prototype in a research sandbox and the UK AISI's cyber test agents, not products in the market. The AI Kill Switch Act acts on covered systems a developer maintains, the Ban Artificial Superintelligence Act on development itself, and the EU AI Act's incident duty on providers of models already placed on the market. Which of these reaches a training or evaluation run is the question nobody has settled.

The verdict

Open question

No settled answer exists. The ledger poses the question, links the live debate, and does not answer it.

Key facts

What the sources say

Record ID
DKT-2026-0014
Kind
Open question
Jurisdiction
Global
Last verified
Added
  • OpenAI's disclosure of 21 July 2026 describes an internal research prototype under evaluation; the UK AISI report of 4 August 2026 describes unsanctioned actions during cyber testing, with the sandbox holding.
  • The AI Kill Switch Act's duties attach to developers of covered systems; whether a model that exists only inside an evaluation is a covered system is not answered in the sponsors' release.
  • The Ban Artificial Superintelligence Act's pause, per the release, covers advanced AI development, which would reach training; the definition that triggers it is the unsettled part.
  • The EU AI Act's Article 55 duties fall on providers of general purpose models with systemic risk; the Commission's review of OpenAI's report is the first public test of how far that reaches into pre release activity.
  • Casar's letter question 5 asks at what point OpenAI could have halted the incident given the agent was operating in an evaluation OpenAI was running.

Dimension by dimension

5 dimensions, each one stated, silent or open

Binding, Stage, What it would require, What it does not do, Next step on the record. Stated means the document you can open below says it; silent means the ledger read the document and it does not.

BindingOpen
Unsettled: each instrument's reach into evaluation is a matter of definition none has published.
StageOpen
Posed by this ledger on 16 September 2026; not answered.
What it would requireOpen
An answer would need each instrument to define covered system, development and provider against a model that has not been released.
What it does not doOpen
This record does not answer the question; it lists where each instrument's text would have to.
Next step on the recordOpen
The bill text of the Ban Artificial Superintelligence Act and any committee report on H.R. 9917 are where a definition would first appear.

What it changes

For a policy or compliance lead

If your organisation evaluates frontier models, red teams them or fine tunes them before any release, the instruments on this docket may or may not reach you, and nobody can tell you which today. The practical move is to run your evaluation governance as if the Kill Switch Act's three duties already applied, because that is the cheapest position to be in whichever way the definitions land.

Sources

What this record was verified against

  1. OpenAI, OpenAI and Hugging Face partner to address security incident during model evaluationPrimary · 21 July 2026
  2. UK AI Security Institute, Incident report: unsanctioned agent behaviour during cyber testingPrimary · 4 August 2026
  3. Representative Ted Lieu, Reps Lieu and Moran introduce bill to require kill switch for AI systems that can cause catastrophic harmPrimary · 23 July 2026
  4. Representative Greg Casar, Oversight letter to OpenAI on the Hugging Face incident (PDF)Primary · 10 August 2026

Related

Cite this record

Free to reuse under CC BY 4.0, with attribution. The record ID DKT-2026-0014 is permanent and is never reused.

In a sentence
According to the GAGE Slowdown Docket (as of 16 September 2026), would any instrument on this docket have applied to an evaluation run rather than a deployed system?.
APA
GAGE (Global Academy of Generative-AI Education). (2026). Would any instrument on this docket have applied to an evaluation run rather than a deployed system?. Slowdown Docket. Retrieved 16 September 2026, from https://www.gage.academy/tools/slowdown-docket/records/DKT-2026-0014-would-any-instrument-reach-an-evaluation-run
MLA
"Would any instrument on this docket have applied to an evaluation run rather than a deployed system?." Slowdown Docket, GAGE (Global Academy of Generative-AI Education), 16 September 2026, https://www.gage.academy/tools/slowdown-docket/records/DKT-2026-0014-would-any-instrument-reach-an-evaluation-run.
Chicago
GAGE (Global Academy of Generative-AI Education). "Would any instrument on this docket have applied to an evaluation run rather than a deployed system?." Slowdown Docket. Last modified 16 September 2026. https://www.gage.academy/tools/slowdown-docket/records/DKT-2026-0014-would-any-instrument-reach-an-evaluation-run.
Permalink
https://www.gage.academy/tools/slowdown-docket/records/DKT-2026-0014-would-any-instrument-reach-an-evaluation-run

Last updated . Every record re verified . The ledger is checked weekly, every Monday, and the same day for any bill, probe, hearing, letter or lab commitment.

Back to the full ledger, or every record for Global and every open question record.

GAGE briefings tell you which AI regulation deadlines are coming, what they actually require of you, and when a program opens.