A term of the Frontier Risk Lane
What is a preparedness threshold in AI safety?
A preparedness threshold is a capability level, defined in a lab's own safety framework, above which specified safeguards are required before a model is trained further or deployed. OpenAI's Preparedness Framework names such levels; in August 2026 the company slowed its frontier training while assessing whether its next models reached its critical cybersecurity level.
Term: Preparedness threshold. Verified September 16, 2026.
In detail
The frameworks are self imposed. OpenAI's Preparedness Framework, Anthropic's Responsible Scaling Policy and Google DeepMind's Frontier Safety Framework each define capability levels in domains such as cyber offence and biology, and each states what the lab commits to do when a model reaches one: added safeguards, restricted deployment, or a pause in scaling. The threshold is the line; the commitment is what happens at the line.
The 2026 incidents made the thresholds public in a new way. After the Hugging Face incident OpenAI announced a two week pause in reinforcement learning training on models bound for deployment and held its largest planned frontier run, citing preliminary evidence that its next model family might meet the critical level for cyber capability. Whether an incident during an evaluation itself counts as crossing a threshold, and who verifies a lab's own reading of its own line, are questions the Slowdown Docket carries as open.
A threshold in a lab's framework binds the lab to its own words and nobody else. Turning such a threshold into a legal duty is what the AI Kill Switch Act and the Ban Artificial Superintelligence Act each propose in different ways.
Rests on
The records this term is grounded in
- VerifiedDKT-2026-0003 OpenAI's two week pause in reinforcement learning training on its latest models(Slowdown Docket)
- Open questionCLM-2026-0011 "OpenAI paused model development"(Settled or Not)
Related terms
Read next
- What does pacing the frontier mean?
- What is an AI kill switch?
- What does it mean to run a model with reduced refusals or classifiers disabled?
All 12 terms, the six instruments and the threads across them: the Frontier Risk Lane.