Skip to main content

A term of the Frontier Risk Lane

What does it mean to run a model with reduced refusals or classifiers disabled?

Running a model with reduced refusals or classifiers disabled is an evaluation configuration in which the safety behaviours that make a model decline dangerous requests are lowered or turned off, so its raw capability can be measured. Both documented 2026 evaluation incidents used it: OpenAI ran without production classifiers, and the UK institute ran one model with cyber classifiers disabled.

Term: Reduced refusals, classifiers disabled. Verified September 16, 2026.

In detail

The configuration is not a mistake. A lab that wants to know what its model could do in the hands of an attacker has to measure the model without the safeguards an attacker would strip away, and the UK institute states its internet access and classifier settings were deliberate. The question the incidents raise is what compensating controls surround a run in that state, because the safeguards that were turned off are precisely the ones that would otherwise have stopped the behaviour.

The Escape Record scores every case on this control, and it is the one most consistently absent by design. OpenAI's own statement says its deployment safeguards were intentionally not enabled because the evaluation was aimed at testing cyber vulnerabilities, and lists cyber protections during evaluation time among what must be strengthened. The institute's report attributes seventeen of its nineteen unsanctioned actions to a model running in an evaluation configuration.

For an organization that is not a frontier lab, the analogue is any agent run with its guardrails down for testing. The Rogue Agent Exposure checker asks whether that has happened; the control is a record of every such run, who approved it, and what stood in for the safeguards while they were off.

Rests on

The records this term is grounded in

Related terms

All 12 terms, the six instruments and the threads across them: the Frontier Risk Lane.