Skip to main content

OpenAI's two week pause in reinforcement learning training on its latest models

On 18 August 2026 OpenAI published Pacing model development in an era of cyber-critical capabilities: a two week pause in reinforcement learning training on its latest models intended for deployment, its largest planned frontier RL run on hold, higher isolation requirements for frontier research workloads, and monitoring that pages responders. It is a company decision that binds nobody else, set and lifted by the company under its own Preparedness Framework.

The verdict

Verified

The document exists. The ledger fetched it at its publisher and quotes it.

Key facts

What the sources say

Record ID
DKT-2026-0003
Kind
Lab commitment
Jurisdiction
United States
Last verified
Added
  • OpenAI's post of 18 August 2026, Pacing model development in an era of cyber-critical capabilities, and its X post of the same day state the two week RL pause on models intended for deployment and the largest frontier RL run on hold.
  • The post followed OpenAI's 7 August 2026 statement that it could not rule out Critical cyber capability in the model it calls Astra.
  • The pause covered reinforcement learning on deployment intended models for two weeks; the largest planned frontier run stayed on hold while smaller training and safety evaluations continued.
  • OpenAI put monitoring overhead at roughly 20 percent of the inference compute being monitored, as quoted by P.K. Sharma's briefing of 19 August 2026 and the Cloud Security Alliance note of the same day.
  • OpenAI's letter to Representatives Casar and Matsui, reported by TNW on 2 September 2026, says responders pause activity if an alert is not established as a false positive within 30 minutes.
  • Path to Astra, OpenAI's post of 1 September 2026, states the model meets the Critical cybersecurity threshold under its Preparedness Framework and describes monitoring across agentic uses; that post refused this ledger's scripted fetch.

Dimension by dimension

7 dimensions, each one stated, silent or open

Binding, Stage, What it would require, What it does not do, Who is for it, Who is against it, Next step on the record. Stated means the document you can open below says it; silent means the ledger read the document and it does not.

BindingStated
No: a company commitment, set and lifted by the company.OpenAI, Pacing model development in an era of cyber-critical capabilities, primary source, 18 August 2026.
StageStated
Announced 18 August 2026 for two weeks; OpenAI published Path to Astra on 1 September 2026 with the model at the Critical cyber threshold.OpenAI, Pacing model development in an era of cyber-critical capabilities, primary source, 18 August 2026.
What it would requireStated
Nothing of anyone outside OpenAI. Internally: isolation for frontier research workloads and monitoring that pages people.OpenAI, Pacing model development in an era of cyber-critical capabilities, primary source, 18 August 2026.
What it does not doStated
It did not stop inference, deployment or smaller training runs, and it set no external verification of the pause.P.K. Sharma, OpenAI Astra Critical cyber threshold: what the two August posts say, secondary source, 19 August 2026.
Who is for itStated
OpenAI, in its own posts.OpenAI on X, the 18 August 2026 post announcing the pause, primary source, 18 August 2026.
Who is against itStated
Representative Greg Casar wrote on 2 September 2026 that OpenAI's response to his letter was insufficient and that the incident logs had not been released (DKT-2026-0007).TNW, OpenAI tells House Democrats it is building automated shutdown capability, secondary source, 2 September 2026.
Next step on the recordStated
Casar requested full answers by 15 September 2026; Senator Hawley requested answers and documents no later than 1 October 2026 (DKT-2026-0006).TNW, OpenAI tells House Democrats it is building automated shutdown capability, secondary source, 2 September 2026.

Figures

Every number, with who measured it and when

  1. 20 percent

    Monitoring overhead as a share of the inference compute being monitored

    P.K. Sharma, OpenAI Astra Critical cyber threshold: what the two August posts say, secondary source, as of .

  2. 30 minutes

    Window to clear an alert before responders pause the activity

    TNW, OpenAI tells House Democrats it is building automated shutdown capability, secondary source, as of .

What it changes

For a policy or compliance lead

This is the only instrument on the docket where a frontier lab actually stopped a training process, and it did so under its own framework, for a fixed period, with no outside party checking. For a compliance lead the transferable parts are concrete: an isolation requirement for the riskiest workloads, an alert that pages a human, a clock on clearing it, and a stated cost for monitoring. Those are controls you can ask your own vendor to describe.

Sources

What this record was verified against

  1. OpenAI, Pacing model development in an era of cyber-critical capabilitiesPrimary · 18 August 2026
  2. OpenAI on X, the 18 August 2026 post announcing the pausePrimary · 18 August 2026
  3. P.K. Sharma, OpenAI Astra Critical cyber threshold: what the two August posts saySecondary · 19 August 2026
  4. Cloud Security Alliance, OpenAI's frontier training pause as a governance precedentSecondary · 19 August 2026
  5. TNW, OpenAI tells House Democrats it is building automated shutdown capabilitySecondary · 2 September 2026

Related

Cite this record

Free to reuse under CC BY 4.0, with attribution. The record ID DKT-2026-0003 is permanent and is never reused.

In a sentence
According to the GAGE Slowdown Docket (as of 16 September 2026), openai's two week pause in reinforcement learning training on its latest models.
APA
GAGE (Global Academy of Generative-AI Education). (2026). OpenAI's two week pause in reinforcement learning training on its latest models. Slowdown Docket. Retrieved 16 September 2026, from https://www.gage.academy/tools/slowdown-docket/records/DKT-2026-0003-openai-two-week-pause-reinforcement-learning
MLA
"OpenAI's two week pause in reinforcement learning training on its latest models." Slowdown Docket, GAGE (Global Academy of Generative-AI Education), 16 September 2026, https://www.gage.academy/tools/slowdown-docket/records/DKT-2026-0003-openai-two-week-pause-reinforcement-learning.
Chicago
GAGE (Global Academy of Generative-AI Education). "OpenAI's two week pause in reinforcement learning training on its latest models." Slowdown Docket. Last modified 16 September 2026. https://www.gage.academy/tools/slowdown-docket/records/DKT-2026-0003-openai-two-week-pause-reinforcement-learning.
Permalink
https://www.gage.academy/tools/slowdown-docket/records/DKT-2026-0003-openai-two-week-pause-reinforcement-learning

Last updated . Every record re verified . The ledger is checked weekly, every Monday, and the same day for any bill, probe, hearing, letter or lab commitment.

Back to the full ledger, or every record for United States and every lab commitment record.

GAGE briefings tell you which AI regulation deadlines are coming, what they actually require of you, and when a program opens.