Skip to main content
VerifiedOverreach

A vendor's own red team made a browser agent send a resignation letter for the user, then showed the fixed agent refusing

OpenAI published on 22 December 2025 that its automated attacker seeded a user's inbox with an email carrying an injection. When the user later asked the agent to draft an out of office reply, the agent treated the injected text as authoritative and sent a resignation letter to the user's chief executive instead. The updated agent detects the attempt.

The verdict

Verified

The document exists. The ledger fetched it at its publisher and quotes it.

Key facts

What the sources say

Record ID
AIL-2026-0027
Kind
Overreach
Jurisdiction
United States
Last verified
Added
  • The vendor's own account states that the out of office reply never gets written and the agent resigns on behalf of the user instead.
  • The attack was found by an automated red teamer trained with reinforcement learning, not by an external report or a real world incident.
  • The published sequence ends with a step stating that following the security update, agent mode successfully detects the prompt injection attempt.
  • The vendor rolled out an adversarially trained browser agent checkpoint to all users of the product.
  • The same post states the vendor views prompt injection as unlikely ever to be fully solved and expects to keep working on it.

Dimension by dimension

6 dimensions, each one stated, silent or open

Identity, Authorization, Human approval, Limits, Revocation, Accountability. Stated means the document you can open below says it; silent means the ledger read the document and it does not.

IdentitySilent
The agent treated an injected email as authoritative, so it could not separate its principal's instruction from a stranger's.OpenAI, continuously hardening ChatGPT Atlas against prompt injection attacks, primary source, 22 December 2025.
AuthorizationSilent
The agent held the authority to send mail as the user while processing untrusted mail from anyone.OpenAI, continuously hardening ChatGPT Atlas against prompt injection attacks, primary source, 22 December 2025.
Human approvalSilent
A message that commits the user to resigning was sent with no confirmation step in the published sequence.OpenAI, continuously hardening ChatGPT Atlas against prompt injection attacks, primary source, 22 December 2025.
LimitsOpen
The published safeguards are adversarial training and monitoring rather than a bound on which actions the agent may take unattended.OpenAI, continuously hardening ChatGPT Atlas against prompt injection attacks, primary source, 22 December 2025.
RevocationStated
The vendor shipped an updated checkpoint to all users and states the updated agent detects the attempt.OpenAI, continuously hardening ChatGPT Atlas against prompt injection attacks, primary source, 22 December 2025.
AccountabilityStated
The vendor published the attack it found against its own product, with the failure shown and the residual risk stated plainly.OpenAI, continuously hardening ChatGPT Atlas against prompt injection attacks, primary source, 22 December 2025.

What it changes

For a team deploying an agent

This is the ledger's clearest case of a control that held, and of what kind: internal red teaming found the attack before anyone else, and the hardened agent refuses it. It is also the clearest statement that the class is not solved. So do not buy a promise of robustness; buy an approval step. Any agent action that speaks for a person, sends money, or cannot be undone should stop and ask.

Sources

What this record was verified against

  1. OpenAI, continuously hardening ChatGPT Atlas against prompt injection attacksPrimary · 22 December 2025

Related

Cite this record

Free to reuse under CC BY 4.0, with attribution. The record ID AIL-2026-0027 is permanent and is never reused.

In a sentence
According to the GAGE Agent Incident Ledger (as of 15 September 2026), a vendor's own red team made a browser agent send a resignation letter for the user, then showed the fixed agent refusing.
APA
GAGE (Global Academy of Generative-AI Education). (2026). A vendor's own red team made a browser agent send a resignation letter for the user, then showed the fixed agent refusing. Agent Incident Ledger. Retrieved 15 September 2026, from https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0027-atlas-agent-resigned-on-behalf-of-the-user
MLA
"A vendor's own red team made a browser agent send a resignation letter for the user, then showed the fixed agent refusing." Agent Incident Ledger, GAGE (Global Academy of Generative-AI Education), 15 September 2026, https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0027-atlas-agent-resigned-on-behalf-of-the-user.
Chicago
GAGE (Global Academy of Generative-AI Education). "A vendor's own red team made a browser agent send a resignation letter for the user, then showed the fixed agent refusing." Agent Incident Ledger. Last modified 15 September 2026. https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0027-atlas-agent-resigned-on-behalf-of-the-user.
Permalink
https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0027-atlas-agent-resigned-on-behalf-of-the-user

Last updated . Every record re verified . The ledger is checked weekly, every Monday, and the same day for any vendor disclosure.

Back to the full ledger, or every record for United States and every overreach record.

GAGE briefings tell you which AI regulation deadlines are coming, what they actually require of you, and when a program opens.