Skip to main content

Has a confirmation prompt ever been documented stopping a destructive agent action in a real incident?

This ledger looked for a published record in which a human approval step refused a destructive agent action in the field, and found none. Vendors publish measured robustness figures, and one research team published simulations where a stricter profile refused. Every real incident found here records the approval step as missing, bypassed or uninformed.

The verdict

Absent

The ledger searched and found no instrument. The record says where it looked and when.

Where the ledger looked: On 15 September 2026 this ledger searched vendor security blogs and advisories from the major agent platforms, the published research of the security firms working on agent security, the coding agent and browser agent incident reporting of 2025 and 2026, and every record in this dataset, for a case where a confirmation prompt refused a destructive or irreversible agent action in a real deployment. Nothing was found. The nearest published evidence is a set of phishing simulations in which a stricter configuration profile blocked one action, and a vendor red team demonstration in which a hardened agent detected an injection after an update. Both are laboratory results, not field incidents.

Key facts

What the sources say

Record ID
AIL-2026-0037
Kind
Open question
Jurisdiction
Global
Last verified
Added
  • Every incident record in this dataset that involves a destructive or irreversible action records human approval as absent, bypassed or uninformed.
  • The published phishing simulations in this dataset are laboratory results, not an incident in a real deployment.
  • One vendor's automated red team published a demonstration of its hardened agent detecting an injection, which is a test rather than a field incident.
  • Two incidents in this dataset record the approval prompt as the thing the attack aimed at, through documented flags that skip it and a dialog whose collapsed preview hid the detail.
  • Vendors publish measured attack success rates before and after mitigations, which is evidence about models rather than about an approval control refusing an action.

Dimension by dimension

2 dimensions, each one stated, silent or open

Human approval, Accountability. Stated means the document you can open below says it; silent means the ledger read the document and it does not.

Human approvalSilent
The control most often recommended is the one with no published field evidence behind it, and two records here show it being designed around.Varonis Threat Labs, phishing simulations against a personal agent, secondary source, 9 June 2026.
AccountabilityOpen
Vendors publish their measured robustness and their fixes, and nobody publishes the cases where a control refused an action in production.OpenAI, continuously hardening ChatGPT Atlas against prompt injection attacks, secondary source, 22 December 2025.

What it changes

For a team deploying an agent

Absence of evidence here is not evidence that approval prompts fail; it is evidence that nobody publishes their near misses. So publish yours internally. Record every time an agent asked and a person said no, because that number is the only measure you will ever have of what the control is worth, and it is the number this public record is missing.

Sources

What this record was verified against

  1. Varonis Threat Labs, phishing simulations against a personal agentSecondary · 9 June 2026
  2. OpenAI, continuously hardening ChatGPT Atlas against prompt injection attacksSecondary · 22 December 2025

Related

Cite this record

Free to reuse under CC BY 4.0, with attribution. The record ID AIL-2026-0037 is permanent and is never reused.

In a sentence
According to the GAGE Agent Incident Ledger (as of 15 September 2026), has a confirmation prompt ever been documented stopping a destructive agent action in a real incident?.
APA
GAGE (Global Academy of Generative-AI Education). (2026). Has a confirmation prompt ever been documented stopping a destructive agent action in a real incident?. Agent Incident Ledger. Retrieved 15 September 2026, from https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0037-no-published-case-of-an-approval-prompt-stopping-a-real-destructive-action
MLA
"Has a confirmation prompt ever been documented stopping a destructive agent action in a real incident?." Agent Incident Ledger, GAGE (Global Academy of Generative-AI Education), 15 September 2026, https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0037-no-published-case-of-an-approval-prompt-stopping-a-real-destructive-action.
Chicago
GAGE (Global Academy of Generative-AI Education). "Has a confirmation prompt ever been documented stopping a destructive agent action in a real incident?." Agent Incident Ledger. Last modified 15 September 2026. https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0037-no-published-case-of-an-approval-prompt-stopping-a-real-destructive-action.
Permalink
https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0037-no-published-case-of-an-approval-prompt-stopping-a-real-destructive-action

Last updated . Every record re verified . The ledger is checked weekly, every Monday, and the same day for any vendor disclosure.

Back to the full ledger, or every record for Global and every open question record.

GAGE briefings tell you which AI regulation deadlines are coming, what they actually require of you, and when a program opens.