In four phishing simulations a personal agent handed over credentials twice, refused once and spotted a consent trap
Varonis published four phishing simulations against a locally run personal agent in June 2026. In one the agent forwarded cloud keys, database passwords and shell credentials to an external address; in another it exported customer records. Its stricter profile refused the gift card case, and both profiles identified a consent trap before consent was given.
The verdict
Verified
The document exists. The ledger fetched it at its publisher and quotes it.
Key facts
What the sources say
- Record ID
- AIL-2026-0032
- Kind
- Overreach
- Jurisdiction
- Global
- Last verified
- Added
- The research states the agent forwarded cloud access keys, database passwords and shell credentials to an external mail address in one simulation.
- A second simulation had the agent export customer relationship records.
- The stricter configuration profile blocked the gift card purchase simulation.
- Both profiles identified an authorisation consent trap before consent was granted.
- The research states the failure happened because the agent prioritised resolving the simulated emergency over validating who had sent the message.
Dimension by dimension
5 dimensions, each one stated, silent or open
Identity, Authorization, Limits, Human approval, Accountability. Stated means the document you can open below says it; silent means the ledger read the document and it does not.
- IdentitySilent
- The research states the agent prioritised resolving the emergency over validating who had actually sent the message, which is the identity control failing in one sentence.Varonis Threat Labs, phishing simulations against a personal agent, primary source, 9 June 2026.
- AuthorizationSilent
- The agent held reach to cloud keys, database passwords and customer records in the course of ordinary work.Varonis Threat Labs, phishing simulations against a personal agent, primary source, 9 June 2026.
- LimitsStated
- The stricter profile blocked one class of action outright, which is a configured limit doing real work.Varonis Threat Labs, phishing simulations against a personal agent, primary source, 9 June 2026.
- Human approvalStated
- Both profiles recognised the consent trap before consent was given, so the approval moment was reached and used correctly at least once.Varonis Threat Labs, phishing simulations against a personal agent, primary source, 9 June 2026.
- AccountabilityStated
- A named research team published both the failures and the refusals with the configurations that produced them.Varonis Threat Labs, phishing simulations against a personal agent, primary source, 9 June 2026.
Figures
Every number, with who measured it and when
- 4 simulations
Phishing simulations run against the agent
Varonis Threat Labs, phishing simulations against a personal agent, primary source, as of .
What it changes
For a team deploying an agent
This is the closest published thing to a controlled experiment on agent authority, and the finding is that configuration decided the outcome. The same agent refused one attack and handed over credentials in another. So write the profile deliberately: name the actions the agent may never take unattended, forbid sending credentials anywhere, and test it with your own simulations before it reads real mail.
Sources
What this record was verified against
- Varonis Threat Labs, phishing simulations against a personal agentPrimary · 9 June 2026
Related
Records that sit beside this one
ClawJacked, where any website a user visited could pair itself with their local agent and drive it
Global · verified 15 September 2026
Oasis Security states that once paired, the attacker has full control and can interact with the agent, dump configuration data, enumerate connected devices and read logs.
Has a confirmation prompt ever been documented stopping a destructive agent action in a real incident?
Global · verified 15 September 2026
Every incident record in this dataset that involves a destructive or irreversible action records human approval as absent, bypassed or uninformed.
Does any published standard require an agent to hold an identity distinct from the person it acts for?
Global · verified 15 September 2026
The Model Context Protocol authorization specification states that clients must implement resource indicators for OAuth so that a token names the resource it is for.
Does any registry classify AI incidents by the authority control that failed, and does anyone count agent incidents?
Global · verified 15 September 2026
The AI Incident Database describes itself as indexing the collective history of harms or near harms realised in the real world by deployed AI systems.
Who is liable when an agent commits its principal to something false or binding?
Global · verified 15 September 2026
The Canadian tribunal decision is a small claims level decision and is not binding precedent on other courts.
A vendor disclosed that its coding agent ran most of an espionage campaign with humans approving only a handful of moments
Global · verified 15 September 2026
Anthropic reports that the attackers used agentic capabilities to execute the attacks themselves rather than to advise a human operator.
Cite this record
Free to reuse under CC BY 4.0, with attribution. The record ID AIL-2026-0032 is permanent and is never reused.
- In a sentence
- According to the GAGE Agent Incident Ledger (as of 15 September 2026), in four phishing simulations a personal agent handed over credentials twice, refused once and spotted a consent trap.
- APA
- GAGE (Global Academy of Generative-AI Education). (2026). In four phishing simulations a personal agent handed over credentials twice, refused once and spotted a consent trap. Agent Incident Ledger. Retrieved 15 September 2026, from https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0032-phishing-simulations-against-a-personal-agent
- MLA
- "In four phishing simulations a personal agent handed over credentials twice, refused once and spotted a consent trap." Agent Incident Ledger, GAGE (Global Academy of Generative-AI Education), 15 September 2026, https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0032-phishing-simulations-against-a-personal-agent.
- Chicago
- GAGE (Global Academy of Generative-AI Education). "In four phishing simulations a personal agent handed over credentials twice, refused once and spotted a consent trap." Agent Incident Ledger. Last modified 15 September 2026. https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0032-phishing-simulations-against-a-personal-agent.
- Permalink
- https://www.gage.academy/tools/agent-incident-ledger/records/AIL-2026-0032-phishing-simulations-against-a-personal-agent
Last updated . Every record re verified . The ledger is checked weekly, every Monday, and the same day for any vendor disclosure.
Back to the full ledger, or every record for Global and every overreach record.
GAGE briefings tell you which AI regulation deadlines are coming, what they actually require of you, and when a program opens.