Skip to main content

Indirect prompt injection

An attack in which malicious instructions are embedded inside content an agent processes as data, such as a webpage, a document, a ticket, or a third-party chatbot's response, rather than delivered directly by the agent's own user. The agent's user need not know the attack exists, and the attacker need not have any relationship with the agent's user or operator. Named directly in the Model AI Governance Framework for Agentic AI (IMDA v1.5, p.40) as the risk surfaced inside the Google x Singapore Government sandbox.

Defined in 4 GAGE programs, which carry 13 distinct definitions of it. The wording above is taught in Agentic AI Governance: Applied Mastery.

How each discipline defines it

The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.

Agentic AI Governance: Applied Mastery

An attack in which malicious instructions are embedded inside content an agent processes as data, such as a webpage, a document, a ticket, or a third-party chatbot's response, rather than delivered directly by the agent's own user. The agent's user need not know the attack exists, and the attacker need not have any relationship with the agent's user or operator. Named directly in the Model AI Governance Framework for Agentic AI (IMDA v1.5, p.40) as the risk surfaced inside the Google x Singapore Government sandbox.

AI Literacy & Professional Conduct

An attack in which the adversary hides instructions inside content the AI later reads on someone else's behalf (a web page, email, document, image, or other format the AI processes), so an innocent user triggers the attack by doing normal work (Greshake et al., 2023). As AI tools process more content types (images, PDFs, audio), the potential carrier formats for indirect injection expand accordingly.

AI Governance: Applied Mastery

Prompt injection in which the malicious instruction is hidden inside external content the system ingests (a document, email, web page, or ticket), so the system follows a command the attacker never typed into the conversation. The failure most teams forget because the instruction does not arrive through the chat box.

AI Data Governance: The Data Chair

A form of prompt injection where the attacker plants adversarial instructions inside content (a document, webpage, email, ticket) that an AI system is expected to retrieve and process later, on behalf of a different, legitimate user, without the attacker ever interacting with the AI system directly.

AI Literacy & Professional Conduct

A security-relevant failure mode in which instructions embedded in a retrieved document or external content manipulate the model's output, causing it to suppress caveats, assert things the document's author wanted stated, or override the user's original intent. Especially relevant in workflows that ingest documents, web pages, or user-supplied files.

Where it is taught

The exact lessons this term appears in. The first 7 topics of every program are free with a free account.

Terms it appears with

Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.