Adversarial manipulation
A failure mode in which a bad output is deliberately induced by an attacker, through a jailbreak that talks the model past its own guardrails or a prompt injection hidden in content the model reads, rather than arising on its own. Distinct from ordinary fabrication because the trigger is chosen and repeatable by anyone who finds it. Deep treatment and defenses belong to Topic 4.3.
Defined in 2 GAGE programs, which carry 2 distinct definitions of it. The wording above is taught in AI Governance: Applied Mastery.
How each discipline defines it
The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.
A failure mode in which a bad output is deliberately induced by an attacker, through a jailbreak that talks the model past its own guardrails or a prompt injection hidden in content the model reads, rather than arising on its own. Distinct from ordinary fabrication because the trigger is chosen and repeatable by anyone who finds it. Deep treatment and defenses belong to Topic 4.3.
Deliberately causing an AI system to misbehave by crafting its inputs, most commonly through prompt injection (input designed to override the system's instructions). A rising threat as AI systems are connected to real tools and data.
Where it is taught
The exact lessons this term appears in. The first 7 topics of every program are free with a free account.
- The license to govern: explaining to a skeptic exactly how your model fails · Build Before You Govern, AI Governance: Applied Mastery
- AI Risk Management · Governance and Responsible AI, Business AI Transformation
Terms it appears with
Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.