Skip to main content

Adversarial manipulation

A failure mode in which a bad output is deliberately induced by an attacker, through a jailbreak that talks the model past its own guardrails or a prompt injection hidden in content the model reads, rather than arising on its own. Distinct from ordinary fabrication because the trigger is chosen and repeatable by anyone who finds it. Deep treatment and defenses belong to Topic 4.3.

Defined in 2 GAGE programs, which carry 2 distinct definitions of it. The wording above is taught in AI Governance: Applied Mastery.

How each discipline defines it

The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.

AI Governance: Applied Mastery

A failure mode in which a bad output is deliberately induced by an attacker, through a jailbreak that talks the model past its own guardrails or a prompt injection hidden in content the model reads, rather than arising on its own. Distinct from ordinary fabrication because the trigger is chosen and repeatable by anyone who finds it. Deep treatment and defenses belong to Topic 4.3.

Business AI Transformation

Deliberately causing an AI system to misbehave by crafting its inputs, most commonly through prompt injection (input designed to override the system's instructions). A rising threat as AI systems are connected to real tools and data.

Where it is taught

The exact lessons this term appears in. The first 7 topics of every program are free with a free account.

Terms it appears with

Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.