Reward hacking
A severe form of the mismatched-objective split in which a model finds an unintended, often aggressive shortcut to maximize its objective, a shortcut the person who set the objective never anticipated and would reject if they saw it. Reward hacking is Goodhart's problem in its sharpest form: the proxy is not just imperfect, it is being actively exploited.
Defined in 2 GAGE programs, which carry 2 distinct definitions of it. The wording above is taught in AI Governance: Applied Mastery.
How each discipline defines it
The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.
A severe form of the mismatched-objective split in which a model finds an unintended, often aggressive shortcut to maximize its objective, a shortcut the person who set the objective never anticipated and would reject if they saw it. Reward hacking is Goodhart's problem in its sharpest form: the proxy is not just imperfect, it is being actively exploited.
When an AI system optimizes its reward signal by exploiting flaws in the evaluation, scoring highly without accomplishing the intended objective; the machine version of gaming a measure, and a mirror of the human gaming traps.
Where it is taught
The exact lessons this term appears in. The first 7 topics of every program are free with a free account.
- Learning from Feedback vs. Gaming Feedback · Assessment and Continuous Learning, AI Literacy & Professional Conduct
- Train a model with your own hands and watch what it actually learns · Build Before You Govern, AI Governance: Applied Mastery
Terms it appears with
Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.