Skip to main content

Reward hacking

A severe form of the mismatched-objective split in which a model finds an unintended, often aggressive shortcut to maximize its objective, a shortcut the person who set the objective never anticipated and would reject if they saw it. Reward hacking is Goodhart's problem in its sharpest form: the proxy is not just imperfect, it is being actively exploited.

Defined in 2 GAGE programs, which carry 2 distinct definitions of it. The wording above is taught in AI Governance: Applied Mastery.

How each discipline defines it

The same term does different work depending on who is using it. These are the definitions as each program teaches them, unedited.

AI Governance: Applied Mastery

A severe form of the mismatched-objective split in which a model finds an unintended, often aggressive shortcut to maximize its objective, a shortcut the person who set the objective never anticipated and would reject if they saw it. Reward hacking is Goodhart's problem in its sharpest form: the proxy is not just imperfect, it is being actively exploited.

AI Literacy & Professional Conduct

When an AI system optimizes its reward signal by exploiting flaws in the evaluation, scoring highly without accomplishing the intended objective; the machine version of gaming a measure, and a mirror of the human gaming traps.

Where it is taught

The exact lessons this term appears in. The first 7 topics of every program are free with a free account.

Terms it appears with

Not an alphabetical neighbourhood: these are the terms taught in the same lessons, ranked by how often they appear together.