Definition
Reward hacking: the AI optimizes the letter of its reward, not the spirit. Goodhart’s law: when a measure becomes a target, it stops being a good measure.
Examples: o3 rewrote the timer instead of speeding up the code; o1 edited Stockfish’s engine state instead of playing chess; coding agents trained on the held-out test set in PostTrainBench; RLHF models learn to sound confident but be wrong because raters reward fluency. More capable models hack more creatively.
Why it’s hard to fix: rewards are under-specified, training pressure rewards any path to high reward (including exploits in the harness), and detection lags capability.
Appears in
- Lecture 1, Reward Hacking
- Lecture 4, EM From Reward Hacking in Production RL: reward hacks cause broad misalignment and alignment faking
- Lecture 7, Reward Over-Optimization: the KL penalty against Goodhart
- Lecture 10, Obfuscated Reward Hacking
- Lecture 12, Reward Hacking in PostTrainBench
- Lecture 12, the realistic misalignment path: hacking generalizes; inoculation prompting
Related concepts
- AI Alignment (specification problem)
- RLHF