Definition
RLVR replaces human preferences with programmatic verification (correct math answer, passing unit tests, format checks). No learned reward model.
GRPO advantage
Sample answers to one prompt, score them, use the group-relative z-score.
DeepSeek-R1 was trained with RLVR only and developed long chains of thought and self-verification (βwait, let me checkβ¦β) on its own.
Appears in
- Lecture 2, RLVR
- Lecture 7, Rubrics as Rewards: RL beyond verifiable domains