Definition

RLVR replaces human preferences with programmatic verification (correct math answer, passing unit tests, format checks). No learned reward model.

GRPO advantage

Sample answers to one prompt, score them, use the group-relative z-score.

DeepSeek-R1 was trained with RLVR only and developed long chains of thought and self-verification (β€œwait, let me check…”) on its own.

Appears in