Unifying GRPO, Dr. GRPO, and DAPO: The Group-Standard-Deviation Identity
Recent research reveals that three prominent language model training methods—GRPO, Dr. GRPO, and DAPO—are fundamentally variations of a single mechanism. They all adjust a single metric: the standard deviation of sampled answers to a given prompt. This standard deviation directly correlates with the magnitude of the training update, indicating that disagreement among responses is a crucial driver of learning.