Skip to content

Group relative policy optimization with verifiable rewards

ModuleL12.3 · optional build · Python · Pass 10 · 3 to 4 h
You buildtinyllm/post/grpo.py: grouped advantages, clipped objective, rollout client
Contractgrpo.pyi
Testscourse/tests/L12.3/ (why: hand objective, zero-advantage gradient, and toy reward improvement)
NeedsM11.1 KL, M07.4 uncertainty, L8.1 sampling, L10.5 engine API
Used byC2 post-trained capstone
MilestoneMS-C2
Optional depthPolicy gradients and PPO in ml/03-reinforcement-learning
  • Normalize rewards within each prompt’s sampled group.
  • Clip the probability ratio to limit each update.
  • Verifiable rewards can be checked without a learned reward model.
  • Record rollouts and seeds so a run can be reproduced.
Terminal window
ol start L12.3
ol tests L12.3
ol check L12.3

Some tasks have objective checks, such as exact format or arithmetic correctness. Grouped rollouts let the model learn from those checks without a human preference label for each pair.

For rewards r_i in a group, use A_i = (r_i - mean(r)) / (std(r) + ε). Let q_i = exp(logp_new - logp_old). The clipped policy objective is the negative mean of min(q_i A_i, clip(q_i, 1-εc, 1+εc) A_i), plus a nonnegative KL penalty. With identical rewards the centered advantages are zero, so the policy term and its gradient are zero.

SymbolMeaning
r_iverifiable reward for sample i
A_inormalized within-group advantage
q_inew-to-old policy probability ratio
εcclipping range

Rewards [0, 1, 1] have mean 2/3; their advantages have one negative and two positive values, summing to zero. If all three rewards equal 1, every advantage is zero and the policy gradient vanishes. The tests use this exact group.

Implement group normalization and clipped loss. test_hand_grpo_objective checks the clipped arithmetic, test_grouped_rewards_are_centered checks normalization, and test_zero_advantage_has_zero_policy_gradient checks the zero-variance boundary.

TestWhy it existsExpected result
test_hand_grpo_objectivePins the clipped objectiveMatches the hand-computed scalar
test_grouped_rewards_are_centeredScopes normalization to one promptGroup advantages have zero mean
test_zero_advantage_has_zero_policy_gradientChecks equal-reward boundaryPolicy gradient is zero
PitfallCaught by
Normalize rewards across unrelated promptstest_grouped_rewards_are_centered; mutant s01
Omit clipping from the policy objectivetest_hand_grpo_objective; mutant s02
DirectionModuleHow it uses this
BackM11.1Supplies the KL penalty.
BackM07.4Supports uncertainty-aware evaluation of reward outcomes.
BackL8.1Samples rollout tokens.
BackL10.5Provides the serving API and chat template.
ForwardC2Runs verifiable reward training after SFT and DPO; reports reward, KL, clip fraction, and rollout revision.
Your pieceProduction equivalentWhat it addsWhere to look
grouped rollout objectiveasynchronous rollout poolsMore throughput and explicit rollout revisionsDeepSeekMath GRPO paper
verifiable rewardcalibrated reward and tool executionValidation against reward hacking and distribution shiftml/03-reinforcement-learning