Skip to content

Direct preference optimization and IPO

ModuleL12.2 · optional build · Python · Pass 10 · 2 to 3 h
You buildtinyllm/post/dpo.py: DPO and IPO objectives
Contractdpo.pyi
Testscourse/tests/L12.2/ (why: closed-form preference odds, reference equality, and IPO scaling)
NeedsS-M10b KL-regularized objective, M11.1 KL
Used byC2 post-trained capstone
MilestoneMS-C2
Optional depthIPO and odds-ratio preference optimization
  • DPO optimizes preference odds relative to a frozen reference policy.
  • The reference term prevents unbounded movement away from the SFT policy.
  • IPO changes the loss shape and is not just a different learning rate.
Terminal window
ol start L12.2
ol tests L12.2
ol check L12.2

SFT teaches response form, but does not encode which of two valid responses a user prefers. Paired responses provide a direct training signal without fitting a separate reward model.

For chosen and rejected completions, define Δπ = log π(y+|x) - log π(y-|x) and Δref similarly. DPO loss is -log σ(β(Δπ-Δref)). The reference policy is fixed. At policy equal to reference, the margin is zero and the per-pair loss is log 2.

SymbolMeaning
πtrainable policy
πreffrozen reference policy
βpreference strength / KL scale
σlogistic sigmoid

With Δπ = 1.2, Δref = 0.2, and β = 0.1, the scaled margin is 0.1, so the loss is -log σ(0.1) ≈ 0.6444. If both policies are equal, both deltas match and the loss is log 2 ≈ 0.6931.

Implement DPO and IPO over completion log probabilities. test_hand_dpo_loss checks the worked value, test_equal_policy_reference_is_log_two checks the equal-policy baseline, and test_ipo_variant checks IPO scaling. Never update the reference weights.

TestWhy it existsExpected result
test_hand_dpo_lossPins the DPO objective to the worked calculationLoss near 0.6444
test_equal_policy_reference_is_log_twoVerifies reference correctionEqual policies produce log(2)
test_ipo_variantChecks the alternate IPO objectiveMatches the hand-computed fixture
PitfallCaught by
Reversing chosen and rejectedtest_hand_dpo_loss; mutant s01
Omitting reference log probabilitiestest_hand_dpo_loss; mutant s02
DirectionModuleHow it uses this
BackM11.1Defines the KL term that regularizes policy movement.
BackS-M10bDerives how KL regularization leads to reference-adjusted preference odds.
ForwardC2Applies DPO to preference pairs after L12.1 SFT.
Your pieceProduction equivalentWhat it addsWhere to look
paired preference objectivepreference optimization pipelinesAnnotator disagreement and pair provenanceRafailov et al., Direct Preference Optimization
held-out preference checkonline preference evaluationPosition controls and outcome measures beyond training lossTRL preference trainer