Skip to content

Post-training

This optional module sequence follows C1. It uses supervised demonstrations, pairwise preferences, verifiable rewards, and teacher distributions to adapt a small causal language model.

#ModuleChapterKindPass
1L12.1Supervised fine-tuning: chat rendering, assistant-only loss, packingbuild10
2L12.2Direct preference optimization and IPObuild10
3L12.3Group relative policy optimization with verifiable rewardsbuild10
4L12.4Distillation with forward and reverse KLbuild10