Post-training
This optional module sequence follows C1. It uses supervised demonstrations, pairwise preferences, verifiable rewards, and teacher distributions to adapt a small causal language model.
| # | Module | Chapter | Kind | Pass |
|---|---|---|---|---|
| 1 | L12.1 | Supervised fine-tuning: chat rendering, assistant-only loss, packing | build | 10 |
| 2 | L12.2 | Direct preference optimization and IPO | build | 10 |
| 3 | L12.3 | Group relative policy optimization with verifiable rewards | build | 10 |
| 4 | L12.4 | Distillation with forward and reverse KL | build | 10 |