Training and Post-Training
How a model is made before it is served: RL foundations, pretraining, SFT, preference and RL post-training, LoRA, and how to scope a customer fine-tune.
For: Research-adjacent ML engineers, and anyone selling or running fine-tuning.
Included in: Superstar FDE. Progress made here counts there too.
The stages, their chapters, and what “done” means for each are in path.tsv.
practice/bin/ol learn training # stages and progresspractice/bin/ol learn training next # read the next unfinished stageStages
Section titled “Stages”Generated from path.tsv. Track progress locally with practice/bin/ol learn.
| Stage | Read | Done when |
|---|---|---|
| 1 | Reinforcement learning foundations Reinforcement Learning | You can derive the policy gradient and the PPO clipped objective, and say what the critic is for. |
| 2 | Training and post-training Training & Post-Training | You run a LoRA fine-tune of a small model, evaluate it against the base on a held-out set, and estimate GPU-hours and cost for an 8B customer job before starting. |