Skip to content

Training and Post-Training

How a model is made before it is served: RL foundations, pretraining, SFT, preference and RL post-training, LoRA, and how to scope a customer fine-tune.

For: Research-adjacent ML engineers, and anyone selling or running fine-tuning.

Included in: Superstar FDE. Progress made here counts there too.

The stages, their chapters, and what “done” means for each are in path.tsv.

Terminal window
practice/bin/ol learn training # stages and progress
practice/bin/ol learn training next # read the next unfinished stage

Generated from path.tsv. Track progress locally with practice/bin/ol learn.

StageReadDone when
1Reinforcement learning foundations
Reinforcement Learning
You can derive the policy gradient and the PPO clipped objective, and say what the critic is for.
2Training and post-training
Training & Post-Training
You run a LoRA fine-tune of a small model, evaluate it against the base on a held-out set, and estimate GPU-hours and cost for an 8B customer job before starting.