Skip to content

Capstone 2: post-trained TinyStories model

ModuleC2 · optional practice · Python and Go · Pass 10 · 1 to 2 days
You buildan SFT and preference or verifiable-reward run, a model artifact, and an updated model card
Contractformats/checkpoint.md, formats/eval-results.schema.json, and the L12.* contracts
Testscourse/tests/C2/ (why: records run provenance and prevents unsupported model-card claims)
NeedsC1, L12.1, L12.2, L12.3, L12.4
Used byMS-C2 checks held-out reward improvement
MilestoneMS-C2
Optional depthL12.4 distillation, or compare all three adaptation methods
  • Post-training starts from the C1 checkpoint and keeps a reproducible base reference.
  • The held-out suite is fixed before training and reports uncertainty.
  • The model card distinguishes measured outcomes from intended behavior.
Terminal window
ol start C2
ol milestone MS-C2 --smoke

C1 proves the model can be trained, evaluated, released, and served. This capstone asks whether demonstrations and preference or verifiable-reward data improve a specific behavior while preserving language-model quality.

Use the same held-out examples before and after adaptation. For a score difference d_i = score_after_i - score_before_i, report the mean difference and a paired confidence interval; the comparison preserves prompt difficulty. Do not claim a general safety improvement from a small task-specific dataset. Keep a frozen copy of C1 and record dataset revision, tokenizer, template, seed, optimizer, and objective.

SymbolMeaning
d_iper-example paired score change
nnumber of held-out examples
CIuncertainty interval for the mean paired change

If three held-out reward differences are [0.2, 0.1, 0.3], the mean gain is 0.2. This tiny sample is not enough to claim a reliable improvement; the capstone uses a larger fixture and requires a paired interval excluding zero with p < 0.05.

Train from C1 using SFT, then DPO or GRPO; distillation is a documented alternate. Evaluate reward and language-model regression, compare with paired bootstrap or permutation, and publish an artifact manifest. The model card records intended uses, limitations, data provenance, evaluation setup, and third-party teacher/model licenses. Smoke checks verify paths and schema; the full optional milestone checks statistical gain.

PitfallCaught by
Evaluate on training pairsHeld-out split check; mutant s01
Report only the best seedRun manifest check; mutant s02
Omit base-model provenanceModel-card artifact check; mutant s03
DirectionModuleHow it uses this
BackC1Provides the pretrained checkpoint and baseline evaluation.
BackL12.1Converts demonstrations into an SFT checkpoint.
BackL12.2Applies preference-pair optimization.
BackL12.3Applies verifiable reward optimization.
BackL12.4Provides distillation as an alternate recipe.
ForwardMS-C2Evaluates the post-trained artifact.
ForwardMS-P10Remains the agent pass gate; the capstone is optional.

Larger alignment programs add red-team review, multiple annotator cohorts, data governance, and staged deployment. Those processes are required before treating an offline score as a product benefit.