Skip to content

Supervised fine-tuning: chat rendering, assistant-only loss, packing

ModuleL12.1 · optional build · Python · Pass 10 · 2 to 3 h
You buildtinyllm/post/sft.py: rendering, assistant masks, and document packing
Contractsft.pyi
Testscourse/tests/L12.1/ (why: renderer parity, loss masks, and isolation between packed examples)
NeedsL6.6 LoRA, L0.3 cross-entropy, L10.5 chat template
Used byC2 post-trained capstone
MilestoneMS-C2
Optional depthHugging Face TRL supervised fine-tuning
  • A chat template is a byte-level protocol and must match the serving engine.
  • Only assistant response tokens contribute to the supervised objective.
  • Packed documents need an attention boundary as well as a loss mask.
Terminal window
ol start L12.1
ol tests L12.1
ol check L12.1

C1 gives you a pretrained TinyStories model. Demonstrations teach its response format and task behavior while preserving the serving template already used by L10.5.

For token ids x_0...x_(T-1), causal cross entropy is evaluated only where the boolean assistant mask m_t is true: L = -sum_t m_t log p(x_(t+1)|x_<=t) / max(1, sum_t m_t). System and user tokens provide context but receive no target loss. A document boundary also prevents attention from one packed example into another.

SymbolMeaning
x_ttoken id at position t
m_t1 for an assistant target, otherwise 0
Tpacked sequence length

For messages user: hi, assistant: hello, rendering yields the template’s user prefix, hi, assistant prefix, hello, and end marker. If these occupy token positions 0 through 5, the user and role markers have mask 0, while the hello target positions have mask 1. The test checks exact rendered ids and the resulting mask.

Implement render_chat, assistant_loss_mask, and pack to the stub contract. Packing returns token ids, loss mask, and document ids. test_hand_rendered_chat checks exact rendering, test_assistant_only_mask checks loss targets, and test_packed_documents_do_not_attend_across_boundaries checks packed document ids.

TestWhy it existsExpected result
test_hand_rendered_chatPins the serving byte protocolExact token ids for the worked conversation
test_assistant_only_maskEnsures prompts are context, not targetsOnly assistant response positions contribute loss
test_packed_documents_do_not_attend_across_boundariesPrevents leakage between packed examplesAttention cannot cross document ids
PitfallCaught by
Training on user tokenstest_assistant_only_mask; mutant s01
Adding an EOS to every packed sample without a masktest_packed_documents_do_not_attend_across_boundaries; mutant s02
Replacing each packed document id with the first idtest_packed_documents_do_not_attend_across_boundaries; mutant s02
DirectionModuleHow it uses this
BackL6.6Provides the adapted model parameters for the SFT run.
BackL0.3Supplies the token-level cross-entropy objective.
ForwardC2Starts from this SFT checkpoint before preference optimization and preserves the serving chat contract.
Your pieceProduction equivalentWhat it addsWhere to look
packlength bucketing and packing in TRLHigher device utilization, with careful document masksHugging Face TRL SFT trainer
assistant maskdataset mixture and loss weightingChanges which tokens define the training distributionml/07-training-and-post-training