Skip to content

Part 6: Objectives and Adaptation

One architecture, many objectives. GPT’s causal language modeling (with GPT-2 weight loading), BERT’s masked LM, ELECTRA’s replaced-token detection, and T5’s span corruption (optional); fine-tuning heads for sequences, tokens, and rewards, with the linear-head export the gateway’s usage policy uses; LoRA with PiSSA initialization; and the evaluation harness behind the model-zoo table every architecture in the course reports to.

Course passes: 5 (L6.1 to L6.7, milestone MS-L6, part of gate MS-P5)

Before you start: the just-in-time math M01.4, M07.5, M07.7 and the solve part S-M07d; the transformer of Part 5.

  • The objective decides what the model learns: predicting the next token gives a generator, reconstructing masked tokens gives an encoder, detecting replaced tokens trains on every position.
  • LoRA freezes WW and learns W+BAW + BA with rank rr; PiSSA initializes AA and BB from the top singular vectors of WW.
  • Evaluation is a measurement with error bars: strided perplexity, log-likelihood multiple choice, and paired comparisons with confidence intervals (L6.7).
ModuleTopicKindPass
L6.1GPT decoder-only, causal LM loss, GPT-2 weight loadingbuild5
L6.2BERT encoder and MLM maskingbuild5
L6.3ELECTRA replaced-token detectionbuild5
L6.4T5 span corruption, relative position bucketsbuild5, optional
L6.5Fine-tuning heads (sequence, token, reward) and the linear-head export (D33)build5
L6.6LoRA with PiSSA init and mergebuild5
L6.7LM evaluation harness: strided perplexity, multiple choice, tasks with CIs, paired comparison, and the model zoo (D36)build5
#ModuleChapterKindPass
1L6.1GPT decoder-only, causal LM loss, GPT-2 weight loadingbuild5
2L6.2BERT encoder and MLM maskingbuild5
3L6.3ELECTRA replaced-token detectionbuild5
4L6.4T5 span corruption and relative position bucketsside5
5L6.5Fine-tuning heads and the linear-head exportbuild5
6L6.6LoRA with PiSSA init and mergebuild5
7L6.7LM evaluation harness and the model zoobuild5