Part 6: Objectives and Adaptation
One architecture, many objectives. GPT’s causal language modeling (with GPT-2 weight loading), BERT’s masked LM, ELECTRA’s replaced-token detection, and T5’s span corruption (optional); fine-tuning heads for sequences, tokens, and rewards, with the linear-head export the gateway’s usage policy uses; LoRA with PiSSA initialization; and the evaluation harness behind the model-zoo table every architecture in the course reports to.
Course passes: 5 (L6.1 to L6.7, milestone MS-L6, part of gate MS-P5)
Before you start: the just-in-time math M01.4, M07.5, M07.7 and the solve part S-M07d; the transformer of Part 5.
Key ideas
Section titled “Key ideas”- The objective decides what the model learns: predicting the next token gives a generator, reconstructing masked tokens gives an encoder, detecting replaced tokens trains on every position.
- LoRA freezes and learns with rank ; PiSSA initializes and from the top singular vectors of .
- Evaluation is a measurement with error bars: strided perplexity, log-likelihood multiple choice, and paired comparisons with confidence intervals (
L6.7).
Modules
Section titled “Modules”| Module | Topic | Kind | Pass |
|---|---|---|---|
L6.1 | GPT decoder-only, causal LM loss, GPT-2 weight loading | build | 5 |
L6.2 | BERT encoder and MLM masking | build | 5 |
L6.3 | ELECTRA replaced-token detection | build | 5 |
L6.4 | T5 span corruption, relative position buckets | build | 5, optional |
L6.5 | Fine-tuning heads (sequence, token, reward) and the linear-head export (D33) | build | 5 |
L6.6 | LoRA with PiSSA init and merge | build | 5 |
L6.7 | LM evaluation harness: strided perplexity, multiple choice, tasks with CIs, paired comparison, and the model zoo (D36) | build | 5 |
Chapters
Section titled “Chapters”| # | Module | Chapter | Kind | Pass |
|---|---|---|---|---|
| 1 | L6.1 | GPT decoder-only, causal LM loss, GPT-2 weight loading | build | 5 |
| 2 | L6.2 | BERT encoder and MLM masking | build | 5 |
| 3 | L6.3 | ELECTRA replaced-token detection | build | 5 |
| 4 | L6.4 | T5 span corruption and relative position buckets | side | 5 |
| 5 | L6.5 | Fine-tuning heads and the linear-head export | build | 5 |
| 6 | L6.6 | LoRA with PiSSA init and merge | build | 5 |
| 7 | L6.7 | LM evaluation harness and the model zoo | build | 5 |