Skip to content

Part 5: The Transformer (2017)

Attention is all you need, built from its parts. Scaled dot-product attention with its backward pass, the masks (causal, padding, sliding window, additive), multi-head attention, sinusoidal and learned positions, and the full encoder-decoder transformer with both LayerNorm placements and the Noam schedule.

Course passes: 5 (L5.1 to L5.5, milestone MS-L5, part of gate MS-P5)

Before you start: the solve parts S-M05 (logic) and S-M08 (VJP derivations); rotations and frequency ladders from Precalculus; attention from Part 4.

  • Scaled dot product: softmax(QK⊤/dk+M)V\text{softmax}(QK^{\top}/\sqrt{d_k} + M)V, where the mask MM adds −∞-\infty to forbidden positions.
  • Heads split the model dimension so each head attends with its own projections; the outputs are concatenated and projected back.
  • LayerNorm placement decides trainability: post-LN as in 2017 needs warmup, pre-LN trains without it.
ModuleTopicKindPass
L5.1Scaled dot-product attention (forward and backward)build5
L5.2Masks: causal, padding, sliding window, additivebuild5
L5.3Multi-head attentionbuild5
L5.4Positional encodings: sinusoidal, learnedbuild5
L5.5Encoder-decoder Transformer with both LayerNorm placements (post-LN as in 2017, pre-LN as norm='pre'), Noam schedule, label smoothingbuild5
#ModuleChapterKindPass
1L5.1Scaled dot-product attention (forward and backward)build5
2L5.2Masks: causal, padding, sliding window, additivebuild5
3L5.3Multi-head attentionbuild5
4L5.4Positional encodings: sinusoidal, learnedbuild5
5L5.5Encoder-decoder Transformer: post-LN and pre-LN, Noam, label smoothingbuild5