Skip to content

Part 7: The Modern Decoder Block

What changed between 2017 and a Llama-family model. Pre-LN with RMSNorm, gated MLPs (SwiGLU, GeGLU), rotary position embeddings and context extension (PI, NTK, YaRN, Llama-3 scaling), grouped-query and multi-head latent attention, sliding windows and attention sinks, mixture of experts, and the Llama-family model that loads Hugging Face configs and weights and matches SmolLM2-135M.

Course passes: 5 (L7.1 to L7.9, milestone MS-L7, part of gate MS-P5)

Before you start: the just-in-time math M05.1 (parameter and FLOP counting); Part 5 and Part 6; rotations from Precalculus.

  • RoPE rotates each pair of query and key dimensions by an angle proportional to position, so their dot product depends only on the relative offset.
  • KV bytes are the serving cost: GQA shares key and value heads across query heads; MLA compresses them into a latent and absorbs the up-projection into the query.
  • MoE routes each token to kk of EE experts, so parameters grow without growing per-token compute; balancing the load is the hard part.
ModuleTopicKindPass
L7.1Pre-LN, RMSNormbuild5
L7.2Gated MLPs: SwiGLU, GeGLUbuild5
L7.3RoPE (half and interleaved layouts, partial rotary)build5
L7.4Context extension: PI, NTK, YaRN, Llama-3 scaling; ALiBibuild5
L7.5MQA/GQA attention with cache hook, window, learned sinksbuild5
L7.6Multi-head latent attention (DeepSeek-V2/V3), weight absorptionbuild5
L7.7Sliding window, StreamingLLM sinks, learned sinksbuild5
L7.8Mixture of Experts: routing, sorted dispatch, Switch aux loss, aux-free biasbuild5
L7.9Llama-family model, HF config, weight loading, downloaderbuild5

The ALiBi part of L7.4 is optional (course decision D31).

#ModuleChapterKindPass
1L7.1Pre-LN, RMSNormbuild5
2L7.2Gated MLPs: SwiGLU, GeGLUbuild5
3L7.3RoPE (half and interleaved layouts, partial rotary)build5
4L7.4Context extension: PI, NTK, YaRN, Llama-3 scaling; ALiBi (optional part)build5
5L7.5MQA/GQA attention with cache hook, window, learned sinksbuild5
6L7.6Multi-head latent attention (DeepSeek-V2/V3), weight absorptionbuild5
7L7.7Sliding window, StreamingLLM sinks, learned sinksbuild5
8L7.8Mixture of Experts: routing, sorted dispatch, Switch aux loss, aux-free biasbuild5
9L7.9Llama-family model, HF config, weight loading, downloaderbuild5