Part 3: Recurrent Networks
Sequence models with state. A vanilla RNN with backpropagation through time written by hand, then the LSTM and GRU in torch’s gate order (so weights load from torch checkpoints), bidirectional RNNs with length-aware reversal, and an RNN language model trained with stateful truncated BPTT. ELMo’s bidirectional LM and scalar mix is optional.
Course passes: 4 (L3.1 to L3.6, milestone MS-L3, part of gate MS-P4)
Before you start: the matrix-calculus VJPs (M08.3) and gradient clipping (M10.4); the autograd engine of Part 0.
Key ideas
Section titled “Key ideas”- BPTT is reverse mode through time: the gradient of an early state is a product of Jacobians, which vanishes or explodes; clipping and gating are the cures.
- Gates are learned multiplexers: the LSTM’s forget gate lets gradient flow through the cell state almost unchanged.
- Truncation is a memory budget: stateful TBPTT carries the hidden state across batches but cuts the gradient every steps.
Modules
Section titled “Modules”| Module | Topic | Kind | Pass |
|---|---|---|---|
L3.1 | Vanilla RNN with manual BPTT and truncation | build | 4 |
L3.2 | LSTM (torch gate order i, f, g, o) | build | 4 |
L3.3 | GRU (torch gate order r, z, n) | build | 4 |
L3.4 | Bidirectional RNN with length-aware reversal | build | 4 |
L3.5 | ELMo: biLM, ScalarMix, linear probes | build | 4, optional |
L3.6 | RNN language model with stateful TBPTT | build | 4 |
The worked example for L3.2 is lstm_cell.c from Neural Architectures, whose RNN and LSTM sections these chapters absorb.
Chapters
Section titled “Chapters”| # | Module | Chapter | Kind | Pass |
|---|---|---|---|---|
| 1 | L3.1 | Vanilla RNN with manual BPTT and truncation | build | 4 |
| 2 | L3.2 | LSTM (torch gate order) | build | 4 |
| 3 | L3.3 | GRU (torch gate order) | build | 4 |
| 4 | L3.4 | Bidirectional RNN with length-aware reversal | build | 4 |
| 5 | L3.5 | ELMo: biLM, ScalarMix, linear probes | side | 4 |
| 6 | L3.6 | RNN language model with stateful TBPTT | build | 4 |
Going further
Section titled “Going further”- Olah, Understanding LSTM Networks; Karpathy, The Unreasonable Effectiveness of Recurrent Neural Networks.
- Hochreiter and Schmidhuber, Long Short-Term Memory (1997); Cho et al., GRU; Peters et al., ELMo.