Skip to content

Part 4: Attention Origins

Where attention came from: sequence-to-sequence translation. An encoder-decoder trained with teacher forcing, then Bahdanau’s additive attention and Luong’s dot, general, and concat scores with input feeding, a generic beam search over any step function, and the metrics that grade generated sequences (exact match, BLEU, chrF).

Course passes: 4 (L4.1 to L4.5, milestone MS-L4, part of gate MS-P4)

Before you start: the just-in-time math M07.4 and the solve part S-M07c; the recurrent layers of Part 3.

  • The bottleneck: a fixed-size encoder state cannot hold a long sentence; attention lets each decoder step read a weighted average of all encoder states.
  • Scores to weights: a score per source position, a softmax, a weighted sum; the 2017 transformer keeps exactly this and drops the recurrence.
  • Search is separate from the model: beam search keeps the kk best prefixes by summed log-probability, with length normalization.
ModuleTopicKindPass
L4.1Encoder-decoder with teacher forcingbuild4
L4.2Bahdanau additive attentionbuild4
L4.3Luong attention (dot, general, concat) + input feedingbuild4
L4.4Beam search (generic over step_fn)build4
L4.5Sequence metrics: exact match, BLEU, chrFbuild4
#ModuleChapterKindPass
1L4.1Encoder-decoder with teacher forcingbuild4
2L4.2Bahdanau additive attentionbuild4
3L4.3Luong attention and input feedingbuild4
4L4.4Beam search, generic over a step functionbuild4
5L4.5Sequence metrics: exact match, BLEU, chrFbuild4