Skip to content

Part 2: Statistical Language Models

The language models before deep learning and the first neural one. An n-gram model with interpolated modified Kneser-Ney smoothing sets the baseline every later model must beat and becomes the corpus pipeline’s perplexity filter in the capstone. Bengio’s 2003 neural probabilistic LM is the first model trained with your autograd on real text, and word2vec (skip-gram with negative sampling) with a PPMI-SVD baseline shows where embeddings come from.

Course passes: 3 (L2.1 to L2.3, milestone MS-L2, part of gate MS-P3)

Before you start: the just-in-time math M03.5, M03.6, M11.4 and the solve parts S-M03b, S-M11b; the tokenizers of Part 1.

  • Smoothing is the whole game for counts: Kneser-Ney discounts every count and redistributes the mass by how many contexts a word continues, not how often it occurs.
  • The NPLM replaces counts with a learned embedding and an MLP, and generalizes to unseen n-grams through similar embeddings.
  • Embeddings factorize co-occurrence: skip-gram with negative sampling implicitly factorizes a shifted PMI matrix, which PPMI plus SVD does explicitly.
ModuleTopicKindPass
L2.1n-gram LM, interpolated modified Kneser-Neybuild3
L2.2Bengio NPLM (2003)build3
L2.3word2vec SGNS + PPMI-SVD baselinebuild3
#ModuleChapterKindPass
1L2.1n-gram language model with interpolated modified Kneser-Neybuild3
2L2.2Bengio’s neural probabilistic language modelbuild3
3L2.3word2vec skip-gram with negative sampling, and the PPMI-SVD baselinebuild3