Skip to content

Part 1: Tokenizers

From bytes to subwords. The tracer’s byte tokenizer (256 ids, no training) gives way to the Tokenizer protocol and three trained families: byte-level BPE (GPT-2 compatible, loading Hugging Face files), WordPiece (BERT), and the Unigram LM (EM, Viterbi, subword sampling). Then BPE is implemented independently in Rust (tl-tok) with a streaming UTF-8 decoder; Python and Rust compare behavior using tokenizer.json and shared fixtures. The part ends with metrics that decide a vocabulary size for the capstone.

Course passes: 3 (L1.1 to L1.6, milestone MS-L1, part of gate MS-P3)

Before you start: the just-in-time math M05.2, M06.2, M07.1, M07.2, M11.2 and the solve parts S-M06b, S-M07b; for L1.5, the Rust primer and the Rust structures ds.05 and ds.06 in Systems Data Structures.

  • A tokenizer is a contract: encode(decode(ids)) == ids and decode(encode(text)) == text for every string, including invalid UTF-8 split across a stream.
  • BPE merges the most frequent adjacent pair until the vocabulary is full; encoding replays the merges by rank. A heap over pairs (ds.06) makes training fast.
  • Unigram starts from a large vocabulary and prunes by likelihood loss; it can sample segmentations, which the capstone ablation (BPE vs Unigram) measures.
  • Compression is the first metric: bytes per token and bits per byte on held-out text (L1.6).
ModuleTopicKindPass
L1.1Tokenizer protocol, char tokenizerbuild3
L1.2Byte-level BPE (GPT-2 compatible): hand-written pre-tokenizer, trainer, HF loaderbuild3
L1.3WordPiece (BERT basic tokenizer + greedy longest match)build3
L1.4Unigram LM tokenizer (EM, Viterbi, subword sampling)build3
L1.5Rust fast BPE (tl-tok) with a streaming UTF-8 decoder and fixture parity with Pythonbuild3
L1.6Tokenizer metricsbuild3
#ModuleChapterKindPass
1L1.1Tokenizer protocol, char tokenizerbuild3
2L1.2Byte-level BPE (GPT-2 compatible): pre-tokenizer, trainer, HF loaderbuild3
3L1.3WordPiece (BERT basic tokenizer + greedy longest match)build3
4L1.4Unigram LM tokenizer (EM, Viterbi, subword sampling)build3
5L1.5Rust fast BPE (tl-tok), byte tokenizer, streaming decoderbuild3
6L1.6Tokenizer metricsbuild3