Part 1: Tokenizers
From bytes to subwords. The tracer’s byte tokenizer (256 ids, no training) gives way to the Tokenizer protocol and three trained families: byte-level BPE (GPT-2 compatible, loading Hugging Face files), WordPiece (BERT), and the Unigram LM (EM, Viterbi, subword sampling). Then BPE is implemented independently in Rust (tl-tok) with a streaming UTF-8 decoder; Python and Rust compare behavior using tokenizer.json and shared fixtures. The part ends with metrics that decide a vocabulary size for the capstone.
Course passes: 3 (L1.1 to L1.6, milestone MS-L1, part of gate MS-P3)
Before you start: the just-in-time math M05.2, M06.2, M07.1, M07.2, M11.2 and the solve parts S-M06b, S-M07b; for L1.5, the Rust primer and the Rust structures ds.05 and ds.06 in Systems Data Structures.
Key ideas
Section titled “Key ideas”- A tokenizer is a contract:
encode(decode(ids)) == idsanddecode(encode(text)) == textfor every string, including invalid UTF-8 split across a stream. - BPE merges the most frequent adjacent pair until the vocabulary is full; encoding replays the merges by rank. A heap over pairs (
ds.06) makes training fast. - Unigram starts from a large vocabulary and prunes by likelihood loss; it can sample segmentations, which the capstone ablation (BPE vs Unigram) measures.
- Compression is the first metric: bytes per token and bits per byte on held-out text (
L1.6).
Modules
Section titled “Modules”| Module | Topic | Kind | Pass |
|---|---|---|---|
L1.1 | Tokenizer protocol, char tokenizer | build | 3 |
L1.2 | Byte-level BPE (GPT-2 compatible): hand-written pre-tokenizer, trainer, HF loader | build | 3 |
L1.3 | WordPiece (BERT basic tokenizer + greedy longest match) | build | 3 |
L1.4 | Unigram LM tokenizer (EM, Viterbi, subword sampling) | build | 3 |
L1.5 | Rust fast BPE (tl-tok) with a streaming UTF-8 decoder and fixture parity with Python | build | 3 |
L1.6 | Tokenizer metrics | build | 3 |
Chapters
Section titled “Chapters”| # | Module | Chapter | Kind | Pass |
|---|---|---|---|---|
| 1 | L1.1 | Tokenizer protocol, char tokenizer | build | 3 |
| 2 | L1.2 | Byte-level BPE (GPT-2 compatible): pre-tokenizer, trainer, HF loader | build | 3 |
| 3 | L1.3 | WordPiece (BERT basic tokenizer + greedy longest match) | build | 3 |
| 4 | L1.4 | Unigram LM tokenizer (EM, Viterbi, subword sampling) | build | 3 |
| 5 | L1.5 | Rust fast BPE (tl-tok), byte tokenizer, streaming decoder | build | 3 |
| 6 | L1.6 | Tokenizer metrics | build | 3 |
Going further
Section titled “Going further”- Hugging Face tokenizers (Rust BPE merges and the pre-tokenizer pipeline) and OpenAI tiktoken.
- Sennrich, Haddow, and Birch, Neural Machine Translation of Rare Words with Subword Units; Kudo, Subword Regularization.