Part 0: Foundations
The model side of the tracer and, from Pass 2, the autograd engine every later model trains with. Pass 1 has one chapter: a byte-level bigram fitted by counting, whose logits use NumPy matmul and whose weights you write as a safetensors checkpoint the Rust engine serves. Pass 2 retrains the same bigram with your autograd (L0.5 takes over bigram.py) behind the same checkpoint contract.
Course passes: 1 (L0.0, gate MS-P1), 2 (L0.1 to L0.6, gate MS-P2).
Before you start: Python and NumPy and matrix multiplication in Python.
| # | Module | Chapter | Kind | Pass |
|---|---|---|---|---|
| 1 | L0.0 | Byte bigram: counts, logits, safetensors v0 | build | 1 |
| 2 | L0.1 | Tensor, broadcasting backward, and no_grad | build | 2 |
| 3 | L0.2 | The op library and gradcheck_all | build | 2 |
| 4 | L0.3 | Losses with a fused backward | build | 2 |
| 5 | L0.4 | Module system and basic layers | build | 2 |
| 6 | L0.5 | Training loop and the autograd bigram | build | 2 |
| 7 | L0.6 | Safetensors for every dtype, atomic checkpoints, and the token stream | build | 2 |