Skip to content

Part 0: Foundations

The model side of the tracer and, from Pass 2, the autograd engine every later model trains with. Pass 1 has one chapter: a byte-level bigram fitted by counting, whose logits use NumPy matmul and whose weights you write as a safetensors checkpoint the Rust engine serves. Pass 2 retrains the same bigram with your autograd (L0.5 takes over bigram.py) behind the same checkpoint contract.

Course passes: 1 (L0.0, gate MS-P1), 2 (L0.1 to L0.6, gate MS-P2).

Before you start: Python and NumPy and matrix multiplication in Python.

#ModuleChapterKindPass
1L0.0Byte bigram: counts, logits, safetensors v0build1
2L0.1Tensor, broadcasting backward, and no_gradbuild2
3L0.2The op library and gradcheck_allbuild2
4L0.3Losses with a fused backwardbuild2
5L0.4Module system and basic layersbuild2
6L0.5Training loop and the autograd bigrambuild2
7L0.6Safetensors for every dtype, atomic checkpoints, and the token streambuild2