Skip to content

tinyllm: Build an LLM from Scratch

  • Primary reference: the course itself: every chapter here is one module you build and ol check grades. Read the system map first.
  • Supplementary: Karpathy, Neural Networks: Zero to Hero (free); Jurafsky and Martin, Speech and Language Processing (free), chapters 3 (n-gram LMs) and 9 (transformers); Raschka, Build a Large Language Model (From Scratch)
  • Prerequisites: the Pass 0 primers (Python and numpy, C) and, just in time, the math each part names (course DESIGN 7.5)
  • Estimated time: the spine of course passes 1 to 11, about 60 weeks part-time; Pass 1 (the tracer) is about 5 weeks
  • An LLM system is a stack of contracts: a checkpoint format between training and serving, a C ABI between Python and Rust and the kernels, an HTTP API between the engine and everything in front of it. Build each side of each contract yourself and nothing in the stack is magic.
  • Start with the thinnest model that exercises every contract: a byte-level bigram fitted by counting, served by a real engine. Every later model (autograd bigram, RNN, transformer, the Llama family) upgrades a component behind an unchanged contract.
  • Numbers are only right if they are tested: every module ships course tests, a reference that passes them, and a stub that fails them.

Follow the course path (practice/bin/ol learn course), not this directory’s order: parts interleave with math, systems, and operations modules. For each chapter run ol start <ID>, read beats 1 to 3 before writing code, implement against the interface in beat 4, and ol check <ID> until green. ol tests <ID> explains what each test checks and why.


A language model is a function from a prefix of token ids to a distribution over the next id. Everything else in the stack is plumbing that makes that function cheap to train, cheap to run, and safe to expose, and each piece of plumbing has a contract you can test in isolation.

Key ideas:

  • Byte tokenizer: 256 ids, no training, no unknown tokens; decoding a stream must handle incomplete UTF-8.
  • Count bigram: P(b∣a)=(cab+α)/(∑xcax+256α)P(b \mid a) = (c_{ab} + \alpha) / (\sum_x c_{ax} + 256\alpha); its negative log-likelihood is the baseline every later model must beat.
  • Logits in Python: the bigram model computes logits in the reference implementation.
  • Checkpoint: Python writes model.safetensors; Rust Candle reads the same weights. Shared fixture files check cross-language behavior.

Key ideas:

  • Autograd and training (p00), tokenizers (p01), statistical LMs (p02), recurrent networks (p03), attention (p04), the 2017 transformer (p05), objectives (p06), the modern decoder block (p07), inference (p08), C kernels (p09), serving in Rust (p10), training at scale (p11), and an optional post-training part (p12).
PartDirectoryChapters so farCourse pass
p00FoundationsL0.0 byte bigram1 (tracer), 2
p01Tokenizersnone yet (B4)3
p02Statistical LMsnone yet (B5)3
p03Recurrent Networksnone yet (B6)4
p04Attention Originsnone yet (B6)4
p05The Transformer (2017)none yet (B7)5
p06Objectives and Adaptationnone yet (B7)5
p07The Modern Decoder Blocknone yet (B7)5
p08Inferencenone yet (B8)6
p09Kernels in Crt.01 the C ABI1 (tracer), 6
p10ServingL10.0 your first endpoint1 (tracer), 7
p11Training at Scalenone yet (B11)9
p12Post-Training (optional)none yet (B12)10
C1, C2Capstonesnone yet (B11, B12)9, 10

Each part README lists its planned modules; chapters arrive with their authoring batches, and the course path lists what is available.

TrackConnection
Linear AlgebraM03.1, the naive matmul in C every logit goes through
Gatewaygw.00, the front door to the engine
Containers and Kubernetesdep.00, the images and charts the engine runs in
LLM Systems & Inferencethe production systems this spine rebuilds in miniature
CompanyPractice
Hugging Facesafetensors, transformers checkpoints, tokenizers
llama.cpp / ggmlC kernels behind a stable ABI, served over an OpenAI-compatible HTTP API
vLLM, SGLangPython model code over custom kernels, a separate serving engine