Skip to content

Part 8: Inference

Generating text fast and exactly. The sampler and logit processors in the course’s fixed op order, the KV cache with incremental decode and an incremental UTF-8 detokenizer, the paged KV block pool in C (rt.04) and the Python paged cache over it, a radix prefix cache, quantization (int8, int4, fp8, KV), speculative decoding, and constrained decoding from regex and JSON schema to token masks.

Course passes: 6 (rt.04, L8.1 to L8.7, milestone MS-L8, part of gate MS-P6)

Before you start: the just-in-time math M07.6, M09.3, M09.4 and the solve part S-M09b; the C structures ds.01 to ds.03 and the radix tree ds.07 in Systems Data Structures; the Llama-family model of Part 7.

  • Decoding is memory-bound: each new token reads every weight and the whole KV cache once, so bytes moved, not FLOPs, set the speed.
  • Paging stores the KV cache in fixed-size blocks with a block table per sequence; shared prefixes share blocks by reference count and copy on write.
  • Sampling is specified to the bit (spec/sampling.md): penalties, temperature, top-k, top-p, min-p, an f64 softmax, one uniform draw, inverse CDF, so Python and Rust emit identical tokens.
  • Speculative decoding drafts several tokens cheaply and verifies them in one forward pass; with greedy acceptance the output is unchanged.
ModuleTopicKindPass
rt.04Optional standalone C paged KV block pool: refcount, copy on write, chained block hashes, prefix index, LRU, export and import (format v1); parity via fixturesoptional6
L8.1Sampling and logit processors (spec/sampling.md)build6
L8.2KV cache, incremental decode, incremental UTF-8 detokenizer, generatebuild6
L8.3Paged KV cache implemented in Pythonbuild6
L8.4Radix prefix cache (block-granular, LRU leaf eviction, locks)build6
L8.5Quantization: int8 per-channel, int4 group (packed), fp8, KV quantbuild6
L8.6Speculative decoding: n-gram, prompt-lookup, and model draftsbuild6
L8.7Constrained decoding: regex to DFA to token masks, JSON-schema subsetbuild6
#ModuleChapterKindPass
1L8.1Sampling and logit processorsbuild6
2L8.2KV cache, incremental decode, incremental UTF-8 detokenizer, generatebuild6
3L8.3Paged KV cache in Pythonside6
4L8.4Radix prefix cache (block-granular, LRU leaf eviction, locks)build6
5L8.5Quantization: int8, int4 group (packed), fp8, KV quantbuild6
6L8.6Speculative decoding: n-gram, prompt-lookup, and model draftsbuild6
7L8.7Constrained decoding: regex to DFA to token masks, JSON-schema subsetbuild6
8rt.04Paged KV block pool (format v1, optional C)side6