Skip to content

Part 10: Serving

The Rust engine: an OpenAI-compatible HTTP server that streams completions from a checkpoint through Candle tensor operations. Pass 1 has one chapter here: a std-only HTTP/1.1 and SSE server (no async runtime, no dependencies) serving the byte bigram, with one hand-written OTLP span per request. Pass 7 rebuilds it on tokio and hyper with continuous batching, chunked prefill, disaggregation, speculative decoding, and tool calls, behind the same API.

Course passes: 1 (L10.0, gate MS-P1), 7 (L10.1 to L10.9, load.01, load.02, gate MS-P7).

Before you start: the Rust and HTTP and SSE primers; the API contract openai-subset.v0.yaml and the engine role in spec/cli-roles.md.

#ModuleChapterKindPass
1L10.0Your first endpoint: a std-only Rust HTTP/1.1 + SSE serverbuild1
2L10.1Candle model runner (Llama + bigram, int4), Rust KV cache, sampler and PCG32build7
3L10.2Continuous batching scheduler with priority and preemptionbuild7
4L10.3Chunked prefill (mixed prefill + decode batches)build7
5L10.4Block manager with prefix cache (none, hash, radix)build7
6L10.5tl-serve: OpenAI-compatible HTTP + SSE on tokio/hyperbuild7
7L10.6Disaggregated prefill/decode, KV transfer, heartbeat clientbuild7
8L10.7Serving metrics, SLO histograms, OTel spans and propagationbuild7
9L10.8Speculative decoding in the enginebuild7
10L10.9Tool calls and constrained JSON decoding in the enginebuild7
11load.01Open-loop load generator, log-linear histogram, Go PCG32 portbuild7
12load.02Run comparison and regression gateside7