Inference Performance
How weights and requests become tokens, fast and cheaply: engine internals, quantization, the cross-language framework map, serving and load testing, and the platforms around the engine.
For: Performance engineers (MTS), ML engineers who own serving, and anyone who sizes deployments.
Included in: Superstar FDE. Progress made here counts there too.
The stages, their chapters, and what “done” means for each are in path.tsv.
practice/bin/ol learn inference-performance # stages and progresspractice/bin/ol learn inference-performance next # read the next unfinished stageStages
Section titled “Stages”Generated from path.tsv. Track progress locally with practice/bin/ol learn.
| Stage | Read | Done when |
|---|---|---|
| 1 | Inference engine internals LLM Systems & Inference | You can explain prefill vs decode, paged KV cache, continuous batching, and speculative decoding to both a researcher and a customer. |
| 2 | Quantization Quantization: Math → Code | You can pick a quantization format (FP8, AWQ, GPTQ, NF4, GGUF) for a given GPU and quality bar and say what it does to memory, speed, and accuracy. |
| 3 | Inference frameworks across languages Inference Frameworks: The Cross-Language Landscape | You can place llama.cpp, vLLM, SGLang, TensorRT-LLM, candle, and mistral.rs on the deployment map and say when each is the right tool. |
| 4 | Serving and load Serving, Capacity & Load Testing: Math → Code | You benchmark vLLM and SGLang on the same model, produce TTFT and TPOT vs concurrency curves, and pick a config that meets a stated SLO at the lowest cost per 1M tokens. |
| 5 | Serving platforms and reproducibility LLM Serving Platforms: How They Work and When to Use Them | You can say which layer (router, gateway, orchestration, managed platform, engine) a customer problem lives in, and why batched inference is not bitwise reproducible by default. |