Skip to content

Inference Performance

How weights and requests become tokens, fast and cheaply: engine internals, quantization, the cross-language framework map, serving and load testing, and the platforms around the engine.

For: Performance engineers (MTS), ML engineers who own serving, and anyone who sizes deployments.

Included in: Superstar FDE. Progress made here counts there too.

The stages, their chapters, and what “done” means for each are in path.tsv.

Terminal window
practice/bin/ol learn inference-performance # stages and progress
practice/bin/ol learn inference-performance next # read the next unfinished stage

Generated from path.tsv. Track progress locally with practice/bin/ol learn.

StageReadDone when
1Inference engine internals
LLM Systems & Inference
You can explain prefill vs decode, paged KV cache, continuous batching, and speculative decoding to both a researcher and a customer.
2Quantization
Quantization: Math → Code
You can pick a quantization format (FP8, AWQ, GPTQ, NF4, GGUF) for a given GPU and quality bar and say what it does to memory, speed, and accuracy.
3Inference frameworks across languages
Inference Frameworks: The Cross-Language Landscape
You can place llama.cpp, vLLM, SGLang, TensorRT-LLM, candle, and mistral.rs on the deployment map and say when each is the right tool.
4Serving and load
Serving, Capacity & Load Testing: Math → Code
You benchmark vLLM and SGLang on the same model, produce TTFT and TPOT vs concurrency curves, and pick a config that meets a stated SLO at the lowest cost per 1M tokens.
5Serving platforms and reproducibility
LLM Serving Platforms: How They Work and When to Use Them
You can say which layer (router, gateway, orchestration, managed platform, engine) a customer problem lives in, and why batched inference is not bitwise reproducible by default.