Skip to content

Inference Frameworks: The Cross-Language Landscape

Where you run an LLM. This maps the inference/serving ecosystem across languages — the C/C++ foundation (llama.cpp/ggml), the Python serving tier (vLLM, TGI, SGLang, TensorRT-LLM), the Rust equivalents (candle, mistral.rs, burn, ratchet, luminal), and the emerging Zig stack (ZML, llama.cpp.zig). The throughline: every framework is the same three layers — a kernel/tensor core, a runtime, and a server — assembled in a different language for a different deployment target.

Parent topic: LLM Systems & Inference. Quantization formats these frameworks consume are in Quantization: Math → Code. Runnable examples: code/.

For deployment choices and operating costs, see LLM Serving Platforms: Anyscale, Ray, vLLM, SGLang, RouteLLM, and OpenRouter.

  • Every inference framework is three layers: (1) a kernel/tensor core (ggml, CUTLASS, candle-core, MLIR/XLA), (2) a runtime (graph execution + KV cache + sampling), (3) a server (batching scheduler + an OpenAI-compatible HTTP API). Frameworks differ in which layers they own and which they borrow
  • GGUF is the lingua franca of portable inference: a single quantized model file (from llama.cpp’s ecosystem) runs under llama.cpp, candle, mistral.rs, and Zig bindings alike
  • Language picks the deployment target: Python for max-throughput GPU serving, C for “runs anywhere with no runtime”, Rust for a single safe static binary at the edge, Zig for minimal-dependency native compilation and effortless C interop
  • The OpenAI-compatible HTTP API is the universal contract — swap llama.cpp’s llama-server for vLLM for mistral.rs for ZML’s LLMD and clients don’t change
  • Run the same GGUF model under llama.cpp (llama-server), candle, and mistral.rs; compare tokens/sec, memory, and binary/footprint
  • For each framework, identify its three layers and which it borrows (e.g. mistral.rs borrows candle’s kernels; Ollama borrows llama.cpp’s runtime)
  • Build one example from code/ per language to compare each language’s runtime and file interfaces

There is no single “inference framework” — there is a stack, and each project stakes out a slice of it. Python frameworks (vLLM) maximize datacenter GPU throughput. C (llama.cpp) maximizes portability and minimizes dependencies so a model runs on a laptop. Rust (candle, mistral.rs) trades a little ecosystem maturity for memory safety and a single static binary. Zig (ZML) bets on compiling the whole model graph to a standalone native artifact via MLIR. Pick the language and you’ve largely picked the deployment story.

LayerJobExamples
Kernels / tensor corematmul, attention, dequant on CPU/GPUggml, CUTLASS/cuBLAS, candle-core, Metal/Vulkan shaders, MLIR+XLA
Runtimeload weights, run the graph, manage KV cache, sample tokensllama.cpp, candle-transformers, mistral.rs engine, ZML
Scheduler / servercontinuous batching, request queue, OpenAI HTTP APIllama-server, vLLM, TGI, SGLang, mistralrs-server, LLMD

Most “frameworks” are a vertical slice of these. Ollama = llama.cpp runtime + a friendly model manager. mistral.rs = candle kernels + its own runtime + an OpenAI server.

The most-deployed inference engine on earth.

  • ggml: a dependency-free C tensor library with a static computation graph and a quantization-first design. Backends: CPU (AVX/NEON), Metal, CUDA, Vulkan, SYCL, ROCm
  • llama.cpp: the transformer runtime + llama-server (OpenAI-compatible). Defines GGUF (the quantized model container) and the k-quant / i-quant formats (see Quantization §5.5)
  • The C API (llama.h) is the integration point for every other language — Python (llama-cpp-python), Rust (llama-cpp-2), Zig (@cImport), Go, etc. all bind to it
  • Family: whisper.cpp, stable-diffusion.cpp, ggml-based projects share the same core
  • Downstream: Ollama, LM Studio, Jan, GPT4All, KoboldCpp all wrap llama.cpp

Why it wins for edge/local: no Python, no CUDA required, a single small binary, runs a 7B model on a laptop CPU. Tradeoff: lower datacenter GPU throughput than vLLM (no PagedAttention-class scheduler historically; continuous batching is newer).

The datacenter-throughput layer, covered in the parent topic §9. In one line each:

  • vLLM — PagedAttention + continuous batching, the general-purpose throughput king
  • TensorRT-LLM — compiled NVIDIA kernels, FP8, max performance on H100
  • TGI — HuggingFace’s production server
  • SGLang — RadixAttention prefix sharing, fast structured output
  • Ray Serve / KServe — orchestration on top of any of the above

These dominate GPU clusters; the C/Rust/Zig tiers below dominate edge, embedded, and dependency-constrained deployments.

4. The Rust Tier (the “Rust equivalents”)

Section titled “4. The Rust Tier (the “Rust equivalents”)”

Rust trades ecosystem breadth for memory safety + a single static binary + no GIL. The current stack:

FrameworkLayer(s)NicheNotes
candlekernels + runtimeminimalist ML, the Rust “PyTorch-lite”CPU/CUDA/Metal/WASM; loads GGUF + safetensors; HuggingFace-backed
mistral.rsruntime + serverturnkey fast inferencebuilds on candle; ISQ (in-situ quant), OpenAI server, vision/tools
burnkernels + runtimebackend-agnostic train and inferWGPU/CUDA/NdArray/LibTorch; own quantization API (Quant §7)
ratchetkernels + runtimeweb/cross-platform GPUwgpu-based, browser-first inference
luminalcompiler + runtimesearch-compiled kernelstiny graph IR, compiles to fast kernels
llama-cpp-2Rust interfacereuse llama.cpp from Rustllama_cpp is the higher-level wrapper

Deprecated: rustformers/llm (and the old llama-rs) are unmaintained — the ecosystem consolidated onto candle + mistral.rs. Use those.

When to pick Rust: edge/embedded, WASM/browser, a CLI you ship as one binary, or a service where you want Rust’s safety end-to-end. See code/candle_generate.rs (GGUF generation with candle) and code/mistralrs_serve.rs (OpenAI server + ISQ).

Zig’s pitch: trivial C interop (@cImport reads C headers directly — no binding generator) and compile the whole thing to a small native artifact. Two distinct approaches:

ProjectApproachNotes
ZMLcompile model graphs via MLIR + OpenXLAZig + XLA → standalone native binaries; NVIDIA/AMD/TPU/Trainium from one codebase; LLMD OpenAI server in a ~2.4 GB image
llama.cpp.zigbindings + build.zig over llama.cppreuse the C engine from Zig via @cImport
llm.zigfrom scratch, no depspedagogical full-stack inference in pure Zig
llama2.zigport of Karpathy’s llama2.cthe cleanest “read the whole inference loop” reference

The two idioms:

  1. Wrap C — @cImport({ @cInclude("llama.h"); }) and call llama.cpp directly. Zig’s interop makes this less code than the equivalent Rust bindgen setup. See code/llama_cpp.zig
  2. Pure Zig kernels — write the hot loops yourself with @Vector SIMD and explicit allocators. See code/quant_dot.zig (INT8 symmetric quantize + SIMD dot, the CPU twin of quantization/code/symmetric_quant.cu)

When to pick Zig: you want C-level control and a tiny dependency-free binary, you’re already calling C libraries (Zig is arguably the best C-interop language), or you want compile-time graph specialization (ZML).

Deployment targetLanguageFrameworkWhy
GPU cluster, max throughputPythonvLLM / TensorRT-LLMPagedAttention, FP8, mature scheduler
Laptop / CPU / “just runs”Cllama.cpp (llama-server)no deps, GGUF, every backend
Single safe static binary, edgeRustmistral.rs / candlememory-safe, OpenAI server, ISQ
Browser / WASMRustcandle / ratchetnear-native in-browser inference
Multi-vendor native compileZigZMLMLIR/XLA → standalone binary, AMD/TPU/Trainium
Reuse C engine, minimal binaryZigllama.cpp.zig@cImport, no binding generator
Train + infer in one stackRustburnbackend-agnostic, own quant API

C/C++PythonRustZig
Flagshipllama.cppvLLMcandle / mistral.rsZML
Kernel coreggml, CUTLASSTriton, CUTLASScandle-coreMLIR/XLA, @Vector
Quant formatsGGUF k/i-quantsGPTQ/AWQ/FP8GGUF, ISQGGUF (via C), native
Serverllama-servervLLM/TGI/SGLangmistralrs-serverLLMD
Best atportabilitythroughputsafe single binarynative compile + C interop
File exchangecheckpoint filessafetensors, tokenizer.jsonsafetensors, tokenizer.jsonserialized model formats
ConceptConnected TrackApplication
GGUF k-quants, ISQ, FP8Quantization: Math → CodeThe formats these frameworks load
Continuous batching, KV cache, SLOsLLM Systems & InferenceWhat the runtime/server layers implement
Polyglot idioms (C, Rust, Zig)Polyglot PracticeSame algorithm, different language tradeoffs
K8s, autoscaling, OpenAI APICloud NativeDeploying any of these servers
CompanyHow This AppearsDifficulty
Meta / ggml-orgllama.cpp, ggml, GGUF — the C foundationExpert
HuggingFacecandle, ratchet, TGI — Rust + Python tiersExpert
ZML (Paris)Zig + MLIR/XLA multi-vendor inferencePhD-level
Ollama / LM Studioproductizing llama.cpp for local useAdvanced
NVIDIATensorRT-LLM, CUTLASS kernels under the runtimesExpert