AI Platform Engineering
The infrastructure layer underneath modern AI products — but taught as patterns, not products. Frameworks, protocols, caches, and orchestrators churn constantly; the patterns underneath them barely move. Master the handful of fundamental patterns and every new tool becomes a recognizable instance of something you already understand.
The math analogy: once you understand limits, derivatives, and integrals, every named technique downstream is a special case. The same is true here. Cache-aside, consistent hashing, idempotent retry, event-sourced replay, the worker pool, backpressure, the decorator — learn the pattern and how it works, and “what’s trending this quarter” becomes derivative. We name the trendy tools (vLLM, Temporal, Redis, gRPC) as instances of patterns, and we note how companies apply them, but the pattern is the lesson.
Prerequisites: Deep Learning (transformers, training loops), LLM Systems & Inference (serving, KV cache, parallelism), System Design, Cloud Native, Data Engineering (OLTP vs OLAP, storage), Functional Programming. Comfortable with Concurrency & Systems.
Prerequisite Graph
Section titled “Prerequisite Graph”graph LR
DL[Deep Learning] --> TF[Training & Frameworks]
LLM[LLM Systems] --> TF
CN[Cloud Native] --> TF
SD[System Design] --> RPC[RPC & Protocols]
RPC --> SSE[Streaming & SSE]
LLM --> SSE
SD --> DDC[Distributed Data & Caching]
DE[Data Engineering] --> DDC
DDC --> ORC[Orchestration & Workers]
TF --> ORC
RPC --> ORC
FP[Functional Programming] --> CDP[Coding & Design Patterns]
CDP --> ORC
CDP --> TF
TF --> RAG[Retrieval & RAG]
DDC --> RAG
RAG --> AUTH[Authorization & Access Control]
TF --> EVAL[LLM Evaluation]
RAG --> EVAL
TF --> EDGE[Edge, Realtime & On-Device]
SSE --> EDGE
EVAL --> ROUTE[Model Routing & Cascades]
TF --> ROUTE
LLM --> ROUTE
SSE --> GWY[Gateway]
ROUTE --> GWY
GWY --> AGT[Agent SDK]
ORC --> AGT
RAG --> AGT
The Split: Three Different “Distributed” Problems
Section titled “The Split: Three Different “Distributed” Problems”A recurring confusion is that “distributed X” is one topic. It is three, and this track keeps them separate because the patterns differ:
| Problem | Lives in | The pattern family |
|---|---|---|
| Distributed training — one model across many GPUs | 01 Training & Frameworks | Parallelism & sharding (data/tensor/pipeline, FSDP/ZeRO) |
| Distributed data — one dataset across many nodes | 04 Distributed Data & Caching | Partitioning, replication, caching, invalidation |
| Distributed orchestration — one workflow across many failures | 05 Orchestration & Workers | Durable execution, checkpoints, the worker pool |
Topics
Section titled “Topics”| # | Topic | Pattern lens | Primary Reference | Time |
|---|---|---|---|---|
| 01 | Training & Frameworks | Parallelism, sharding, adapters | PyTorch + JAX + vLLM docs | 4-6 weeks |
| 02 | RPC & Protocols | Contracts, serialization, evolution | gRPC + Protobuf docs | 2-3 weeks |
| 03 | Streaming & SSE | Push, backpressure, cancellation | HTML SSE spec + MDN | 1-2 weeks |
| 04 | Distributed Data & Caching | Partition, replicate, cache, invalidate | Redis + Cassandra docs + DDIA | 3-4 weeks |
| 05 | Durable Orchestration & Workers | Durable execution, workers, observability, profiling | Temporal + DBOS docs | 2-3 weeks |
| 06 | Coding & Design Patterns | Decorators, facade, closures, currying | Refactoring Guru + Mostly Adequate Guide | 2-3 weeks |
| 07 | Retrieval & RAG | Encoders, chunking, hybrid fusion, multimodal | RAG paper + BM25 + pgvector | 3-4 weeks |
| 08 | Authorization & Access Control | RBAC/ABAC/ReBAC/NGAC, the pushdown ladder | NIST RBAC/ABAC/NGAC + Zanzibar | 2-3 weeks |
| 09 | LLM Evaluation | Cross-entropy/perplexity/bits-per-byte, judges | MacKay + HELM + Ragas | 2-3 weeks |
| 10 | Edge, Realtime & On-Device Inference | Deployment spectrum, streaming encoders, efficiency architectures | llama.cpp + Mistral 7B + Mamba | 2-3 weeks |
| 11 | Model Routing & Cascades | Oracle vs router, cascades, escalation, cache-aware switching, decision models | RouteLLM + FrugalGPT + LLMRouterBench | 1-2 weeks |
| 12 | Gateway | The reverse proxy whose payload is a stream: keys, trace propagation, then limits, routing, caching (course) | OpenAI API reference + W3C Trace Context | course passes 1, 7, 10 |
| 13 | Agent SDK | The model proposes, the program disposes: types, tools, the loop, gates, durable runs (course) | Building effective agents + ReAct (free) | course pass 10 |
Key Takeaways
Section titled “Key Takeaways”- Learn the pattern, not the product. vLLM is continuous batching + paged cache; Temporal is event-sourced durable execution + worker pool; Redis is an in-memory hash table you must invalidate. Name the pattern and the tool is interchangeable.
- “Distributed” is three problems, not one — training (parallelism), data (partition/replicate/cache), orchestration (durable execution). Different failure modes, different patterns; this track keeps them apart on purpose.
- Caching is the second-hardest problem in CS, and invalidation is why. A cache is a lie you tell for speed; every caching pattern (cache-aside, write-through, write-back, TTL) is a different answer to “when does the lie expire?”
- Durable execution is checkpointing for workflows. Temporal persists every step (the checkpoint), replays history to recover, runs your code on a pool of workers, and gives you distributed observability over long-running processes for free.
- Profiling beats guessing, always. Whether it’s a GPU kernel, a slow query, a cache miss rate, or a workflow’s tail latency — measure first. The bottleneck is rarely where intuition says.
- Design patterns are how these systems are built. Decorators wrap workflows and retries; the facade hides a serving stack behind one call; closures and currying are how JAX, middleware, and configuration actually work. The patterns in topic 06 recur in every topic above.
- RAG is information retrieval wearing a neural coat. Retrieve-then-generate; hybrid (dense + BM25 + graph) fused with RRF; bi-encode to find and cross-encode to rank. Retrieval quality — not the model — is the usual ceiling.
- Authorization is a complexity-ladder problem. Flat RBAC is a lookup; nested groups and hierarchies are graph reachability (pushdown, not finite-state) — which is why ReBAC/Zanzibar exist, and why secure RAG filters by permission during retrieval.
- Evaluation starts in information theory. Cross-entropy = bits to predict the next token = compression; perplexity and bits-per-byte are re-normalizations. Intrinsic loss steers training; extrinsic evals (benchmarks, judges, RAG faithfulness) decide if it’s good.
- Inference is a spectrum from datacenter to phone, and “realtime” is an architecture choice. llama.cpp + GGUF is the end-to-end local path; streaming encoders must be causal/chunked; and Mistral’s SWA, GQA, MoE, and Mamba/SSM variants are four different escapes from the vanilla transformer’s cost curve.
- Routing is approximating an oracle you cannot have. No model is best at everything, so a pool beats its best member, but only on paper. A real router chooses before the outcome exists, has to price the cache it abandons when it switches, and should escalate on category or a verifier, not on model confidence alone.
How to Use This Track
Section titled “How to Use This Track”- Coming from ML / research: 01 Training & Frameworks → 04 Data & Caching → 05 Orchestration — see the platform around your model.
- Coming from backend / distributed systems: 02 RPC → 04 Data & Caching → 05 Orchestration, then loop to training.
- Strengthening fundamentals: start with 06 Coding & Design Patterns — it’s the vocabulary the other topics are written in.
- Building a GenAI / RAG product: 07 Retrieval & RAG + 01 §embeddings + 03 Streaming & SSE + 04 §caching, then 08 Authorization for secure/multi-tenant retrieval and 09 Evaluation to measure it.
- Cutting an inference bill: 09 Evaluation first, then 11 Model Routing & Cascades and 01 §small language models. Build the harness before you change which model answers.
See Study Plan for the AI Platform Engineering schedule.