Skip to content

AI Platform Engineering

The infrastructure layer underneath modern AI products — but taught as patterns, not products. Frameworks, protocols, caches, and orchestrators churn constantly; the patterns underneath them barely move. Master the handful of fundamental patterns and every new tool becomes a recognizable instance of something you already understand.

The math analogy: once you understand limits, derivatives, and integrals, every named technique downstream is a special case. The same is true here. Cache-aside, consistent hashing, idempotent retry, event-sourced replay, the worker pool, backpressure, the decorator — learn the pattern and how it works, and “what’s trending this quarter” becomes derivative. We name the trendy tools (vLLM, Temporal, Redis, gRPC) as instances of patterns, and we note how companies apply them, but the pattern is the lesson.

Prerequisites: Deep Learning (transformers, training loops), LLM Systems & Inference (serving, KV cache, parallelism), System Design, Cloud Native, Data Engineering (OLTP vs OLAP, storage), Functional Programming. Comfortable with Concurrency & Systems.

graph LR
    DL[Deep Learning] --> TF[Training & Frameworks]
    LLM[LLM Systems] --> TF
    CN[Cloud Native] --> TF
    SD[System Design] --> RPC[RPC & Protocols]
    RPC --> SSE[Streaming & SSE]
    LLM --> SSE
    SD --> DDC[Distributed Data & Caching]
    DE[Data Engineering] --> DDC
    DDC --> ORC[Orchestration & Workers]
    TF --> ORC
    RPC --> ORC
    FP[Functional Programming] --> CDP[Coding & Design Patterns]
    CDP --> ORC
    CDP --> TF
    TF --> RAG[Retrieval & RAG]
    DDC --> RAG
    RAG --> AUTH[Authorization & Access Control]
    TF --> EVAL[LLM Evaluation]
    RAG --> EVAL
    TF --> EDGE[Edge, Realtime & On-Device]
    SSE --> EDGE
    EVAL --> ROUTE[Model Routing & Cascades]
    TF --> ROUTE
    LLM --> ROUTE
    SSE --> GWY[Gateway]
    ROUTE --> GWY
    GWY --> AGT[Agent SDK]
    ORC --> AGT
    RAG --> AGT

The Split: Three Different “Distributed” Problems

Section titled “The Split: Three Different “Distributed” Problems”

A recurring confusion is that “distributed X” is one topic. It is three, and this track keeps them separate because the patterns differ:

ProblemLives inThe pattern family
Distributed training — one model across many GPUs01 Training & FrameworksParallelism & sharding (data/tensor/pipeline, FSDP/ZeRO)
Distributed data — one dataset across many nodes04 Distributed Data & CachingPartitioning, replication, caching, invalidation
Distributed orchestration — one workflow across many failures05 Orchestration & WorkersDurable execution, checkpoints, the worker pool
#TopicPattern lensPrimary ReferenceTime
01Training & FrameworksParallelism, sharding, adaptersPyTorch + JAX + vLLM docs4-6 weeks
02RPC & ProtocolsContracts, serialization, evolutiongRPC + Protobuf docs2-3 weeks
03Streaming & SSEPush, backpressure, cancellationHTML SSE spec + MDN1-2 weeks
04Distributed Data & CachingPartition, replicate, cache, invalidateRedis + Cassandra docs + DDIA3-4 weeks
05Durable Orchestration & WorkersDurable execution, workers, observability, profilingTemporal + DBOS docs2-3 weeks
06Coding & Design PatternsDecorators, facade, closures, curryingRefactoring Guru + Mostly Adequate Guide2-3 weeks
07Retrieval & RAGEncoders, chunking, hybrid fusion, multimodalRAG paper + BM25 + pgvector3-4 weeks
08Authorization & Access ControlRBAC/ABAC/ReBAC/NGAC, the pushdown ladderNIST RBAC/ABAC/NGAC + Zanzibar2-3 weeks
09LLM EvaluationCross-entropy/perplexity/bits-per-byte, judgesMacKay + HELM + Ragas2-3 weeks
10Edge, Realtime & On-Device InferenceDeployment spectrum, streaming encoders, efficiency architecturesllama.cpp + Mistral 7B + Mamba2-3 weeks
11Model Routing & CascadesOracle vs router, cascades, escalation, cache-aware switching, decision modelsRouteLLM + FrugalGPT + LLMRouterBench1-2 weeks
12GatewayThe reverse proxy whose payload is a stream: keys, trace propagation, then limits, routing, caching (course)OpenAI API reference + W3C Trace Contextcourse passes 1, 7, 10
13Agent SDKThe model proposes, the program disposes: types, tools, the loop, gates, durable runs (course)Building effective agents + ReAct (free)course pass 10
  • Learn the pattern, not the product. vLLM is continuous batching + paged cache; Temporal is event-sourced durable execution + worker pool; Redis is an in-memory hash table you must invalidate. Name the pattern and the tool is interchangeable.
  • “Distributed” is three problems, not one — training (parallelism), data (partition/replicate/cache), orchestration (durable execution). Different failure modes, different patterns; this track keeps them apart on purpose.
  • Caching is the second-hardest problem in CS, and invalidation is why. A cache is a lie you tell for speed; every caching pattern (cache-aside, write-through, write-back, TTL) is a different answer to “when does the lie expire?”
  • Durable execution is checkpointing for workflows. Temporal persists every step (the checkpoint), replays history to recover, runs your code on a pool of workers, and gives you distributed observability over long-running processes for free.
  • Profiling beats guessing, always. Whether it’s a GPU kernel, a slow query, a cache miss rate, or a workflow’s tail latency — measure first. The bottleneck is rarely where intuition says.
  • Design patterns are how these systems are built. Decorators wrap workflows and retries; the facade hides a serving stack behind one call; closures and currying are how JAX, middleware, and configuration actually work. The patterns in topic 06 recur in every topic above.
  • RAG is information retrieval wearing a neural coat. Retrieve-then-generate; hybrid (dense + BM25 + graph) fused with RRF; bi-encode to find and cross-encode to rank. Retrieval quality — not the model — is the usual ceiling.
  • Authorization is a complexity-ladder problem. Flat RBAC is a lookup; nested groups and hierarchies are graph reachability (pushdown, not finite-state) — which is why ReBAC/Zanzibar exist, and why secure RAG filters by permission during retrieval.
  • Evaluation starts in information theory. Cross-entropy = bits to predict the next token = compression; perplexity and bits-per-byte are re-normalizations. Intrinsic loss steers training; extrinsic evals (benchmarks, judges, RAG faithfulness) decide if it’s good.
  • Inference is a spectrum from datacenter to phone, and “realtime” is an architecture choice. llama.cpp + GGUF is the end-to-end local path; streaming encoders must be causal/chunked; and Mistral’s SWA, GQA, MoE, and Mamba/SSM variants are four different escapes from the vanilla transformer’s cost curve.
  • Routing is approximating an oracle you cannot have. No model is best at everything, so a pool beats its best member, but only on paper. A real router chooses before the outcome exists, has to price the cache it abandons when it switches, and should escalate on category or a verifier, not on model confidence alone.
  1. Coming from ML / research: 01 Training & Frameworks → 04 Data & Caching → 05 Orchestration — see the platform around your model.
  2. Coming from backend / distributed systems: 02 RPC → 04 Data & Caching → 05 Orchestration, then loop to training.
  3. Strengthening fundamentals: start with 06 Coding & Design Patterns — it’s the vocabulary the other topics are written in.
  4. Building a GenAI / RAG product: 07 Retrieval & RAG + 01 §embeddings + 03 Streaming & SSE + 04 §caching, then 08 Authorization for secure/multi-tenant retrieval and 09 Evaluation to measure it.
  5. Cutting an inference bill: 09 Evaluation first, then 11 Model Routing & Cascades and 01 §small language models. Build the harness before you change which model answers.

See Study Plan for the AI Platform Engineering schedule.