Infrastructure
How distributed systems actually run: the containers and orchestration that schedule them, the messaging and coordination layers that move and order their data, the workers that process it, and the search and storage engines underneath — the operational substrate beneath every application.
Prerequisites: Some professional experience shipping software, plus Concurrency & Systems. Pairs with Cloud Native (orchestration depth), Batch & Streaming (the data-engineering view of the same Kafka), and Diagramming & Documentation (you cannot draw a deployment or a consumer topology honestly until you understand its operational caveats).
Why this track exists
Section titled “Why this track exists”Most engineers can write a service. Far fewer can answer what happens when it runs at scale, under failure, on shared infrastructure — where the state lives, how messages are delivered exactly once (or not), what bounds parallelism, which broker coordinates the cluster, and how a query reaches an inverted index. This track treats the runtime substrate as a first-class subject: the math of queues and delivery semantics, the architecture of the engines (Kubernetes, Kafka/KRaft/Redpanda, Lucene/Elasticsearch), and the caveats that cause real outages.
Topics 01-03 stay anchored to one running example — Streamflow, an event-driven order platform — so the containers you schedule in 01 are the consumer groups you scale in 03 over the log you model in 02.
Prerequisite Graph
Section titled “Prerequisite Graph”graph LR
K8S[01 Containers, Kubernetes & Workloads] --> MSG[02 Messaging & Distributed Queueing]
MSG --> WRK[03 Distributed Workers: merged into AI Platform 05]
K8S --> WRK
MSG --> SRCH[04 Search & Indexing]
SRCH --> APACHE[05 The Apache Stack]
WRK --> APACHE
Topics
Section titled “Topics”| # | Topic | Primary Reference | Time |
|---|---|---|---|
| 01 | Containers, Kubernetes & Workloads | Kubernetes docs (free) + 12-Factor App (free) | 1-2 weeks |
| 02 | Messaging & Distributed Queueing | Kafka docs (free) + DDIA Ch 11 | 1-2 weeks |
| 03 | Distributed Workers | merged into AI Platform 05 as Depth: partitioned consumers, with the Kafka consumers in its side quest | 1-2 weeks |
| 04 | Search & Indexing | Apache Lucene docs (free) + Elasticsearch: The Definitive Guide (free) | 1-2 weeks |
| 05 | The Apache Stack | Apache project docs (free) + DDIA | 1 week |
Quick Start
Section titled “Quick Start”- Asked to “add Kubernetes support”? Start with 01 — the reconciliation model and the caveats (limits/OOM, probes, SIGTERM, PDB) nobody warns you about.
- Designing an async architecture? 02 gives you the dynamics (queue vs log, push vs pull, delivery semantics, ZooKeeper vs KRaft vs Redpanda); 03 gives you the worker operations (rebalancing, idempotency, DLQ, lag-based scaling).
- Building search? 04 takes you from the inverted index and BM25 math down to Lucene/Elasticsearch/Solr — and cross-links the RAG side to Retrieval & RAG.
- Mapping the ecosystem? 05 is the bird’s-eye view of the Apache data stack — what each project is and which track goes deeper.
- Data/streaming role? 02-03 pair directly with Batch & Streaming.
See Study Plan for the schedule.