Skip to content

Infrastructure

How distributed systems actually run: the containers and orchestration that schedule them, the messaging and coordination layers that move and order their data, the workers that process it, and the search and storage engines underneath — the operational substrate beneath every application.

Prerequisites: Some professional experience shipping software, plus Concurrency & Systems. Pairs with Cloud Native (orchestration depth), Batch & Streaming (the data-engineering view of the same Kafka), and Diagramming & Documentation (you cannot draw a deployment or a consumer topology honestly until you understand its operational caveats).

Most engineers can write a service. Far fewer can answer what happens when it runs at scale, under failure, on shared infrastructure — where the state lives, how messages are delivered exactly once (or not), what bounds parallelism, which broker coordinates the cluster, and how a query reaches an inverted index. This track treats the runtime substrate as a first-class subject: the math of queues and delivery semantics, the architecture of the engines (Kubernetes, Kafka/KRaft/Redpanda, Lucene/Elasticsearch), and the caveats that cause real outages.

Topics 01-03 stay anchored to one running example — Streamflow, an event-driven order platform — so the containers you schedule in 01 are the consumer groups you scale in 03 over the log you model in 02.

graph LR
    K8S[01 Containers, Kubernetes & Workloads] --> MSG[02 Messaging & Distributed Queueing]
    MSG --> WRK[03 Distributed Workers: merged into AI Platform 05]
    K8S --> WRK
    MSG --> SRCH[04 Search & Indexing]
    SRCH --> APACHE[05 The Apache Stack]
    WRK --> APACHE
#TopicPrimary ReferenceTime
01Containers, Kubernetes & WorkloadsKubernetes docs (free) + 12-Factor App (free)1-2 weeks
02Messaging & Distributed QueueingKafka docs (free) + DDIA Ch 111-2 weeks
03Distributed Workersmerged into AI Platform 05 as Depth: partitioned consumers, with the Kafka consumers in its side quest1-2 weeks
04Search & IndexingApache Lucene docs (free) + Elasticsearch: The Definitive Guide (free)1-2 weeks
05The Apache StackApache project docs (free) + DDIA1 week
  1. Asked to “add Kubernetes support”? Start with 01 — the reconciliation model and the caveats (limits/OOM, probes, SIGTERM, PDB) nobody warns you about.
  2. Designing an async architecture? 02 gives you the dynamics (queue vs log, push vs pull, delivery semantics, ZooKeeper vs KRaft vs Redpanda); 03 gives you the worker operations (rebalancing, idempotency, DLQ, lag-based scaling).
  3. Building search? 04 takes you from the inverted index and BM25 math down to Lucene/Elasticsearch/Solr — and cross-links the RAG side to Retrieval & RAG.
  4. Mapping the ecosystem? 05 is the bird’s-eye view of the Apache data stack — what each project is and which track goes deeper.
  5. Data/streaming role? 02-03 pair directly with Batch & Streaming.

See Study Plan for the schedule.