Skip to content

Data Engineering

Building the systems that move, store, and shape data so analysts, ML models, and applications can use it reliably. From storage formats and warehouses through batch/stream processing to orchestration and data quality.

Prerequisites: Algorithms, System Design basics, SQL

graph LR
    SQL[SQL + System Design] --> F[Foundations]
    F --> SW[Storage & Warehousing]
    F --> BS[Batch & Streaming]
    SW --> OM[Orchestration & Modeling]
    BS --> OM
    OM --> ML[ML Infrastructure / Serving]
    SW --> CP[05 Corpus Pipeline]
    OM --> CP
#TopicPrimary ReferenceTime
01FoundationsThe Data Engineering Cookbook (free) + Fundamentals of Data Engineering2-3 weeks
02Storage & WarehousingDDIA Ch 3 + The Data Warehouse Toolkit (Kimball)3-4 weeks
03Batch & Streaming ProcessingStreaming Systems + Spark / Kafka docs (free)3-4 weeks
04Orchestration & Modelingdbt docs (free) + Airflow docs (free)2-3 weeks
05Corpus PipelineFineWeb + datatrove (free)course passes 3, 8
  • Data engineering is a lifecycle: generation → ingestion → storage → transformation → serving, with governance, security, and orchestration as undercurrents throughout
  • The defining tradeoff is OLTP vs OLAP: row-oriented systems for transactions, column-oriented systems for analytics — storage layout follows access pattern
  • Modern stacks are ELT, not ETL: load raw data cheaply into a lake/warehouse, then transform in place with SQL — compute and storage are decoupled and elastic
  • Idempotency and exactly-once semantics are the hardest and most important properties of any pipeline; design for replay and failure from the start
  • The lakehouse (open table formats over object storage) is collapsing the old lake/warehouse divide

A small end-to-end Streamflow orders pipeline runs through this track, so the concepts have working code, not just prose:

StageFileTopic
Transform raw → clean (Spark, bronze→silver)spark/orders_bronze_to_silver.py03
Model clean → tested marts (dbt, silver→gold)dbt/04
Orchestrate the whole DAG (Airflow)airflow/streamflow_orders_dag.py04

It uses the same Streamflow event-driven platform as the Infrastructure and Diagramming tracks, and runs locally on DuckDB with no cloud setup.

  1. New to data: start with Foundations, then Storage & Warehousing
  2. Coming from backend/systems: skim Foundations, focus on Batch & Streaming
  3. Analytics/BI background: Storage & Warehousing → Orchestration & Modeling (dbt)
  4. ML/MLOps focus: Foundations + Batch & Streaming, then bridge to LLM Systems

See Study Plan for the 10-14 week data engineering schedule.