Data Engineering
Building the systems that move, store, and shape data so analysts, ML models, and applications can use it reliably. From storage formats and warehouses through batch/stream processing to orchestration and data quality.
Prerequisites: Algorithms, System Design basics, SQL
Prerequisite Graph
Section titled “Prerequisite Graph”graph LR
SQL[SQL + System Design] --> F[Foundations]
F --> SW[Storage & Warehousing]
F --> BS[Batch & Streaming]
SW --> OM[Orchestration & Modeling]
BS --> OM
OM --> ML[ML Infrastructure / Serving]
SW --> CP[05 Corpus Pipeline]
OM --> CP
Topics
Section titled “Topics”| # | Topic | Primary Reference | Time |
|---|---|---|---|
| 01 | Foundations | The Data Engineering Cookbook (free) + Fundamentals of Data Engineering | 2-3 weeks |
| 02 | Storage & Warehousing | DDIA Ch 3 + The Data Warehouse Toolkit (Kimball) | 3-4 weeks |
| 03 | Batch & Streaming Processing | Streaming Systems + Spark / Kafka docs (free) | 3-4 weeks |
| 04 | Orchestration & Modeling | dbt docs (free) + Airflow docs (free) | 2-3 weeks |
| 05 | Corpus Pipeline | FineWeb + datatrove (free) | course passes 3, 8 |
Key Takeaways
Section titled “Key Takeaways”- Data engineering is a lifecycle: generation → ingestion → storage → transformation → serving, with governance, security, and orchestration as undercurrents throughout
- The defining tradeoff is OLTP vs OLAP: row-oriented systems for transactions, column-oriented systems for analytics — storage layout follows access pattern
- Modern stacks are ELT, not ETL: load raw data cheaply into a lake/warehouse, then transform in place with SQL — compute and storage are decoupled and elastic
- Idempotency and exactly-once semantics are the hardest and most important properties of any pipeline; design for replay and failure from the start
- The lakehouse (open table formats over object storage) is collapsing the old lake/warehouse divide
Runnable Example
Section titled “Runnable Example”A small end-to-end Streamflow orders pipeline runs through this track, so the concepts have working code, not just prose:
| Stage | File | Topic |
|---|---|---|
| Transform raw → clean (Spark, bronze→silver) | spark/orders_bronze_to_silver.py | 03 |
| Model clean → tested marts (dbt, silver→gold) | dbt/ | 04 |
| Orchestrate the whole DAG (Airflow) | airflow/streamflow_orders_dag.py | 04 |
It uses the same Streamflow event-driven platform as the Infrastructure and Diagramming tracks, and runs locally on DuckDB with no cloud setup.
Quick Start
Section titled “Quick Start”- New to data: start with Foundations, then Storage & Warehousing
- Coming from backend/systems: skim Foundations, focus on Batch & Streaming
- Analytics/BI background: Storage & Warehousing → Orchestration & Modeling (dbt)
- ML/MLOps focus: Foundations + Batch & Streaming, then bridge to LLM Systems
See Study Plan for the 10-14 week data engineering schedule.