Skip to content

Foundations

  • Data engineering exists to make data usable downstream — it is plumbing in service of analytics and ML, not an end in itself
  • The data engineering lifecycle (generation → ingestion → storage → transformation → serving) is the mental map for the entire field
  • The single most important distinction is OLTP vs OLAP — it dictates storage layout, file format, and system choice
  • Choose components for the access pattern and SLA, not for novelty; most failures come from mismatched tools
  • Build a tiny end-to-end pipeline (ingest an API → land in object storage → transform with SQL → query) before going deep on any one stage
  • For every system you learn, classify it on the lifecycle and as OLTP or OLAP
  • Read DDIA Ch 1-2 for the vocabulary (reliability, scalability, maintainability, data models) that the rest of the track assumes

A data engineer’s job is to take data that was produced for one purpose (running an application, a sensor, a log) and make it reliably available for an entirely different purpose (analytics, ML, reporting). Every design decision flows from the gap between how data is generated and how it will be consumed: transactional vs analytical, real-time vs batch, raw vs modeled.

Key ideas:

  • Generation: source systems — application databases, APIs, event streams, IoT, files. You usually don’t control these
  • Ingestion: get data in — batch (scheduled pulls) or streaming (continuous). Push vs pull, CDC (change data capture) from databases
  • Storage: where it lands — object storage, lakes, warehouses. The substrate the other stages sit on
  • Transformation: cleaning, joining, aggregating, modeling — raw data into useful shapes
  • Serving: deliver to consumers — BI dashboards, ML features, reverse ETL back into apps
  • Undercurrents (cut across all stages): security, data management/governance, DataOps, data architecture, orchestration, software engineering

The foundational split

OLTPOLAP
PurposeRun the applicationAnalyze the business
AccessMany small reads/writesFew large scans/aggregates
LayoutRow-orientedColumn-oriented
LatencyMillisecondsSeconds to minutes
ExamplesPostgres, MySQL, DynamoDBSnowflake, BigQuery, Redshift, DuckDB

Why columnar wins for analytics: queries touch a few columns over many rows; storing columns together means reading only what you need, plus far better compression (similar values adjacent). This is the single most important idea in analytical storage.

Key ideas:

  • ETL (extract → transform → load): transform before loading, on a dedicated engine. The old default when storage and compute were expensive
  • ELT (extract → load → transform): land raw data first, transform in the warehouse with SQL. The modern default — cheap object storage + elastic compute make it cheaper to keep raw data and transform on demand
  • Why ELT won: decoupled storage/compute, raw data is replayable, transformation logic lives in version-controlled SQL (dbt), schema-on-read flexibility

Key ideas:

  • Row formats: Avro, JSON, CSV — good for write-heavy, record-at-a-time, schema evolution (Avro)
  • Columnar formats: Parquet, ORC — the analytics workhorses; column pruning, predicate pushdown, dictionary/run-length compression
  • Serialization: Protobuf/Avro for compact typed wire formats; Arrow for zero-copy in-memory columnar interchange between tools
  • Compression: Snappy (fast), Zstd (balanced), Gzip (small) — trade CPU for size

Key ideas:

  • Schema-on-write (warehouse): enforce structure at load time — safe, rigid
  • Schema-on-read (lake): store raw, interpret at query time — flexible, risky
  • Normalization (3NF) for OLTP to avoid update anomalies; denormalization (star schema) for OLAP to avoid join cost
  • Semi-structured: JSON/nested types are first-class in modern warehouses — handle them without flattening everything

From DDIA — shared with System Design

Key ideas:

  • CAP / PACELC: under partition, choose consistency or availability; even without partitions, latency vs consistency
  • Consistency models: strong, eventual, read-your-writes — pick per use case
  • Reliability, scalability, maintainability: the three properties DDIA argues every data system is judged on
  • Idempotency: re-running a step must not change the result — the property that makes retries and backfills safe

Key ideas:

  • Batch: bounded data, process on a schedule, high throughput, simple semantics. Default unless you need freshness
  • Streaming: unbounded data, process continuously, low latency, harder semantics (ordering, late data, exactly-once)
  • The honest rule: most “real-time” requirements are actually “fresh enough” — don’t pay the complexity tax of streaming unless the latency SLA truly demands it

ConceptWhat it isWhy it matters
LifecycleGeneration → serving + undercurrentsThe map of the whole field
OLTP/OLAPTransactional vs analyticalDictates storage and system choice
ELTLoad raw, transform in warehouseModern default, decoupled compute
ColumnarParquet/ORCAnalytics speed + compression
CDCCapture DB changes as a streamLow-impact ingestion from OLTP
IdempotencySafe re-executionMakes pipelines recoverable
ConceptConnected TrackApplication
CAP, replication, consistencySystem DesignShared distributed-systems core
Compression, entropyInformation TheoryWhy columnar compresses well
Feature pipelinesLLM SystemsFeeding data to models
CompanyHow This AppearsDifficulty
DatabricksLakehouse, lifecycle, Spark — this is their productExpert
AmazonRedshift, S3, Glue, data lake architectureAdvanced
NetflixMassive batch + streaming data platformExpert
PalantirData integration and modeling at scaleAdvanced
SnowflakeCloud warehouse, ELT, separation of storage/computeAdvanced