Skip to content

The Apache Stack

  • The ASF is a de-facto standard library for distributed data. When a hard distributed-systems problem becomes common enough, an Apache project usually crystallizes as the open, vendor-neutral reference implementation — and “Apache X” becomes a quality and longevity signal, not just a name.
  • Much of the stack is open re-implementation of Google papers: GFS→HDFS, Bigtable→HBase, MapReduce→Hadoop, Dataflow→Beam/Flink. Knowing the paper tells you the architecture before you read a line of docs.
  • Pick the project by the problem, not the brand. Coordination, messaging, storage, OLAP, search, orchestration, and graph are seven distinct problem families; this topic groups the stack by which one each project owns.
  • The frontier is moving toward separation and interchange: storage/compute split (Pulsar’s BookKeeper, the lakehouse), the death of a separate coordinator (Kafka’s KRaft), open table formats (Iceberg’s rise), and zero-copy columnar interchange (Arrow/ADBC). Track these — they are where the coupled architectures of the 2010s are being unbundled.
  • This is the bird’s-eye map. Every box here is covered in depth by a sibling topic; the value of this page is knowing which box to reach for and why, not re-deriving the internals.
  • For each category below, name the one problem it solves out loud before reading the projects. If you can’t, you’ll collect trivia instead of a mental index.
  • Place Streamflow (our running ingest-and-analytics platform) onto the map: which Apache project sits at each stage, and what would you swap if a requirement changed (sub-second latency, multi-region writes, ad-hoc OLAP)?
  • When two projects look interchangeable (Kafka vs Pulsar, Flink vs Spark, Iceberg vs Hudi), find the one architectural decision that separates them — that difference is the whole reason both exist.
  • Follow each cross-link to the deep topic; do not try to learn Kafka or Iceberg from this page.

The Apache Software Foundation is best understood as a de-facto standard library for distributed data: a place where the recurring hard problems of large-scale systems — coordinate this cluster, move this firehose of events, store this petabyte cheaply, answer this aggregation in 50ms, search this corpus, schedule this DAG — each get an open, governed, vendor-neutral reference implementation. The ASF doesn’t write the code; it provides the governance model (meritocracy, the “Apache Way,” a foundation that holds the trademark and IP) that lets a project outlive any single company. That is why “Apache X” is a signal: it means a community, not a vendor, owns the roadmap.

The right way to hold the whole stack in your head is to group projects by the problem they solve, not by their age or popularity:

graph TD
    subgraph COORD["Coordination & metadata"]
        ZK[ZooKeeper -- ZAB]
        KRAFT[Kafka KRaft -- self-managed]
    end
    subgraph MSG["Messaging & streaming"]
        KAFKA[Kafka]
        PULSAR[Pulsar + BookKeeper]
        FLINK[Flink]
        SPARK[Spark]
        BEAM[Beam]
    end
    subgraph STORE["Storage & tables"]
        HDFS[HDFS]
        CASS[Cassandra]
        HBASE[HBase]
        ICE[Iceberg / Hudi]
        PARQ[Parquet / ORC]
    end
    subgraph OLAP["OLAP / real-time analytics"]
        DRUID[Druid]
        PINOT[Pinot]
        KYLIN[Kylin]
    end
    subgraph SEARCH["Search"]
        LUCENE[Lucene]
        SOLR[Solr]
    end
    subgraph ORCH["Orchestration & interchange"]
        AIRFLOW[Airflow]
        ARROW[Arrow]
    end
    subgraph GRAPH["Graph"]
        AGE[Apache AGE]
    end
    MSG --> STORE
    STORE --> OLAP
    STORE --> SEARCH
    ORCH -.schedules/moves.-> MSG
    ORCH -.schedules/moves.-> STORE
    COORD -.coordinates.-> MSG
    COORD -.coordinates.-> STORE

Read this map as a pipeline: events arrive through messaging, land in storage (raw files, wide-column rows, or open tables), get served to OLAP engines and search indexes, all coordinated and orchestrated, with graph as a specialized lens when relationships are the data. Our running example, Streamflow, threads the whole map: clickstream events → Kafka → Flink (real-time aggregates) and Spark (nightly batch) → Iceberg tables on object storage → Druid for sub-second dashboards and Solr for the product search box, all scheduled by Airflow.

The unglamorous foundation: who is the leader, which broker owns which partition, what is the current cluster config. This is a consensus problem, and getting it wrong corrupts everything above it.

ProjectWhat it ISThe one problemArchitecture ideaGo deeper
ZooKeeperA replicated, hierarchical key-value store (znodes) for coordinationReliable cluster metadata, leader election, locksZAB (ZooKeeper Atomic Broadcast): a leader-based total-order broadcast protocol, Paxos-adjacent; strongly consistent small-data storeMessaging & Queueing
Kafka KRaftKafka’s built-in Raft quorum that replaces ZooKeeperRemove the external coordinator dependencyMetadata becomes an internal Raft-replicated log on the brokers themselves — one fewer system to operate, faster failover, far higher partition countsMessaging & Queueing

The move away from ZooKeeper is the story here. For a decade, ZooKeeper was the coordinator under Kafka, HBase, Solr, and more. Operating a separate ZK ensemble is real toil, and its consistency model (small data, watch-based) doesn’t scale to millions of partitions. Kafka’s KRaft mode (production-default since Kafka 3.x, ZooKeeper removed in 4.0) folds coordination into the brokers via Raft. The non-Apache contrast is etcd — also Raft-based, also a consistent metadata store, but it became the coordination substrate of the cloud-native world (it is the backing store for Kubernetes; see Containers & Kubernetes). Same problem, two lineages: ZooKeeper from the Hadoop era, etcd from the Kubernetes era.

Move an unbounded stream of events between systems, then compute over it. This is where the most active design tension in the stack lives.

ProjectWhat it ISThe one problemArchitecture ideaGo deeper
KafkaA distributed, partitioned, replicated commit logDurable, replayable decoupling of producers from consumersAppend-only partitioned log; ordering within a partition; consumers read at their own offset. Storage and serving are coupled in the brokerDE: Batch & Streaming
PulsarA pub-sub messaging + streaming systemSame decoupling, but with elastic, independently-scaled storageCompute/storage separation: stateless brokers serve; Apache BookKeeper (a separate log-storage service of “bookies”) holds the durable ledgers. Scale serving and storage independently; built-in tiered storage and multi-tenancyDE: Batch & Streaming
FlinkA true streaming, stateful processing engineLow-latency, exactly-once stateful computation over unbounded streamsEvent-at-a-time processing; keyed managed state; Chandy-Lamport aligned checkpoints for exactly-once recovery tied to source offsetsDE: Batch & Streaming
SparkA general distributed compute engine (batch + micro-batch streaming)Optimized batch and “good enough” streaming on one engineLazy DAG of transformations; Catalyst optimizer + Tungsten codegen; Structured Streaming treats a stream as an unbounded table (micro-batch by default)DE: Batch & Streaming
BeamA unified programming model (not an engine)Write one pipeline, run it on many enginesThe Dataflow model (what/where/when/how) as a portable API; runners execute it on Flink, Spark, or Google DataflowDE: Batch & Streaming
Storm / SamzaFirst-generation stream processors (legacy)Early real-time stream processingStorm: at-least-once tuple processing; Samza: Kafka-native, local state. Mostly superseded by Flink; you’ll meet them in old systemsDE: Batch & Streaming

The defining contrasts. Kafka vs Pulsar: Kafka couples log storage to the broker, so adding storage means adding (and rebalancing) brokers; Pulsar splits them via BookKeeper, so you scale serving and storage on separate axes — at the cost of operating more moving parts. Kafka answers back with tiered storage (offloading cold log segments to object storage), narrowing the gap. Flink vs Spark: Flink is streaming-first (sub-second, heavy stateful logic, CEP); Spark is batch-first with micro-batch streaming bolted on cleanly — reach for Flink when latency and state dominate, Spark when you want one engine for big batch plus simpler streaming. Beam sits above both: a portability layer for teams that refuse to marry an engine.

Where the bytes actually live, and how files become tables with guarantees.

ProjectWhat it ISThe one problemArchitecture ideaGo deeper
Hadoop / HDFSA distributed filesystem (the historical foundation)Store petabytes across commodity disksGFS re-implementation: NameNode (metadata) + DataNodes (blocks), 3x replication. Largely displaced by cloud object storage (S3/GCS), but the conceptual ancestor of the whole lakeDE: Storage & Warehousing
CassandraA masterless wide-column storeAlways-on, write-heavy, multi-region writesDynamo lineage: consistent-hash ring, no SPOF, tunable quorum consistency; LSM-tree storage. Query-first modeling, partition key = shard keyAIPE: Distributed Data & Caching
HBaseA distributed sorted map on HDFSRandom read/write over huge sparse tablesBigtable re-implementation: sorted rows split into regions, LSM (memstore + HFiles); strong per-row consistency. Reach for it when you need Bigtable semantics on-premDE: Storage & Warehousing
IcebergAn open table format over data filesACID, time travel, schema/partition evolution on object storageSnapshot + manifest design: each commit writes a new immutable metadata snapshot pointing at manifest lists → manifests → data files. Hidden partitioning decouples layout from queries. The emerging open standardDE: Storage & Warehousing
HudiAn open table format specialized for upsertsFast incremental upserts + CDC on the lakeCopy-on-write vs merge-on-read tables; record-level indexes for mutate-heavy ingestDE: Storage & Warehousing
Delta LakeA table format (Linux Foundation, not Apache)ACID transactions via a write-ahead logTransaction log (_delta_log) of ordered JSON commits; Databricks-native origin. Listed here because it competes head-on with Iceberg/HudiDE: Storage & Warehousing
Parquet / ORCColumnar file formats (the bytes on disk)Compact, scan-efficient analytical storageColumnar layout + per-column encoding/compression + row-group statistics (min/max, zone maps) for predicate pushdown. The substrate Iceberg/Hudi/Delta organizeDE: Storage & Warehousing

The layering that confuses everyone: Parquet/ORC are file formats (how one file is encoded); Iceberg/Hudi/Delta are table formats (a metadata layer that turns a directory of those files into a transactional table). Iceberg’s rise is the headline — its engine-agnostic snapshot/manifest design (and its REST catalog) made it the format every major engine and cloud vendor now supports, and the point where the lake/warehouse divide finally collapses into the lakehouse.

Sub-second aggregations over fresh data — the gap between a streaming engine and a slow warehouse.

ProjectWhat it ISThe one problemArchitecture ideaGo deeper
DruidA real-time analytics databaseLow-latency slice-and-dice over event streamsSegment-based columnar storage, time-partitioned; ingests from Kafka and serves queries; pre-aggregation at ingest. Powers live dashboardsDE: Storage & Warehousing
PinotA real-time OLAP datastoreUltra-low-latency, high-QPS user-facing analyticsColumnar segments with rich indexing (inverted, star-tree, range); built for serving analytics to end users at scale (e.g. “who viewed your profile”)DE: Storage & Warehousing
KylinA distributed OLAP engineSub-second queries on huge dimensional dataPre-computed OLAP cubes over Hadoop/Hive; trades storage + build time for instant cube lookups. More legacy/batch-cube oriented than Druid/PinotDE: Storage & Warehousing

When to reach here: when your dashboard query is too latency-sensitive or too high-QPS for a warehouse, but too analytical (group-by over millions of rows) for an OLTP store. Druid and Pinot both ingest directly from Kafka, closing the loop from §2 — Streamflow’s live dashboards sit here.

Full-text relevance ranking — a different problem from both OLAP (aggregation) and key-value (point lookup).

ProjectWhat it ISThe one problemArchitecture ideaGo deeper
LuceneA full-text search library (the engine core)Index a corpus and rank documents by relevanceThe inverted index (term → posting list of docs), plus tokenization/analysis and scoring (BM25). Not a server — a JAR you embedSearch & Indexing
SolrA search server built on LuceneOperate Lucene as a distributed, queryable serviceAdds sharding, replication, a REST/HTTP API, and faceting over Lucene. (Elasticsearch/OpenSearch are the other Lucene-based servers)Search & Indexing

The key relationship: Lucene is the engine; Solr (and Elasticsearch/OpenSearch) are servers wrapped around it. This is why “search” interview questions about inverted indexes, analyzers, and BM25 are really Lucene questions regardless of which server runs in production. Streamflow’s product search box is Solr in front of a Lucene index.

Schedule the pipeline, and move columnar data between systems without re-serializing it.

ProjectWhat it ISThe one problemArchitecture ideaGo deeper
AirflowA workflow schedulerExpress and run dependency-ordered batch pipelinesDAGs as Python code; a scheduler triggers tasks when upstreams succeed; rich retries/backfills. The default batch orchestratorDE: Orchestration & Modeling
BeamUnified pipeline model (see §2)Engine-portable data processingCross-listed: Beam is both a streaming and an orchestration-adjacent abstractionDE: Batch & Streaming
ArrowAn in-memory columnar interchange formatMove/share columnar data with zero copies and no serializationA standardized in-memory columnar layout so Spark, pandas, DuckDB, and Parquet readers share buffers directly. ADBC is the emerging Arrow-native DB connectivity standard (a columnar JDBC/ODBC)DE: Storage & Warehousing

Why Arrow matters more than it looks: most data-system overhead is serializing/deserializing between processes. Arrow defines one in-memory columnar layout everyone agrees on, so handing a table from Spark to pandas to DuckDB becomes a pointer pass, not a re-encode. Arrow Flight (and Flight SQL / ADBC) extends this to the wire. This is the quiet, frontier-relevant glue under the modern stack.

When the relationships are the data and traversals dominate.

ProjectWhat it ISThe one problemArchitecture ideaGo deeper
Apache AGEA Postgres extension for graphsRun graph (Cypher) queries alongside relational tablesopenCypher inside Postgres — one database for relational + (with pgvector) vector + graph, instead of bolting on a separate Neo4jAIPE: Distributed Data & Caching

When to reach here: multi-hop traversals (“friends-of-friends,” dependency chains, GraphRAG subgraphs) that would be self-join hell in plain SQL — but you’d rather extend the Postgres you already run than operate a dedicated graph database.

Lineage & Governance (the thread tying it together)

Section titled “Lineage & Governance (the thread tying it together)”

Two things explain why this stack looks the way it does:

  • Open re-implementation of Google papers. A striking share of the foundation is the open-source world rebuilding what Google published: GFS → HDFS, Bigtable → HBase (and conceptually Cassandra, blending Bigtable’s data model with Dynamo’s distribution), MapReduce → Hadoop, Dataflow → Beam/Flink. Read the paper and you have the architecture for free; the Apache project is the community-maintained implementation.
  • The ASF governance model. The foundation holds the trademark and IP, enforces the “Apache Way” (public, merit-based, consensus-driven development), and requires a diverse community before a project graduates from the Incubator to top-level. The payoff: vendor neutrality and longevity. “Apache X” signals that no single company can unilaterally relicense or abandon it — which is exactly why enterprises standardize on it. (Contrast the rug-pulls in the BSL/SSPL world; the ASF model is the structural answer to that risk.)
Your problemReach forNot
Coordinate a cluster / store cluster metadataKRaft (new), etcd (cloud-native), ZooKeeper (legacy)A general DB
Durable, replayable event backboneKafka (default)A traditional message queue
Independently scale messaging storage vs serving, multi-tenantPulsar (BookKeeper split)Kafka, if ops simplicity matters more
Sub-second stateful streaming, exactly-once, CEPFlinkSpark Structured Streaming
Big batch + “good enough” micro-batch on one engineSparkFlink
Engine-portable pipeline codeBeamBinding to one engine
ACID tables / time travel on cheap object storageIceberg (open standard); Hudi (upsert-heavy); Delta (Databricks)A raw directory of Parquet
Always-on, write-heavy, multi-region writesCassandraA single-leader RDBMS
Bigtable-style sorted random access on-premHBaseA relational store
Sub-second, high-QPS user-facing analyticsPinot (or Druid)A cloud warehouse
Live operational dashboards over event streamsDruidBatch warehouse queries
Full-text relevance searchSolr / Elasticsearch (Lucene under both)A LIKE '%...%' query
Schedule dependency-ordered batch jobsAirflowCron + glue scripts
Zero-copy columnar data interchangeArrow (+ ADBC)Re-serializing between every system
Multi-hop graph queries beside relational dataApache AGEA separate graph DB you must operate
ConceptConnected TrackHow
Kafka/Pulsar internals, KRaft, ZABMessaging & QueueingThe log and its coordinator, in depth
Flink/Spark/Beam, exactly-once, watermarksDE: Batch & StreamingThe processing engines, in depth
HDFS, Cassandra/HBase, Iceberg/Hudi/Parquet, Druid/Pinot, ArrowDE: Storage & WarehousingStorage engines and table/file formats
Airflow DAGs, dbt, modelingDE: Orchestration & ModelingThe scheduler and the medallion pipeline
Lucene inverted index, Solr, BM25Search & IndexingFull-text relevance internals
Cassandra, Apache AGE, etcd-style coordinationAIPE: Distributed Data & CachingWide-column, graph-on-Postgres
etcd as the K8s coordinatorContainers & KubernetesThe cloud-native contrast to ZooKeeper
Consensus, replication, CAP, Dynamo/BigtableSystem DesignThe theory under every box on the map
CompanyWhere the stack shows upFocus
LinkedInCreated Kafka, Samza, Pinot; heavy LuceneEvent backbone + user-facing OLAP
NetflixCreated Iceberg; massive Kafka + FlinkLakehouse tables, streaming at scale
Confluent / StreamNativeKafka / Pulsar as productsKRaft, tiered storage, BookKeeper
DatabricksSpark (created here), Delta, IcebergLakehouse, structured streaming
UberCreated Hudi; Flink, PinotUpsert-heavy lake, real-time analytics
Apple / BloombergLarge Cassandra and Solr fleetsAlways-on storage, enterprise search
Any data-platform / staff infra roleKnowing which box and whyChoosing and operating the stack