Corpus Pipeline
Overview
Section titled “Overview”- Primary references: Penedo et al., The FineWeb Datasets (free); Soldaini et al., Dolma (free); Hugging Face, datatrove (free, Apache-2.0)
- Supplementary: Lee et al., Deduplicating Training Data Makes Language Models Better (free); Broder, On the Resemblance and Containment of Documents (1997, MinHash); Rae et al., Scaling Language Models: Methods, Analysis & Insights from Training Gopher (the Gopher quality rules, free); Leskovec, Rajaraman, and Ullman, Mining of Massive Datasets chapter 3 (free)
- Prerequisites: Data Engineering Foundations, Storage & Warehousing (columnar files), the asyncio primer
lang.08, the responsible-AI topics on licensing and PII - Estimated time: 3 to 4 weeks in course Pass 3, plus one module in Pass 8 when the pipeline becomes a durable workflow
Key Takeaways
Section titled “Key Takeaways”- Data quality bounds model quality. Filtering, deduplication, and decontamination change a model’s loss and its benchmark numbers more than most architecture choices at small scale.
- A pipeline is a composition of streaming stages over one
Doctype: fetch, normalize, filter, exact dedup, near dedup, PII scrub, shard, tokenize. Each stage is a generator, so memory stays flat on any corpus size. - Deduplication has two halves: exact (hashes, confirmed by sort-merge) and near (MinHash signatures, LSH banding, union-find clusters). Decontamination is near-dedup against your eval sets.
- Every output byte traces to a licensed source. The ledger records license and provenance per source; a shard whose source is unknown or not allowlisted fails the build.
- Determinism is a feature: the same config gives the same output hash regardless of worker count, which is what makes a durable, resumable pipeline possible later.
How to Study
Section titled “How to Study”Read the FineWeb paper end to end (it documents each filter’s measured effect), then skim datatrove’s pipeline stages. Work the MinHash and LSH chapter of Mining of Massive Datasets by hand for a 3-band example. In the course, build data.01 to data.08 in Pass 3 after the tokenizers, and data.09 in Pass 8 once the durable engine exists.
Concepts & Techniques
Section titled “Concepts & Techniques”Core Insight
Section titled “Core Insight”A training corpus is the output of a program, and like any program it has bugs, tests, and provenance. Treating the corpus as code means stages with typed interfaces, tests on hand-labelled fixtures, a manifest with content hashes, and a ledger that answers “where did this text come from and may we use it” for every shard.
1. Fetch and stages
Section titled “1. Fetch and stages”Key ideas:
- Async fetch with resume via HTTP
Range, checksums, and quarantine on mismatch (data.01). - Stages are
Callable[[Iterator[Doc]], Iterator[Doc]], composed withcompose(data.02): Unicode normalization, language ID, Gopher rules, repetition filters, and a perplexity filter that plugs in theL2.1n-gram model in C1.
2. Deduplication and decontamination
Section titled “2. Deduplication and decontamination”Key ideas:
- Exact: paragraph hashes, a Bloom screen (
ds.08), then a sort-merge confirm so no true document is dropped (data.03). - Near: MinHash with 128 permutations, LSH with bands of rows (threshold about ), union-find clusters, process-parallel and worker-count invariant (
data.04). - Decontamination: drop any document that shares a 13-gram with a protected eval or validation set, and record the count.
3. Privacy, shards, and tokens
Section titled “3. Privacy, shards, and tokens”Key ideas:
- PII scrub with typed placeholders and audit spans (
data.05; policy inethics.02). - Parquet shards and a manifest, with a document-hash train/val split (
data.06). - Tokenize and pack to the llm.c
.binformat read byL0.6TokenStream(data.07).
4. Ledger, datasheet, and the durable pipeline
Section titled “4. Ledger, datasheet, and the durable pipeline”Key ideas:
- Ledger verification against
formats/ledger.schema.jsonand a generated datasheet (data.08; policy inethics.01). CorpusBuildruns the stages as subprocess activities on the learner’s durable engine; a killed worker resumes and every shard is written once (data.09).
Course modules
Section titled “Course modules”| Module | Topic | Kind | Pass |
|---|---|---|---|
data.01 | Async fetch with resume, checksums, license capture (ledger rows) | build | 3 |
data.02 | Extract, normalize, quality filters (generator stages) | build | 3 |
data.03 | Exact dedup: paragraph hashes, Bloom screen, sort-merge confirm | build | 3 |
data.04 | Near-dup MinHash + LSH + union-find, process-parallel; decontamination against protected eval and validation sets | build | 3 |
data.05 | PII scrub with typed placeholders and audit spans | build | 3 |
data.06 | Parquet shards, manifest, document-hash train/val split (uses pyarrow) | build | 3 |
data.07 | Tokenize and pack to llm.c .bin | build | 3 |
data.08 | Licensing ledger verification and datasheet | build | 3 |
data.09 | Pipeline as a durable workflow CorpusBuild | build | 8 |
Milestone MS-corpus closes the pipeline: shards, manifest, ledger, and .bin files pass the format conformance suite with a deterministic output hash across two runs and worker counts 1 and 4.
Chapters
Section titled “Chapters”| # | Module | Chapter | Kind | Pass |
|---|---|---|---|---|
| 1 | data.01 | Async fetch with resume, checksums, license capture | build | 3 |
| 2 | data.02 | Extract, normalize, quality filters (generator stages) | build | 3 |
| 3 | data.03 | Exact dedup: paragraph hashes, Bloom screen, sort-merge confirm | build | 3 |
| 4 | data.04 | Near-duplicate dedup: MinHash, LSH, union-find, and decontamination | build | 3 |
| 5 | data.05 | PII scrub with typed placeholders and audit spans | build | 3 |
| 6 | data.06 | Parquet shards, manifest, and the document-hash split | build | 3 |
| 7 | data.07 | Tokenize and pack to llm.c .bin | build | 3 |
| 8 | data.08 | Licensing ledger verification and the datasheet | build | 3 |
| 9 | data.09 | Pipeline as a durable workflow CorpusBuild | build | 8 |
Connections to Other Tracks
Section titled “Connections to Other Tracks”| Track | Connection |
|---|---|
| Storage & Warehousing | Parquet and columnar layout behind data.06 |
| Orchestration & Modeling | Airflow and dbt as the going-further for data.09 |
| Responsible AI | data licensing (ethics.01) and privacy (ethics.02) |
| Systems Data Structures | the Bloom filter (ds.08) |
| tinyllm Part 1 | the tokenizer that data.07 runs |
| Durable Orchestration & Workers | the engine CorpusBuild runs on |
Company Relevance
Section titled “Company Relevance”| Company | Practice |
|---|---|
| Hugging Face | FineWeb and datatrove: documented, reproducible web-scale filtering |
| Allen Institute for AI | Dolma: an open corpus with its toolkit and datasheet |
| Every frontier lab | decontamination against eval sets before reporting benchmark numbers |