Skip to content

Corpus Pipeline

  • Data quality bounds model quality. Filtering, deduplication, and decontamination change a model’s loss and its benchmark numbers more than most architecture choices at small scale.
  • A pipeline is a composition of streaming stages over one Doc type: fetch, normalize, filter, exact dedup, near dedup, PII scrub, shard, tokenize. Each stage is a generator, so memory stays flat on any corpus size.
  • Deduplication has two halves: exact (hashes, confirmed by sort-merge) and near (MinHash signatures, LSH banding, union-find clusters). Decontamination is near-dedup against your eval sets.
  • Every output byte traces to a licensed source. The ledger records license and provenance per source; a shard whose source is unknown or not allowlisted fails the build.
  • Determinism is a feature: the same config gives the same output hash regardless of worker count, which is what makes a durable, resumable pipeline possible later.

Read the FineWeb paper end to end (it documents each filter’s measured effect), then skim datatrove’s pipeline stages. Work the MinHash and LSH chapter of Mining of Massive Datasets by hand for a 3-band example. In the course, build data.01 to data.08 in Pass 3 after the tokenizers, and data.09 in Pass 8 once the durable engine exists.


A training corpus is the output of a program, and like any program it has bugs, tests, and provenance. Treating the corpus as code means stages with typed interfaces, tests on hand-labelled fixtures, a manifest with content hashes, and a ledger that answers “where did this text come from and may we use it” for every shard.

Key ideas:

  • Async fetch with resume via HTTP Range, checksums, and quarantine on mismatch (data.01).
  • Stages are Callable[[Iterator[Doc]], Iterator[Doc]], composed with compose (data.02): Unicode normalization, language ID, Gopher rules, repetition filters, and a perplexity filter that plugs in the L2.1 n-gram model in C1.

Key ideas:

  • Exact: paragraph hashes, a Bloom screen (ds.08), then a sort-merge confirm so no true document is dropped (data.03).
  • Near: MinHash with 128 permutations, LSH with bb bands of rr rows (threshold about (1/b)1/r(1/b)^{1/r}), union-find clusters, process-parallel and worker-count invariant (data.04).
  • Decontamination: drop any document that shares a 13-gram with a protected eval or validation set, and record the count.

Key ideas:

  • PII scrub with typed placeholders and audit spans (data.05; policy in ethics.02).
  • Parquet shards and a manifest, with a document-hash train/val split (data.06).
  • Tokenize and pack to the llm.c .bin format read by L0.6 TokenStream (data.07).

4. Ledger, datasheet, and the durable pipeline

Section titled “4. Ledger, datasheet, and the durable pipeline”

Key ideas:

  • Ledger verification against formats/ledger.schema.json and a generated datasheet (data.08; policy in ethics.01).
  • CorpusBuild runs the stages as subprocess activities on the learner’s durable engine; a killed worker resumes and every shard is written once (data.09).
ModuleTopicKindPass
data.01Async fetch with resume, checksums, license capture (ledger rows)build3
data.02Extract, normalize, quality filters (generator stages)build3
data.03Exact dedup: paragraph hashes, Bloom screen, sort-merge confirmbuild3
data.04Near-dup MinHash + LSH + union-find, process-parallel; decontamination against protected eval and validation setsbuild3
data.05PII scrub with typed placeholders and audit spansbuild3
data.06Parquet shards, manifest, document-hash train/val split (uses pyarrow)build3
data.07Tokenize and pack to llm.c .binbuild3
data.08Licensing ledger verification and datasheetbuild3
data.09Pipeline as a durable workflow CorpusBuildbuild8

Milestone MS-corpus closes the pipeline: shards, manifest, ledger, and .bin files pass the format conformance suite with a deterministic output hash across two runs and worker counts 1 and 4.

#ModuleChapterKindPass
1data.01Async fetch with resume, checksums, license capturebuild3
2data.02Extract, normalize, quality filters (generator stages)build3
3data.03Exact dedup: paragraph hashes, Bloom screen, sort-merge confirmbuild3
4data.04Near-duplicate dedup: MinHash, LSH, union-find, and decontaminationbuild3
5data.05PII scrub with typed placeholders and audit spansbuild3
6data.06Parquet shards, manifest, and the document-hash splitbuild3
7data.07Tokenize and pack to llm.c .binbuild3
8data.08Licensing ledger verification and the datasheetbuild3
9data.09Pipeline as a durable workflow CorpusBuildbuild8
TrackConnection
Storage & WarehousingParquet and columnar layout behind data.06
Orchestration & ModelingAirflow and dbt as the going-further for data.09
Responsible AIdata licensing (ethics.01) and privacy (ethics.02)
Systems Data Structuresthe Bloom filter (ds.08)
tinyllm Part 1the tokenizer that data.07 runs
Durable Orchestration & Workersthe engine CorpusBuild runs on
CompanyPractice
Hugging FaceFineWeb and datatrove: documented, reproducible web-scale filtering
Allen Institute for AIDolma: an open corpus with its toolkit and datasheet
Every frontier labdecontamination against eval sets before reporting benchmark numbers