Skip to content

Corpus files: raw documents, clean shards, the shard manifest

The corpus pipeline (python/corpus/) turns fetched sources into clean, deduplicated, PII-scrubbed Parquet shards that the tokenizer step (data.07) packs into token streams. Paths are relative to /artifacts ([paths].artifacts). The pipeline’s inputs are fixed by a corpus config (corpus-config.schema.json) and its sources by the data ledger.

corpus/raw/<source_id>/<yyyymmdd>/part-<nnnnn>.jsonl.zst: zstd-compressed JSON Lines, one document per line, UTF-8:

{"url": "https://example.org/story/1", "fetched_at": "2026-10-09T12:00:00Z", "license_spdx": "CDLA-Sharing-1.0", "text": "Once upon a time..."}
FieldTypeMeaning
urlstringwhere the document came from (for a single-file source, the file URL plus #<line>)
fetched_atRFC 3339 UTCfetch time
license_spdxstringSPDX identifier from the ledger row of the source
textstringthe extracted text, before normalization

<yyyymmdd> is the fetch date. A rerun of fetch downloads nothing whose checksum already matches the ledger.

corpus/<dataset>/<version>/shard-<iiiii>-of-<nnnnn>.parquet, for example shard-00000-of-00064.parquet. Parquet with zstd compression, at most 64 MiB per row group, one row per kept document, columns in this order:

ColumnArrow typeMeaning
idstring<source_id>:<k>, where k is the document’s 0-based position in its source’s raw files read in path order
textstringthe final text: normalized (Unicode NFC), filtered, PII scrubbed
source_idstringthe ledger’s source id
urlstringfrom the raw document
license_spdxstringfrom the raw document
langstringISO 639-1 code from the language filter (und when undetermined)
n_charsint32Unicode code points in text
sha256fixed_size_binary(32)SHA-256 of the UTF-8 bytes of text
minhash_clusterint64the near-duplicate cluster id (union-find root, data.04): the smallest row index of the cluster in id order; equal to the row’s own index when it has no near duplicate
pii_redactionsint32placeholders the PII scrub inserted (data.05)
splitstringtrain or val

Order and determinism. Rows are sorted by id (byte order) across the whole dataset and cut into shards of shard_rows rows (the corpus config); nnnnn is the shard count. The output (every shard and the manifest) is byte-identical across runs and across worker counts.

Split. A document is val when the first 8 bytes of its sha256, read as a big-endian unsigned integer, modulo 1000 are below the config’s val_permille; otherwise train. Splitting by content hash keeps exact duplicates on one side.

Placeholders. The PII scrub replaces each match with a typed placeholder: <EMAIL>, <PHONE>, <CARD>, <IP>, <KEY>.

corpus/<dataset>/<version>/_MANIFEST.json, schema corpus-manifest.schema.json. It records what was kept, what was dropped and why, and which ledger and config produced it:

{"dataset": "tinystories", "version": "v1", "n_docs": 2119719, "n_shards": 64,
"shards": [{"name": "shard-00000-of-00064.parquet", "sha256": "<hex>", "rows": 33121}],
"filters": {"lang": 1203, "gopher": 8812, "repetition": 77, "ppl": 0},
"dedup": {"exact_dropped": 5120, "near_dropped": 18044, "decontaminated": 12,
"jaccard_threshold": 0.8, "num_perm": 128, "bands": 16, "ngram": 13},
"pii": {"emails": 3, "phones": 0, "cards": 0, "ips": 1, "keys": 0},
"ledger_ref": "corpus/LEDGER.jsonl", "config_sha256": "<hex>"}

n_docs equals the sum of rows; filters counts documents each quality filter dropped; dedup.decontaminated counts documents dropped for sharing an ngram-token n-gram with a protected eval or validation set (data.04).