Skip to content

Licensing ledger verification and the datasheet

Moduledata.08 · build · Python · Pass 3 · 2 to 3 h
You buildpython/corpus/ledger.py: read_ledger, latest, row_errors, permitted_uses, verify, check, reconcile, datasheet, the LedgerError exception, and the tables ALLOWLIST, EX_DATAERR
Contractcourse/contracts/py/corpus/ledger.pyi · the rows: formats/ledger.schema.json · the template: templates/DATASHEET.md · exit codes: spec/subprocess-activity.md
Testscourse/tests/data.08/ (what they check: section 4)
Needsdata.05 the PII kinds · data.06 read_shards and the manifest · data.07 the tokens manifest (or --ref-deps) · reading: data.01 (it appends the ledger rows), ethics.01 (the license policy)
Used byethics.03’s check judges the datasheet’s licenses with your permitted_uses · later: dur.12 refuses to release a model whose sources are not licensed for training · ops.08 purges a revoked source
MilestoneMS-corpus
Optional depthGebru et al., Datasheets for Datasets (free); the SPDX license expression syntax (free); Longpre et al., The Data Provenance Initiative (free)
  • Every row of training text traces to a ledger row whose license permits training, under the same license string; anything that does not trace fails the build (test_every_shard_row_traces_to_its_source).
  • License expressions are evaluated, not pattern-matched: OR is the union of what each license permits, AND the intersection, and anything outside the allowlist is unknown (test_expressions).
  • A licensing failure exits 65, a data error, which the durable engine never retries: no retry will change a license (test_unknown_license_is_a_data_error).
  • The ledger is append-only and its last row per source wins, so a revocation is one appended line (test_latest_row_wins); the datasheet is generated from the same manifest and ledger, so its numbers cannot drift from the data (test_datasheet).
Terminal window
ol start data.08 # stubs ledger.py into python/corpus/
ol tests data.08 # the course test catalog
ol tdd red data.08 # rung R3: your tests first, failing against the stubs
ol check data.08 # exit code is the verdict
ol check data.08 --ref-deps # only if data.05, data.06, or data.07 is not passing yet
ol mutate data.08 # how many planted bugs your tests catch
ol diff data.08 # after passing: your code against the reference

Your CLI gains {corpus} ledger verify and {corpus} datasheet (MS-corpus fixes their flags); copy the generated datasheet to docs/DATASHEET.md and write its Motivation section yourself.


After data.07 you have a corpus, shards, and token streams, and a ledger that data.01 filled as it fetched: one row per source with its URL, license, and hash. Nothing checks that the two agree. A shard row could come from a source the ledger never heard of, a non-commercial source could claim “train”, a source revoked last week could still be in the shards, and the kept counts still say what fetch saw, not what survived the filters. When C1 releases a model, dur.12 must answer “may this model be trained on this data?”, and when someone asks what is in your corpus, you need a datasheet whose numbers come from the data, not from memory. This module verifies the ledger against the shards, reconciles its counts, and generates the datasheet.

SymbolMeaningType / shape
U(ℓ)U(\ell)the set of uses license ℓ\ell permits, a subset of {train, eval}frozenset[str]
ℓ1 OR ℓ2\ell_1 \text{ OR } \ell_2the licensee may pick either licenseexpression
ℓ1 AND ℓ2\ell_1 \text{ AND } \ell_2both licenses apply at onceexpression
rowone ledger line for one sourcedict
c(s)c(s)rows in the shards whose source_id is ssint

The allowlist is a policy, written as data. ALLOWLIST maps SPDX license ids to the uses they permit: permissive and attribution licenses (CC0-1.0, CC-BY-4.0, CDLA-Permissive-2.0, CDLA-Sharing-1.0, MIT, Apache-2.0, …) permit train and eval; non-commercial and no-derivatives licenses (CC-BY-NC-4.0, CC-BY-ND-4.0, …) permit eval only, because training a model is the kind of reuse they exclude. ethics.01 is where that policy is argued; this module enforces it. An id not in the table is unknown, never guessed.

Expressions. SPDX writes combined licenses as expressions. MIT OR Apache-2.0 (dual licensing: you choose) permits U(MIT)∪U(Apache-2.0)U(\text{MIT}) \cup U(\text{Apache-2.0}). MIT AND CC-BY-NC-4.0 (both apply, as when a dataset bundles two works) permits U(MIT)∩U(CC-BY-NC-4.0)U(\text{MIT}) \cap U(\text{CC-BY-NC-4.0}). AND binds tighter than OR, so A OR B AND C is U(A)∪(U(B)∩U(C))U(A) \cup (U(B) \cap U(C)). Operators are upper case; parentheses and WITH exceptions are outside this subset, so they make the expression unknown.

What verification checks. verify(ledger, corpus, use) returns every problem, an empty list when the corpus may be used for use:

  1. every ledger row follows formats/ledger.schema.json (row_errors; a bad row is a problem, and its source then has no valid row);
  2. for every source’s latest valid row: its license is known; its allowed_uses claim nothing its license does not permit;
  3. for a source with shard rows: it is not revoked, use is in its allowed_uses, and every one of its rows carries its license string;
  4. every shard row’s source_id has a ledger row;
  5. every source’s kept equals c(s)c(s).

Sources with no rows in the shards (an eval-only benchmark) need not permit training. check raises LedgerError(problems); its exit_code is EX_DATAERR = 65.

Exit 65 means do not retry. spec/subprocess-activity.md gives the durable engine (data.09) one rule: 75 (and most other codes) means try again, 65 means the input is wrong and retrying is pointless. A license problem is the textbook data error; exiting 75 would make CorpusBuild retry until its budget runs out and then dead-letter a build that could never pass.

Append-only, last row wins. The ledger is history. data.01 appends a row per fetch; reconcile(ledger, corpus) appends, for every source, a copy of its latest row with kept =c(s)= c(s), dropped =max⁡(0,ndocs−c(s))= \max(0, n_\text{docs} - c(s)), and the filters applied; a revocation (drill ops.08) appends a row with revoked: true. latest keeps the last row per source_id. Nothing is rewritten, so the ledger shows what was decided and when.

The datasheet. Gebru et al. ask fixed questions of every dataset: motivation, composition, collection, preprocessing, uses, distribution. datasheet() fills templates/DATASHEET.md with what the files know: documents, splits, languages, PII counts (Composition); every source with URL, license, retrieval time, and hash (Collection); every drop count and dedup parameter, the decontamination count, and the token counts (Preprocessing); each source’s allowed uses (Uses); the config hash and the ledger path (Distribution and maintenance). It leaves Motivation to you: why the dataset exists is not in any file.

What each expression permits.

License expressionEvaluationPermits
CC-BY-4.0listedtrain, eval
CC-BY-NC-4.0listedeval
MIT OR CC-BY-NC-4.0{t,e}∪{e}\{t, e\} \cup \{e\}train, eval
MIT AND CC-BY-NC-4.0{t,e}∩{e}\{t, e\} \cap \{e\}eval
CC-BY-NC-4.0 OR CC-BY-ND-4.0 AND MIT{e}∪({e}∩{t,e})\{e\} \cup (\{e\} \cap \{t, e\})eval
Apache-2.0 WITH LLVM-exceptionWITH is outside the subsetunknown: exit 65

One corpus, two sources. The shards hold stories:0, stories:1 (CC-BY-4.0) and papers:0 (CC-BY-NC-4.0). The ledger:

source_idlicense_spdxallowed_usesn_docskept
storiesCC-BY-4.0train, eval22
papersCC-BY-NC-4.0train11

Walk the checks for use = "train". stories: license known, claims {t,e}⊆{t,e}\{t, e\} \subseteq \{t, e\}, not revoked, train allowed, both rows say CC-BY-4.0, kept 2=c2 = c. No problem. papers: license known, but claims train, which CC-BY-NC-4.0 does not permit: one problem. (Train is in its claimed allowed_uses, so check 3 passes; the lie is caught at check 2.) Every shard row traces to a ledger row. verify returns one problem naming papers, and check raises LedgerError with exit code 65.

This is test_hand_example.

# python/corpus/ledger.py (the full contract is contracts/py/corpus/ledger.pyi)
EX_DATAERR: int # 65
ALLOWLIST: dict[str, frozenset[str]]
class LedgerError(Exception): # .exit_code == 65, .problems
def __init__(self, problems: list[str]) -> None
def read_ledger(path: Path) -> list[dict]
def latest(rows) -> dict[str, dict]
def row_errors(row) -> list[str]
def permitted_uses(license_spdx: str, allowlist=ALLOWLIST) -> frozenset[str] | None
def verify(ledger: Path, corpus: Path, *, use="train", allowlist=ALLOWLIST) -> list[str]
def check(ledger: Path, corpus: Path, *, use="train", allowlist=ALLOWLIST) -> None
def reconcile(ledger: Path, corpus: Path, *, filters_applied=()) -> list[dict]
def datasheet(ledger: Path, corpus: Path, tokens: Path | None = None) -> str
TestKINDChecksWhy it matters downstream
test_hand_exampleunitthe section 3 expressions and the one problem in the two-source corpusyou and the tests agree on the rules
test_expressionsunitunion, intersection, AND before OR; parentheses, WITH, lower case, unknown ids are unknowna license is never guessed
test_unknown_license_is_a_data_errorboundarycheck raises LedgerError, exit code 65, problems name the sourcedata.09 and dur.12 stop instead of retrying
test_rows_follow_the_schemaconformancerow_errors agrees with ledger.schema.json on a valid row and eleven violationsthe ledger is a contract too
test_bad_rows_are_reportedboundarya malformed row is a problem, and its source’s rows no longer traceno crash, no silent pass
test_every_shard_row_traces_to_its_sourceconformancerows from an unlisted source, or under another license string, failthe design’s traceability rule
test_latest_row_winsunitan appended revocation applies; reversed order passesrevocation is one appended line
test_use_is_checked_only_where_rows_existunitan eval-only source passes until its rows enter a training corpus; a train-only source fails for evalbenchmarks may live in the ledger
test_kept_counts_and_reconcileunitfetch’s kept fails; reconcile appends corrected rows (kept, dropped, filters) and leaves historythe ledger says what survived
test_read_ledgerboundaryblank lines skipped; a non-object or broken line raises naming itno silently skipped row
test_datasheetunitthe template’s sections in order; split, PII, source, drop, and token numbers from the files; Motivation left to youthe datasheet cannot drift from the data

Your tests (rung R3, red then green). Under python/tests/data-08-ledger/, failing first against the stubs: the hand-example expressions; an unknown license raising LedgerError with exit code 65; a clean corpus passing check; untraceable rows and a license mismatch failing; a non-commercial source claiming train; the latest row and revocation; schema violations; reconcile; read_ledger rejecting non-objects; the datasheet’s seven headings. ol mutate data.08 grades them: 0.70 of the mutants, including the one behind Pitfall 1.

PitfallSymptomCaught by
1. exiting 75 (or 1) on a license failureCorpusBuild retries a build that can never pass, then dead-letters ittest_unknown_license_is_a_data_error (mutant s05)
2. checking only the ledger, not the shard rowsrows from an unlisted source, or relicensed on the way, reach trainingtest_every_shard_row_traces_to_its_source (mutants s08, s09)
3. reading the first row of a sourcea revocation appended later is ignoredtest_latest_row_wins (mutant s10)
4. trusting fetch’s keptthe datasheet and the release gate count documents the filters removedtest_kept_counts_and_reconcile (mutant s12)
5. OR as intersection, AND as uniondual-licensed sources rejected, bundled non-commercial sources acceptedtest_expressions, test_hand_example (mutants s01, s02)
6. skipping unknown parts of an expression, or stripping parenthesesMIT OR GPL-3.0-only passes as MIT; (MIT) passes untestedtest_expressions (mutants s03, s04)
7. dropping invalid rows silentlya malformed row disappears instead of failing the buildtest_bad_rows_are_reported (mutant s07)
8. rewriting the ledger in reconcilethe history of what was fetched and decided is losttest_kept_counts_and_reconcile (mutant s13)
DirectionModuleHow it uses this
Backdata.05the PII kinds and their manifest names, reported in the datasheet
Backdata.06read_shards gives every row’s source and license; the manifest gives the counts
Backdata.07the tokens manifest gives the token count of each split
Backdata.01fetch appends the rows this module verifies
Backethics.01the license policy behind ALLOWLIST
Forwarddur.12ModelRelease runs check(use="train") before export and fails the release on exit 65 (Pass 9)
Forwardops.08the data-incident drill appends a revocation and purges the derived shards until verify passes again (Pass 11)
Forwardethics.03its check runs your permitted_uses over the ethics.01 allowlist on every license the datasheet lists; the model card links the datasheet (Pass 9)

MS-corpus already runs ledger verify through your CLI; ethics.03 (Pass 9) is the registered call site.

Your pieceProduction equivalentWhat it addsWhere to look
ALLOWLIST, permitted_usesthe SPDX license list, ScanCodeparsing full SPDX expressions with exceptions and parentheses; detecting licenses from file textnexB/scancode-toolkit, spdx/license-list-data
verifythe Data Provenance Initiative’s explorerper-dataset license, source, and creator audits across thousands of fine-tuning datasetsLongpre et al. 2023
the ledgerdataset cards and lineage stores (OpenLineage, Hugging Face dataset cards)lineage events per job, queryable across pipelinesthe OpenLineage spec
datasheetHugging Face dataset cards, Data Statements (Bender and Friedman)community templates, YAML metadata for search, language-variety statementshuggingface_hub DatasetCard