Skip to content

Data licensing and the ledger

Moduleethics.01 · practice · docs · Pass 3 · 3 to 5 h
You builddocs/data/LICENSE_POLICY.md (your licensing policy, eight fixed sections) and docs/data/license-allowlist.toml (the machine-readable decisions data.08 enforces)
Contractthe ledger row: formats/ledger.schema.json; the allowlist format is defined in section 4 of this chapter
Testscourse/tests/ethics.01/check runs test_license_policy.py against your two files (what each test checks: section 4)
Needsnothing: this is policy, written before the pipeline it governs
Used byno code call site (a policy): data.01 records a ledger row per source, data.08 verifies every row and shard against your allowlist, ethics.03 reports it in the datasheet, and dur.12 refuses to release a model whose ledger does not verify
MilestoneMS-P3
Optional depthLongpre et al., The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI (2023, free); the SPDX License List and the SPDX specification, annex D “License expressions” (spdx.org, free); Creative Commons, “About CC licenses” (free); Community Data License Agreement texts (cdla.dev, free)

This chapter teaches you to make provenance auditable. It is not legal advice: whether a license permits training a model is decided by a lawyer for your situation. What the course fixes is that every decision is written down, uses names a program can check, and is checked on every build.

  • Text you did not write is somebody’s, and without a license all rights are reserved: “no license” means “no”, not “free” (test_unknown_licenses_are_refused).
  • A license is named by its SPDX identifier, the one spelling the ledger, the allowlist, and the datasheet share (test_every_id_is_an_spdx_identifier).
  • An allowlist is a policy in data: each license allowed for named uses with its obligations, or refused with a reason (test_every_entry_is_complete, test_every_review_case_is_decided).
  • Non-commercial and no-derivatives terms are incompatible with training a model you release; the course refuses them for training (test_non_commercial_and_no_derivatives_are_never_trained_on).
  • The ledger ties a license to the sha256 of the exact bytes you fetched, so the decision cannot drift when the URL starts serving something else (test_ledger_section_names_the_fields).
Terminal window
ol start ethics.01 # records the start; you write docs/data/ from section 4
ol tests ethics.01 # read the test catalog first
ol check ethics.01 # exit code is the verdict

Your fetcher (data.01, this pass) is about to download every source in your corpus config, and in Pass 9 a workflow will train a model on the result and publish it (dur.12). Between the two, somebody has to be able to answer, for every document in every shard: where did this come from, under what terms, and do those terms let us train on it and release what we trained? The Data Provenance Initiative audited over 1,800 popular fine-tuning datasets and found licenses missing for most of them on the hosting sites and miscategorized for a large share, usually more permissive than the original authors granted. Once text is mixed into a corpus the answer cannot be reconstructed, so the policy has to exist before the first download, and it has to be data a build can check, not a paragraph in a README.

Copyright gives the author of a text the exclusive right to copy it, adapt it, and distribute it, automatically and without registration. A license is the author’s written permission to do some of those things under conditions. With no license, nothing is permitted beyond what the law allows anyway (fair use and text-and-data-mining exceptions vary by country and are exactly what you would need a lawyer for). Separately, a website’s terms of service can bind you by contract even for text that is not protected. A dataset aggregated from many sources has one license per source; the dataset card’s single license line is a claim about all of them that someone else made.

FamilyExamples (SPDX ids)You mayThe license asks
Public domain dedicationCC0-1.0, PDDL-1.0anythingnothing
PermissiveMIT, Apache-2.0, BSD-3-Clause, CDLA-Permissive-2.0copy, adapt, redistributekeep the copyright and license text (notice)
AttributionCC-BY-4.0, ODC-By-1.0copy, adapt, redistributecredit the source (attribution)
ShareAlikeCC-BY-SA-4.0, CDLA-Sharing-1.0, ODbL-1.0copy, adapt, redistributederived data you share keeps the same license (share-alike)
Non-commercialCC-BY-NC-4.0, CC-BY-NC-SA-4.0use for non-commercial purposesno commercial use
No derivativesCC-BY-ND-4.0, CC-BY-NC-ND-4.0copy unchangedno adaptations
None or unknownNOASSERTION (nobody said), NONE (no license)nothing

Whether a trained model is an “adaptation” of its training text is unsettled. CDLA-Sharing-1.0 was written for this question and says that results computed from the data, a model among them, are not bound by its share-alike term; the Creative Commons licenses do not say. That is why the allowlist records uses separately: a license can be allowed for evaluation (reading text to score a model) and refused for training (shaping the weights of a model you release).

The SPDX License List gives every common license a short, case-sensitive identifier: CC-BY-4.0, not “CC BY 4.0” or “cc-by-4.0”; Apache-2.0, not “Apache 2.0”. A license SPDX does not list is written LicenseRef-<name>. An expression combines identifiers: MIT OR Apache-2.0 (you may choose either: the uses of both together) and CC-BY-4.0 AND MIT (both apply: only the uses both allow). data.08 evaluates expressions against your allowlist; an identifier it does not find is refused.

Your allowlist is your decision. Four rules are not, because the rest of the course depends on them:

  1. Unknown is refused. NOASSERTION is refused explicitly, and NONE can never be allowed.
  2. No NC or ND in training. Your model’s weights are released (dur.12) under a license you choose; a non-commercial term forbids commercial use of what you build from the data and a no-derivatives term forbids adaptations, so neither can be in a training set. Evaluation use is your call.
  3. Obligations are recorded. Every allowed attribution license (CC-BY*, ODC-By) lists attribution; every ShareAlike license (*-SA-*, CDLA-Sharing) lists share-alike. The datasheet (ethics.03) prints them from here.
  4. The course’s own sources are allowed for training: TinyStories, the C1 corpus, is CDLA-Sharing-1.0, and the course fixtures are Apache-2.0. Refusing them is a valid policy, but then the course corpus cannot be built.

Every fetched source gets one row in corpus/LEDGER.jsonl (formats/ledger.schema.json), written by data.01 at the moment the bytes, the URL, and the license you reviewed are all in one place:

FieldRecordsWhy it matters
source_idyour name for the sourceshard rows point back to it
urlwhere the bytes came fromanyone can go and look
license_spdxthe license you reviewed when you pinned the sourcewhat the allowlist is checked against
retrieved_atUTC time of the downloadthe terms in force then
sha256digest of the exact bytes receivedthe license applies to these bytes, not to whatever the URL serves later
allowed_usestrain, eval as the config claimsdata.08 checks the claim against the allowlist
pii_policy, filters_applied, kept, dropped, noteswhat the pipeline didthe datasheet

Rows are appended, never edited: a new version of a source is a new row with a new checksum, and a withdrawn source gets a row with revoked: true.

A license can turn out wrong, an author can ask for removal, a host can change its terms. The policy says in advance what happens: the source is refused from then on, every shard and token stream derived from it is rebuilt without it, and a model trained on it is retrained or withdrawn before its next release. The data-incident drill (ops.08) rehearses exactly this.

A different corpus than yours, to decide by hand. Three sources:

SourceLicense on the cardWhat you find
poems“Public Domain”the site says the poems were dedicated with CC0
wiki-mini“CC BY-SA”the footer says Creative Commons Attribution-ShareAlike 3.0
forum(none)a scrape of a forum; the terms of service reserve all rights

The decisions:

  1. poems: the SPDX identifier is CC0-1.0. Allowed for train and eval, no obligations.
  2. wiki-mini: CC-BY-SA-3.0. Attribution and ShareAlike. If you decide a released model is not an adaptation under this license, allow train and eval with obligations attribution and share-alike; if you are unsure, allow eval only and say so in why. This example chooses eval.
  3. forum: no license and restrictive terms: the ledger records NOASSERTION, refused.

The allowlist entries:

[[allow]]
spdx = "CC0-1.0"
uses = ["train", "eval"]
obligations = []
why = "A public domain dedication: no conditions on use or redistribution."
[[allow]]
spdx = "CC-BY-SA-3.0"
uses = ["eval"]
obligations = ["attribution", "share-alike"]
why = "Evaluation only until we decide whether a released model is a ShareAlike adaptation."
[[refuse]]
spdx = "NOASSERTION"
why = "Nobody recorded a license, so all rights are reserved until someone finds the terms."

Now suppose the corpus config claims allowed_uses = ["train", "eval"] for wiki-mini. data.08 reads its ledger row, finds CC-BY-SA-3.0 allowed for eval only, and fails verification with exit 65: the config claims a use the policy does not grant. The fix is in the config (allowed_uses = ["eval"]), not in the allowlist.

Write two files in your repo.

docs/data/LICENSE_POLICY.md: prose, with these eight ## sections, each long enough to answer its question (the check counts words, from 8 to 25 per section), and no template placeholders (TODO, <your text>) left:

SectionAnswers
## Scopewhich data and which models the policy covers
## Allowed useswhat train and eval mean for you
## Ledgerwhat the ledger records; name the fields source_id, license_spdx, retrieved_at, sha256, allowed_uses
## Allowlistwhere the decisions live and what happens when a row fails them
## Obligationswhat attribution, share-alike, and notice require of you in practice
## Unknown and missing licensesthe default answer and why
## Revocationwhat happens to shards and models when a license is withdrawn
## Owner and reviewwho decides, and how often the list is reviewed

docs/data/license-allowlist.toml:

version = 1
owner = "who decides" # a person or a team
reviewed = 2026-10-09 # a TOML date, no quotes
[[allow]] # one per allowed license
spdx = "CC-BY-4.0" # an SPDX identifier, or LicenseRef-<name>
uses = ["train", "eval"] # a non-empty subset of train, eval
obligations = ["attribution"] # from: attribution, share-alike, notice, no-endorsement
why = "One sentence, at least six words."
[[refuse]] # one per refused license
spdx = "CC-BY-NC-4.0"
why = "One sentence, at least six words."

The review. Decide each of these seven licenses, which you will meet on dataset cards, as an [[allow]] or a [[refuse]]: CC-BY-4.0, CC-BY-SA-4.0, CC-BY-NC-4.0, CC-BY-ND-4.0, ODC-By-1.0, MIT, NOASSERTION. Add CDLA-Sharing-1.0 and Apache-2.0 (rule 4 of section 2.4) and anything else your sources use.

ol check ethics.01 runs course/tests/ethics.01/check from your repo root; it refuses early if either file is missing.

TestKINDChecksWhy it matters downstream
test_policy_has_every_sectionunitthe eight sections, each with enough wordsevery question has an answer
test_policy_has_no_placeholdersunitno TODO, TBD, or <lower-case placeholder> outside codea template is not a decision
test_ledger_section_names_the_fieldsunitthe Ledger section names the five fieldsreaders know where a decision is recorded
test_allowlist_has_an_owner_and_a_review_dateunitversion = 1, an owner, a past TOML date, no unknown keyssomeone maintains it
test_every_entry_is_completeunituses, obligations, and why in the vocabularies; no unknown keysdata.08 reads every field
test_every_id_is_an_spdx_identifierboundaryeach id is on the course’s SPDX list or LicenseRef-...exact-match checks in data.08
test_no_license_is_both_allowed_and_refusedboundaryone decision per licenseno order-dependent answers
test_the_course_sources_are_allowed_for_trainingunitCDLA-Sharing-1.0 and Apache-2.0 allow trainyour corpus can be built
test_every_review_case_is_decidedunitthe seven review licenses are each allowed or refusedno undecided license slips in
test_non_commercial_and_no_derivatives_are_never_trained_onboundaryno -NC/-ND license allows trainthe released model is yours to license
test_unknown_licenses_are_refusedboundaryNOASSERTION refused, NONE never allowedthe default is no
test_obligations_follow_the_license_familyunitattribution and share-alike recorded where the family requires themthe datasheet credits and relicenses correctly
PitfallSymptomCaught by
1. Copying the license as the card prints it (“CC BY 4.0”)data.08’s exact match refuses a source you meant to allowtest_every_id_is_an_spdx_identifier
2. Treating “no license” as “free to use”a scraped source enters the corpus with no permission at alltest_unknown_licenses_are_refused
3. Allowing CC-BY-NC-* for training because the project is a hobbythe released model inherits a restriction its license does not statetest_non_commercial_and_no_derivatives_are_never_trained_on
4. Allowing a license without recording what it asksthe datasheet omits credits; ShareAlike data is redistributed under the wrong termstest_obligations_follow_the_license_family
5. Leaving a review case undecidedthe first source under it fails the build at the worst time, or a permissive default lets it intest_every_review_case_is_decided
6. Listing a license in both tablesthe answer depends on which line a program reads lasttest_no_license_is_both_allowed_and_refused
7. A policy with no owner or review datethe list is never updated after the first weektest_allowlist_has_an_owner_and_a_review_date
8. A Ledger section that never says what is recordednobody can tell which field carries the license decisiontest_ledger_section_names_the_fields
DirectionModuleHow it uses this
Forwarddata.01writes the ledger row of section 2.5 for every source it fetches, with the license you pinned
Forwarddata.08{corpus} ledger verify checks every row and every shard against your allowlist and exits 65 on a refused or unknown license
Forwardethics.02the same structure for personal data: a policy in prose plus decisions a program checks
Forwardethics.03the datasheet lists every source, its license, and the obligations recorded here
Forwarddur.12ModelRelease refuses a model whose training sources are not licensed for train
Your pieceProduction equivalentWhat it addsWhere to look
the allowlistScanCode Toolkit, FOSSA, ClearlyDefinedlicense detection from text, curated license data per package and datasetaboutcode-org/scancode-toolkit
the ledgerData Provenance Initiative’s Data Provenance Explorer and provenance cardslineage across re-releases, license categories per datasetdataprovenance.org
the reviewCommonCanvas, Common Corpus, the Common Pilecorpora built only from openly licensed or public domain text, with per-document licensesGokaslan et al. 2023; Kandpal et al. 2025