Skip to content

Privacy and PII policy

Moduleethics.02 · practice · docs · Pass 3 · 3 to 5 h
You builddocs/data/PII_POLICY.md (your privacy policy, eight fixed sections), docs/data/pii-policy.toml (the kinds, placeholders, audit, and log rules a program checks), and docs/data/pii-review.toml (your answers to twelve review cases)
Contractthe kinds and placeholders are the PII scrub’s: contracts/py/corpus/pii.pyi (KINDS, PLACEHOLDERS); the per-source pii_policy is a field of formats/ledger.schema.json; the file formats are in section 4
Testscourse/tests/ethics.02/check runs test_pii_policy.py against your three files (what each test checks: section 4)
Needsnothing to call · reading: ethics.01 (a policy as prose plus checkable decisions)
Used byno code call site (a policy): data.05 implements the scrub it describes and is graded on recall and precision over the same lookalikes, and gw.08 redacts gateway logs with the same kinds
MilestoneMS-P3
Optional depthCarlini et al., Extracting Training Data from Large Language Models (2021, free); Lukas et al., Analyzing Leakage of Personally Identifiable Information in Language Models (2023, free); NIST Privacy Framework 1.0 (free); GDPR articles 5 and 17 (gdpr-info.eu, free)

Like ethics.01, this is engineering, not legal advice: the course checks that your decisions are written down, consistent with what your pipeline does, and testable, not that they satisfy a particular law.

  • Language models memorize, and rare strings seen more than once (an email in a mirrored page, a key in a copied config file) are the easiest to extract. What enters the corpus can come out of the model (test_data_flows_name_the_ways_in_and_out).
  • A policy names every flow: personal data enters through the corpus, prompts, and evaluation suites, and leaves through generations, logs, traces, and reports.
  • Scrubbing replaces a value with a typed placeholder (<EMAIL>, <CARD>): the text keeps its shape for training and loses the value (test_every_category_is_named_once_with_its_placeholder).
  • A detector is a classifier with a precision and a recall. The lookalikes (an ISBN, a number that fails the Luhn check, a trace id) decide its precision (test_review_cases_are_classified).
  • An audit trail records what was replaced and where, never the value itself (test_audit_never_stores_the_value).
Terminal window
ol start ethics.02 # records the start; you write docs/data/ from section 4
ol tests ethics.02 # read the test catalog first
ol check ethics.02 # exit code is the verdict

data.01 is downloading text written by strangers, and some of it contains their email addresses, phone numbers, and, in copied configuration files, live API keys. In this pass data.05 scrubs them, and in Pass 7 your gateway starts receiving prompts that contain your users’ personal data and writing logs about them. Carlini et al. extracted hundreds of verbatim training sequences from GPT-2, among them names, phone numbers, and email addresses, simply by sampling and ranking; a scrub that runs after training is too late. Before any detector is written, the system needs a written answer to: which kinds of personal data do we look for, what do we replace them with, what do we keep as evidence that we did, how long do logs live, and how does a person get their data removed. data.05 and gw.08 then implement those answers, and the tests of both are written against the same lookalikes you classify here.

Personal data is any information about an identified or identifiable person: a direct identifier (a name with an address, an email, a phone number, a card number) or an indirect one that can be tied to a person with other data (an IP address, which an internet provider can map to a subscriber). Secrets are not personal but are as dangerous: an API key grants whoever reads it the access of its owner. The course detects five kinds, the ones data.05 implements (KINDS in contracts/py/corpus/pii.pyi):

KindWhat data.05 matchesPlaceholder
emaillocal@domain.tld<EMAIL>
phoneNorth American numbers, and + followed by 8 to 15 digits<PHONE>
card13 to 19 digits, first digit 2 to 6, passing the Luhn check<CARD>
ipIPv4 dotted quads with octets 0 to 255, and IPv6 addresses<IP>
keyknown API-key shapes: sk-..., AKIA + 16, ghp_..., xox?-..., AIza..., and the course’s own tl_<id>_<secret><KEY>

Names are personal data too, but no rule-based detector finds them reliably; the policy says so instead of pretending.

FlowDirectionControl
the corpus (fetched sources)inthe scrub before training (data.05)
prompts and uploaded documentsinnot logged; the gateway redacts what it does log (gw.08)
evaluation suitesinwritten by you; no real personal data
generationsoutthe scrub upstream; memorization is reduced by dedup (data.03, data.04)
logsoutredaction with the same kinds, and a retention period
tracesoutlengths and ids only, never prompt or completion text
evaluation reportsoutquote model outputs: the same rules as logs

A flow the policy does not name is a flow nobody controls.

The ledger’s pii_policy field (formats/ledger.schema.json) decides what data.05 does to each source:

  • scrub: replace every detected value by its placeholder and keep the document. “Write to for the slides.” still teaches the model how such a sentence goes.
  • drop: remove every document with a detection. For a source dominated by personal data, a forum dump for example, the documents are not worth keeping.
  • none: skip detection. Only for a named source someone checked and found free of personal data, such as your own synthetic fixtures; it is a claim, so it is never the default.

Data minimization comes before all three: the cheapest personal data to protect is the data you never fetch or log.

A detector decides, for each candidate string, “personal data” or not. Compare its decisions with labels:

SymbolMeaningType
TPTPtrue positives: personal data the detector flaggedcount
FPFPfalse positives: harmless strings it flagged (an ISBN replaced by <CARD>)count
FNFNfalse negatives: personal data it missedcount
precision =TP/(TP+FP)= TP / (TP + FP)of what was flagged, the share that was personal datain [0,1][0, 1]
recall =TP/(TP+FN)= TP / (TP + FN)of the personal data, the share that was flaggedin [0,1][0, 1]

Recall protects people; precision protects the corpus (every false positive destroys real text, and version strings, dates, and identifiers are common in technical writing). data.05 is graded on both: recall of at least 0.98 on emails and cards, precision of at least 0.95, and no false positives on a list of lookalikes.

Every payment card number ends in a check digit chosen so that the Luhn sum is a multiple of 10. With the digits d1d2…dLd_1 d_2 \dots d_L read from the right (d1d_1 is the check digit):

SymbolMeaningType
did_ithe ii-th digit from the rightinteger in 0..90..9
eie_idid_i when ii is odd; 2di2 d_i when ii is even, minus 9 if that exceeds 9integer in 0..90..9
S=∑ieiS = \sum_i e_ithe Luhn suminteger

The number passes when S mod 10=0S \bmod 10 = 0. A random 16-digit number passes with probability 1/10, so the check alone removes 90% of the false positives that “any 16 digits” produces.

An audit span records one replacement: the document id, the kind, and the start and end offsets in the original text (PiiSpan in data.05). It never records the value: an audit log full of values is a second copy of the PII, in a file nobody thought to protect. Logs pass through the same detector, and they expire after a fixed number of days; a log kept forever becomes a corpus. An erasure request (GDPR article 17 is one legal basis) asks you to remove a person’s data: the policy says how you find it (search raw parts and shards for the identifiers they give), what you rebuild (shards, token streams), and what happens to models trained on it.

Luhn. Check 79927398713. From the right, double every second digit:

ii1234567891011
did_i31789372997
eie_i32716 - 9 = 79674918 - 9 = 97

S=3+2+7+7+9+6+7+4+9+9+7=70S = 3 + 2 + 7 + 7 + 9 + 6 + 7 + 4 + 9 + 9 + 7 = 70, a multiple of 10: the number passes. Change the last digit to 4 and S=71S = 71: it fails.

Precision of a naive detector. Among the review cases of section 4, three strings hold 13 or more digits: the card 4539 1488 0343 6467 (Luhn sum 80), the order number 4539 1488 0343 6468 (Luhn sum 81), and the ISBN 978-0-306-40615-7 (13 digits, Luhn sum 51). A detector that flags “13 to 19 digits” as a card has TP=1TP = 1, FP=2FP = 2: precision 1/31/3. Adding the Luhn check removes both false positives: precision 1/11/1, recall unchanged. (data.05 also requires a first digit from 2 to 6, which rules out every ISBN-13: they start with 978 or 979.)

A key or not. trace_id=4bf92f3577b34da6a3ce929d0e0e4736 is 32 random hex digits, the shape of a W3C trace id: it identifies a request, grants nothing, and appears in every trace you will export, so flagging it would scrub your own observability. AKIAIOSFODNN7EXAMPLE is AKIA plus 16 upper-case letters and digits, the shape of an AWS access key id: a key. The difference is a known prefix, not the randomness.

Write three files in your repo.

docs/data/PII_POLICY.md: prose with these eight ## sections (the check counts words, 8 to 30 per section) and no template placeholders left:

SectionAnswers
## Scopewhich data, systems, and models the policy covers
## Data flowsevery way personal data enters and leaves (2.2); name the corpus, prompts, logs, and traces
## Categories and actionsthe five kinds, and when a source is scrubbed, dropped, or left alone
## Placeholderswhat replaces each kind and why the text keeps its shape
## Auditwhat an audit span records, and that it never records the value
## Logs and retentionwhat is logged, what is redacted, how long logs live
## Erasure requestshow a person gets their data removed, and what you rebuild
## Owner and reviewwho decides, and how often the policy is reviewed

docs/data/pii-policy.toml:

version = 1
owner = "who decides"
reviewed = 2026-10-09 # a TOML date, no quotes
default_source_policy = "scrub" # scrub or drop: the ledger pii_policy of a new source
[[category]] # once for each of email, phone, card, ip, key
name = "email"
placeholder = "<EMAIL>" # exactly data.05's placeholder for the kind
why = "One sentence, at least six words."
[audit]
store_values = false
fields = ["doc_id", "kind", "start", "end"] # from doc_id, source_id, kind, start, end
[logs]
redact = ["email", "phone", "card", "ip", "key"]
retain_days = 30 # whole days, 1 to 365

docs/data/pii-review.toml: one line per case, the value one of email, phone, card, ip, key, none:

CaseString
c01Write to ana.lopez@example.org for the slides.
c02Call me at +1 (415) 555-0132 after six.
c03Card on file: 4539 1488 0343 6467
c04Order number 4539 1488 0343 6468
c05ISBN 978-0-306-40615-7
c06Upgraded the engine to v2.10.3 last night.
c07The request came from 203.0.113.42 at noon.
c08Peer 2001:db8:85a3::8a2e:370:7334 dropped the stream.
c09export AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE
c10Meeting moved to 2026-10-09 at 14:30.
c11trace_id=4bf92f3577b34da6a3ce929d0e0e4736
c12Authorization: Bearer tl_k7f3_9s8d7f6g5h4j3k2l1m0nq8w7
c01 = "email"
# ... c02 to c12

ol check ethics.02 runs course/tests/ethics.02/check from your repo root; it refuses early if a file is missing.

TestKINDChecksWhy it matters downstream
test_policy_has_every_sectionunitthe eight sections, each with enough wordsevery question has an answer
test_policy_has_no_placeholdersunitno TODO or <lower-case placeholder> outside codea template promises nothing
test_data_flows_name_the_ways_in_and_outunitthe Data flows section names corpus, prompts, logs, tracesevery flow has a control
test_policy_toml_has_an_owner_and_a_review_dateunitversion = 1, an owner, a past TOML date, no unknown keyssomeone maintains it
test_every_category_is_named_once_with_its_placeholderunitthe five kinds, once each, with data.05’s placeholders and a reasonthe policy matches what the shards contain
test_new_sources_are_scrubbed_or_droppedboundarydefault_source_policy is scrub or drop“none” is a claim, never a default
test_audit_never_stores_the_valueboundarystore_values = false; fields include kind, start, end and nothing that could hold the valuethe audit log is not a second copy
test_logs_redact_every_category_and_expireboundaryevery kind redacted; retention 1 to 365 dayslogs do not become a corpus
test_review_cases_are_classifiedunitthe twelve cases, lookalikes includedthe precision data.05 is graded on
PitfallSymptomCaught by
1. A data-flow list that stops at the corpusprompts in logs and text in traces are nobody’s jobtest_data_flows_name_the_ways_in_and_out
2. Placeholders in the policy that differ from what the scrub writesan auditor cannot match the policy to the shardstest_every_category_is_named_once_with_its_placeholder
3. none as the default source policyevery new source skips detection until someone noticestest_new_sources_are_scrubbed_or_dropped
4. Logging the matched value “for debugging”the audit log holds every email the corpus hadtest_audit_never_stores_the_value
5. Logs with no expirya growing archive of prompts and personal datatest_logs_redact_every_category_and_expire
6. Treating any 16 digits as a cardorder numbers and ISBNs replaced; precision of 1/3 on the reviewtest_review_cases_are_classified (cases c04, c05)
7. Treating any long random hex as a secrettrace ids scrubbed from your own observability datatest_review_cases_are_classified (case c11)
DirectionModuleHow it uses this
Backethics.01the same shape: policy in prose, decisions in data, a check that runs
Forwarddata.05the scrub with these kinds and placeholders, the per-source pii_policy, and audit spans without values; graded on recall, precision, and the lookalikes
Forwarddata.08the datasheet reports the redaction counts per kind
Forwardgw.08gateway log redaction ports the detector list to Go
Forwardethics.03the model card states what personal data the training corpus may still contain
Your pieceProduction equivalentWhat it addsWhere to look
the detector listMicrosoft Presidiorecognizers per entity with context words, named-entity models for names and places, configurable anonymizersmicrosoft/presidio
the scrub policyBigScience ROOTS and StarCoder PII pipelinesPII removal at corpus scale, with a trained detector (StarPII) for names, emails, keys, and passwords in codeLi et al., StarCoder: may the source be with you! (2023), the PII redaction section
memorizationdifferential privacy (DP-SGD), deduplicationprovable bounds on what one training example can change; less memorization from fewer repeatsAbadi et al., “Deep Learning with Differential Privacy” (2016)