Skip to content

PII scrub with typed placeholders and audit spans

Moduledata.05 · build · Python · Pass 3 · 3 to 4 h
You buildpython/corpus/pii.py: luhn_ok, detect, redact, scrub, audit_line, scrub_stage (a stage), manifest_counts, the PiiSpan record, and the tables KINDS, PLACEHOLDERS, MANIFEST_KEYS
Contractcourse/contracts/py/corpus/pii.pyi · the counts it feeds: formats/corpus-manifest.schema.json · the policy field: formats/ledger.schema.json (pii_policy)
Testscourse/tests/data.05/ (what they check: section 4); fixture course/fixtures/data.05/labelled.jsonl (396 sentences with labelled spans and lookalikes)
Needsdata.02 Doc · data.04 near dedup runs first (or --ref-deps) · reading: ethics.02 (the PII policy this stage enforces)
Used bydata.06 sums the per-document counts into the manifest · data.08 reports them in the datasheet · later: gw.08 ports the detector list to Go for log redaction
MilestoneMS-corpus
Optional depthMicrosoft Presidio (analyzers and recognizers, free); Subramani et al., Detecting Personal Information in Training Corpora (free); ISO/IEC 7812-1 (card numbers and the Luhn check); RFC 5952 (IPv6 text form)
  • A detector is a candidate regex plus a validator: the regex finds what looks right, the validator (Luhn plus a network prefix for cards, an address parser for IPv6) throws out lookalikes (test_lookalikes_are_not_pii, test_card_needs_a_network_prefix).
  • Typed placeholders (<EMAIL>, <CARD>) keep the shape of the text, so the model still learns that an address goes there without learning anyone’s address (test_hand_example).
  • The audit trail is the spans (kind and offsets into the original text), never the matched text: an audit log that copies the PII is a second copy of it (test_audit_holds_spans_not_text, test_offsets_are_in_the_original_text).
  • The policy is per source, from the ledger: scrub, drop the document, or skip detection for a source known to be clean (test_policies).
Terminal window
ol start data.05 # stubs pii.py into python/corpus/
ol tests data.05 # the course test catalog
ol tdd red data.05 # rung R3: your tests first, failing against the stubs
ol check data.05 # exit code is the verdict
ol check data.05 --ref-deps # only if data.02 or data.04 is not passing yet
ol mutate data.05 # how many planted bugs your tests catch
ol diff data.05 # after passing: your code against the reference

After data.04 your corpus is deduplicated and decontaminated, and it is about to be written to shards that a model will read thousands of times. Some documents still hold an email address, a phone number, a card number pasted into a forum post, a server’s IP address, or an API key someone committed by mistake. A language model memorizes rare strings that appear verbatim, and a generation can repeat them to anyone who asks the right prefix. Your engine will serve that model through your gateway, so a key in the training data is a key in your API’s output. This module replaces each match with a typed placeholder before the text is sharded, counts what it replaced per document and per kind for the manifest and the datasheet, and records where, without recording what.

SymbolMeaningType / shape
TTa document’s text, indexed by code pointstr
[s,e)[s, e)a span: the code points ss to e−1e - 1 of TTPiiSpan(kind, start, end)
kindone of email, phone, card, ip, keystr
d1…dmd_1 \dots d_mthe decimal digits of a number, dmd_m the rightmostdigits
LLthe Luhn sum (below)int

PII here means five machine-findable kinds. Email addresses, phone numbers, payment card numbers, IP addresses, and API keys. Names and street addresses are personal data too, but finding them needs a named-entity model with its own error rates; the datasheet (data.08) says plainly that they are not scrubbed. ethics.02 is where you decide whether that is acceptable for your corpus.

Candidate, then validate. A regex alone is either too loose (every 16-digit number is a card) or too strict (misses real formats). Each detector here is a permissive regex followed by a check that the match is real:

  • email: a local part of [A-Za-z0-9._%+-], @, dot-separated labels, and a last label of at least two letters, so react@18.3.1 (an npm version spec) is not an address and root@localhost (no top-level domain) is not either.
  • phone: North American layouts (212) 555-0123, 212-555-0123, 212.555.0123, +1 212 555 0123, or + and 8 to 15 digits in groups (the E.164 limit). A bare 7-digit 555-0123 is too ambiguous to call a phone.
  • card: 13 to 19 digits, single spaces or dashes allowed between them, first digit 2 to 6 (the card networks’ ranges), and a valid Luhn checksum.
  • ip: an IPv4 dotted quad with every octet at most 255, not part of a longer dotted run (1.2.3.4.5, v1.2.3.4 are versions); an IPv6 candidate that ipaddress.IPv6Address accepts, with at least two non-empty groups and four hex digits (so the Python slice a[1::2] and the loopback ::1, which is not personal, stay).
  • key: provider shapes (sk-..., AKIA + 16, ghp_ + 36, xoxb-..., AIza + 35) and your gateway’s own tl_<id>_<secret> (D19), never inside a longer word.

The Luhn check. Card numbers carry a check digit. Starting from the rightmost digit, double every second digit (dm−1,dm−3,…d_{m-1}, d_{m-3}, \dots); when a doubled digit exceeds 9, subtract 9 (the same as adding its two digits); sum everything into LL. The number is valid when L mod 10=0L \bmod 10 = 0. A random number passes with probability 1/101/10, which is why the network prefix matters: an ISBN-13 starts with 978 or 979 and a millisecond timestamp with 1, and about one in ten of them passes Luhn.

Overlaps resolve left to right. Detectors run independently, so matches overlap: 1.2.3.4@example.com is an IPv4 and an email, +4222222222222 is a phone and a card. Sort candidates by start, then longer first, then by the order of KINDS, and keep each candidate that starts at or after the end of the last kept one. The result is sorted and non-overlapping, which redact requires.

Typed placeholders. redact replaces each span by <EMAIL>, <PHONE>, <CARD>, <IP>, or <KEY>. A placeholder keeps the sentence grammatical and tells the model what kind of thing stood there. No placeholder matches any detector, so scrubbing twice changes nothing: a retried activity (data.09) is safe.

Audit spans. scrub(doc) returns the new document and the spans, with offsets into the original text. Offsets into the redacted text would shift after every placeholder of a different length and point at the wrong characters. audit_line(id, spans) writes {"id", "spans": [{"kind", "start", "end"}]}: enough to count, sample, and review a decision against the source later, and nothing that is itself personal data.

Per-source policy. The ledger row of every source carries pii_policy (ethics.02 fills it): scrub (default) redacts; drop removes any document that has a span (for sources where PII signals something you do not want, such as a scraped contact page), keeping clean ones; none skips detection for a source known to be clean, such as your own synthetic data. Every yielded document carries meta["pii_redactions"] and meta["pii"] (a count per kind, zeros included), which data.06 sums into the manifest under the names of MANIFEST_KEYS (manifest_counts does the renaming).

The text (70 code points):

Mail ana@example.com or call (212) 555-0123; card 4111 1111 1111 1111.
0 5 20 29 43 50 69

Candidates. email ana@example.com at [5,20)[5, 20): local part ana, domain labels example and com (three letters). phone (212) 555-0123 at [29,43)[29, 43): the North American layout with parentheses. card 4111 1111 1111 1111 at [50,69)[50, 69): 16 digits with spaces. Nothing overlaps.

Luhn on 4111111111111111. Number the digits from the right, starting at 1. The even positions (2, 4, …, 16) are doubled: seven of them are 1 (doubled to 2) and position 16 is the leading 4 (doubled to 8), so 7⋅2+8=227 \cdot 2 + 8 = 22. The eight odd positions are all 1: 8. L=22+8=30L = 22 + 8 = 30, a multiple of 10: valid, and the first digit 4 is a card network’s. The card stands.

Redact. Replace right to left or rebuild left to right; either way the result is

Mail <EMAIL> or call <PHONE>; card <CARD>.

Counts and audit. meta["pii_redactions"] = 3, meta["pii"] = {"email": 1, "phone": 1, "card": 1, "ip": 0, "key": 0}, and the audit line is {"id": "web:7", "spans": [{"kind": "email", "start": 5, "end": 20}, {"kind": "phone", "start": 29, "end": 43}, {"kind": "card", "start": 50, "end": 69}]}: offsets into the line above, not into the redacted text, where the phone would start at 21.

This is test_hand_example; the same text drives test_audit_holds_spans_not_text.

# python/corpus/pii.py (the full contract is contracts/py/corpus/pii.pyi)
KINDS: tuple[str, ...] # ("email", "phone", "card", "ip", "key")
PLACEHOLDERS: dict[str, str] # kind -> "<EMAIL>", ...
MANIFEST_KEYS: dict[str, str] # kind -> "emails", ...
@dataclass(frozen=True, slots=True)
class PiiSpan:
kind: str
start: int
end: int
def luhn_ok(digits: str) -> bool
def detect(text: str) -> list[PiiSpan] # sorted, non-overlapping
def redact(text: str, spans: Sequence[PiiSpan]) -> str
def scrub(doc: Doc) -> tuple[Doc, list[PiiSpan]]
def audit_line(doc_id: str, spans: Sequence[PiiSpan]) -> str
def scrub_stage(docs, *, policy=None, audit=None) -> Iterator[Doc] # a Stage
def manifest_counts(counts: Mapping[str, int] | None) -> dict[str, int] # meta["pii"] under the manifest's names
TestKINDChecksWhy it matters downstream
test_hand_exampleunitthe section 3 spans, Luhn sum, redacted text, and countsyou and the tests agree on the definitions
test_luhnunittextbook valid and invalid numbers, odd lengths, empty and non-digit inputthe validator that keeps cards from being any number
test_labelled_fixtureconformancerecall at least 0.98 on email and card, precision at least 0.95 overall, span by spanthe design’s bar on a labelled set
test_lookalikes_are_not_piiboundaryISBNs, versions, dates, times, UUIDs, commit hashes, scopes, slices, C symbols: no spantechnical text survives intact
test_card_needs_a_network_prefixboundarya Luhn-valid 978… is not a card, the same body behind a 4 is; Amex groupingone random number in ten passes Luhn
test_phonesunitfour North American layouts and international forms; 7 digits and 19 digits are notthe layouts people write
test_ipsunitIPv4 octet range and dotted runs; valid IPv6 forms; ::1, slices, and times are notaddresses, not syntax
test_keysunitevery key shape whole; short and embedded lookalikes are notthe costliest leak in a code corpus
test_overlaps_resolve_left_to_rightboundaryearliest start wins, then the longer; spans never overlapredact gets a valid span list
test_offsets_are_in_the_original_textunitaudit offsets slice the original textthe audit can be checked against the source
test_redact_rejects_bad_spansboundaryunsorted, overlapping, empty, out-of-range spans raise; the tables are as contractedcorruption fails loudly
test_scrub_keeps_the_document_and_upstream_metaunitid, source, minhash_cluster and other meta survive; per-kind counts, zeros includeddata.06 sums them
test_scrub_is_idempotentpropertya second scrub changes nothing and counts 0retried activities are safe
test_policiesunitscrub, drop, none per source; audit called before the drop decision; unknown policy raisesethics.02’s policy is enforced
test_audit_holds_spans_not_textunitthe audit line has kinds and offsets, not the address or the cardthe log is not a second copy of the PII
test_stage_streamspropertypulls one document at a time from an endless streamthe stage runs on any corpus size

Your tests (rung R3, red then green). Write them under python/tests/data-05-pii/ against the stubs first (ol tdd red data.05 must fail), then make them pass: the hand example, the Luhn textbook numbers, offsets into the original, lookalikes that must stay, each kind found once, overlapping candidates giving one span, counts and meta that survive a second scrub, redact’s checks, and the three policies with the audit hook. ol mutate data.05 grades them: 0.70 of the mutants, including the one behind Pitfall 4.

PitfallSymptomCaught by
1. a regex with no validatorversion strings, IPv4-looking runs, and times become <IP>; npm specs become <EMAIL>test_lookalikes_are_not_pii, test_ips (mutants s04, s05, s07)
2. Luhn alone for cardsISBN-13s and timestamps that pass Luhn become <CARD>test_card_needs_a_network_prefix (mutant s03)
3. keeping every candidatetwo placeholders for one region; redact raises or splicestest_overlaps_resolve_left_to_right (mutant s09)
4. offsets into the redacted textthe audit points at the wrong characters after the first placeholdertest_offsets_are_in_the_original_text (mutant s10)
5. ignoring the source’s policya drop source’s documents are kept, a none source is scanned and changedtest_policies (mutants s12, s13)
6. handing the text to the audit hookthe audit log holds the PII it was meant to account fortest_policies (mutant s14), test_audit_holds_spans_not_text
7. doubling from the left, or not subtracting 9half of the real cards fail Luhntest_luhn, test_hand_example (mutants s01, s02)
8. replacing meta instead of adding to itminhash_cluster and the license vanish before the shardstest_scrub_keeps_the_document_and_upstream_meta (mutant s11)
DirectionModuleHow it uses this
Backdata.02Doc, and the stage shape scrub_stage follows
Backdata.04the scrub runs on the near-dedup survivors and keeps their cluster tags
Backethics.02the policy: which kinds, which placeholders, which sources scrub, drop, or skip
Forwarddata.06pii_redactions becomes a shard column; meta["pii"] sums into the manifest’s pii object
Forwarddata.08the datasheet’s Composition section reports the counts by kind
Forwardgw.08the gateway redacts request logs with a Go port of these detectors (a port, not a call)

If you skip this module, ol check data.06 stops with data.06 needs data.05: build it, or rerun with --ref-deps.

Your pieceProduction equivalentWhat it addsWhere to look
detectMicrosoft Presidiorecognizer registry, context words that raise confidence, named-entity models for names and addresses, per-language recognizerspresidio-analyzer/presidio_analyzer/predefined_recognizers/
scrub_stagedatatrove PIIFormatter, BigCode’s PII pipelinecorpus-scale replacement with consistent fake values, a trained key detector for codedatatrove src/datatrove/pipeline/formatters/pii.py; bigcode-project/bigcode-dataset/pii
key shapesGitHub secret scanning, gitleakshundreds of provider patterns with entropy checks and verification against the providergitleaks/config/gitleaks.toml
audit_linedata protection impact assessmentsthe review process the audit spans feed: sampling, appeal, retention limitsethics.02 going further