Skip to content

Privacy and PII

  • Language models memorize, and repeated or rare strings (emails, phone numbers, keys) are the easiest to extract. What is in the corpus can come out of the model.
  • Scrub before training, redact before logging. The corpus pipeline replaces PII with typed placeholders; the gateway applies the same detector list to its logs.
  • A detector is a classifier with a precision and a recall. Measure both on labelled fixtures, and test the lookalikes (ISBNs, version strings) that cause false positives.
  • Data minimization is the cheapest control: what you never collect or keep cannot leak or need erasing.

Read Carlini et al. sections 1 to 5, then the NIST Privacy Framework core. In the course, ethics.02 writes your PII policy (what you detect, replace, log, and keep); data.05 implements the scrub and is graded on recall and precision against it.


Privacy in an ML system is a data-flow property: personal data enters through the corpus and through user requests, and leaves through model outputs and logs. A policy names each flow and its control; the pipeline and the gateway enforce it; tests on labelled fixtures show it works.

Key ideas:

  • In: crawled text, user prompts, uploaded documents. Out: generations, logs, traces, eval reports.
  • Memorization: duplicated sequences are memorized far more often, which is one more reason deduplication (data.03, data.04) runs before training.

Key ideas:

  • Typed placeholders (<EMAIL>, <PHONE>, <CARD>) keep text shape for training while removing the value; Luhn checks separate card numbers from other digit runs.
  • Audit spans record what was replaced and where, without storing the value.
  • Gateway logs reuse the detector list (gw.08), so the same policy covers training data and traffic.
ModuleTopicKindPass
ethics.02Privacy and PII policypractice3
#ModuleChapterKindPass
1ethics.02Privacy and PII policypractice3
TrackConnection
Responsible AIthe track overview and how the six topics connect
Corpus Pipelinethe PII scrub (data.05)
Observabilitywhat traces and logs may contain
Securitysecrets review and the threat model