Skip to content

Bias and Safety Evals

  • Bias is measured as a difference: the same prompt template with only a group term changed, scored on the same metric, with a confidence interval on the gap.
  • Safety evals need a scorer you trust: a deterministic one (a lexicon, a classifier with known precision) before an LLM judge.
  • Small models fail differently, not less. A 10M-parameter story model will not produce dangerous instructions, but it will reproduce stereotypes from its corpus; measure what your model can actually do.
  • An eval is a test: fixed prompts, fixed seeds, a threshold, and a place in the release gate.

Read BBQ and RealToxicityPrompts, then the HELM sections on bias and toxicity metrics. In the course, ethics.04 is a build module: you write paired-template bias evals and a toxicity scorer in Python, run them through EvalSuite, and report the results in the model card.


A fairness or safety claim is only as good as its measurement. Paired templates isolate the effect of one attribute; bootstrap intervals separate a real gap from noise; and running the suite on every release candidate turns a one-time audit into a regression test.

Key ideas:

  • Paired templates: identical prompts differing in one group term; compare log-likelihoods or completions.
  • Uncertainty: a paired bootstrap interval on the difference, as in M07.5; report the interval, not just the point estimate.

Key ideas:

  • Toxicity and refusal rates on fixed prompt sets with seeded sampling.
  • Where it runs: the EvalSuite workflow (dur.11) and the release gate (dur.12); results land in the model card (ethics.03).
ModuleTopicKindPass
ethics.04Bias and safety evalsbuild9
#ModuleChapterKindPass
1ethics.04Bias and safety evalsbuild9
TrackConnection
Responsible AIthe track overview and how the six topics connect
LLM Evaluationeval runners, judges, and A/B experiments
Probability & Statisticsbootstrap intervals and paired tests
Usage Policywhat you do about the failures you find