Bias and Safety Evals
Overview
Section titled “Overview”- Primary references: Parrish et al., BBQ: A Hand-Built Bias Benchmark for Question Answering (free); Gehman et al., RealToxicityPrompts (free)
- Supplementary: Liang et al., Holistic Evaluation of Language Models (HELM) (free); Nangia et al., CrowS-Pairs (free); Weidinger et al., Ethical and social risks of harm from Language Models (free)
- Prerequisites: the model-evaluation harness (
L6.7), Probability & Statistics (confidence intervals), Model and Data Cards - Estimated time: 1 week in course Pass 9
Key Takeaways
Section titled “Key Takeaways”- Bias is measured as a difference: the same prompt template with only a group term changed, scored on the same metric, with a confidence interval on the gap.
- Safety evals need a scorer you trust: a deterministic one (a lexicon, a classifier with known precision) before an LLM judge.
- Small models fail differently, not less. A 10M-parameter story model will not produce dangerous instructions, but it will reproduce stereotypes from its corpus; measure what your model can actually do.
- An eval is a test: fixed prompts, fixed seeds, a threshold, and a place in the release gate.
How to Study
Section titled “How to Study”Read BBQ and RealToxicityPrompts, then the HELM sections on bias and toxicity metrics. In the course, ethics.04 is a build module: you write paired-template bias evals and a toxicity scorer in Python, run them through EvalSuite, and report the results in the model card.
Concepts & Techniques
Section titled “Concepts & Techniques”Core Insight
Section titled “Core Insight”A fairness or safety claim is only as good as its measurement. Paired templates isolate the effect of one attribute; bootstrap intervals separate a real gap from noise; and running the suite on every release candidate turns a one-time audit into a regression test.
1. Measuring bias
Section titled “1. Measuring bias”Key ideas:
- Paired templates: identical prompts differing in one group term; compare log-likelihoods or completions.
- Uncertainty: a paired bootstrap interval on the difference, as in
M07.5; report the interval, not just the point estimate.
2. Measuring safety
Section titled “2. Measuring safety”Key ideas:
- Toxicity and refusal rates on fixed prompt sets with seeded sampling.
- Where it runs: the
EvalSuiteworkflow (dur.11) and the release gate (dur.12); results land in the model card (ethics.03).
Course modules
Section titled “Course modules”| Module | Topic | Kind | Pass |
|---|---|---|---|
ethics.04 | Bias and safety evals | build | 9 |
Chapters
Section titled “Chapters”| # | Module | Chapter | Kind | Pass |
|---|---|---|---|---|
| 1 | ethics.04 | Bias and safety evals | build | 9 |
Connections to Other Tracks
Section titled “Connections to Other Tracks”| Track | Connection |
|---|---|
| Responsible AI | the track overview and how the six topics connect |
| LLM Evaluation | eval runners, judges, and A/B experiments |
| Probability & Statistics | bootstrap intervals and paired tests |
| Usage Policy | what you do about the failures you find |