LLN, CLT, confidence intervals, bootstrap
Overview
Section titled “Overview”| Module | M07.4 · build · Python · Pass 4 · 4 to 5 h |
| You build | python/tinyllm/prob/stats.py: normal_cdf, normal_ppf, t_cdf, t_ppf, standard_error, mean_ci, wilson_interval, quantile, bootstrap_ci |
| Contract | course/contracts/py/tinyllm/prob/stats.pyi · the generator behind rng: spec/pcg32.md |
| Tests | course/tests/M07.4/ (what they check: section 4), golden values from scipy 1.17.1 in course/fixtures/M07.4/scipy_golden.json |
| Needs | no code from earlier modules · reading: M07.1 one uniform per draw, M07.0 expectation, variance, the normal, S-M07b the variance of a sum |
| Used by | later: L4.5 sequence metrics with bootstrap CIs · L6.7 every evaluation metric carries a CI · L8.5 quantization verdicts · L10.7 · ag.12 re-implements the bootstrap in Go · M07.5 hypothesis tests build on it · later: L12.3, ethics.04 |
| Milestone | MS-P4 (sequence models; the pass gate runs ol check on every math module of the pass) |
| Optional depth | Wasserman, All of Statistics, ch. 5 to 8 (convergence, the bootstrap); Efron and Tibshirani, An Introduction to the Bootstrap (1993), ch. 13; Brown, Cai, DasGupta, “Interval Estimation for a Binomial Proportion” (2001) |
Key Takeaways
Section titled “Key Takeaways”- The law of large numbers says a sample mean converges to the true mean; the central limit theorem says the error is close to normal with standard deviation , so quadrupling the data halves the error bar (
test_standard_error_shrinks_like_root_n,test_mean_ci_coverage_skewed_data_by_clt). - When is estimated from the same values, the right quantile is Student’s with degrees of freedom, not the normal one; at the normal quantile covers only about 92% instead of 95% (
test_hand_example_mean_ci,test_mean_ci_coverage_normal_small_n). - For a proportion near 0 or 1 the Wilson interval stays honest where the textbook collapses to a point (
test_hand_example_wilson,test_wilson_matches_scipy). - The percentile bootstrap gives an interval for any statistic, BLEU or a median included, by resampling the data; with one uniform per index it is reproducible across languages (
test_bootstrap_matches_independent_resampling,test_bootstrap_draws_and_determinism).
How to work this chapter
Section titled “How to work this chapter”ol start M07.4 # stubs stats.py into your repool tests M07.4 # read the test catalog firstol check M07.4 # exit code is the verdictol diff M07.4 # after passing: your code against the reference1. Why now
Section titled “1. Why now”Pass 4 is the first time you compare models. Your LSTM language model (L3.6) will report a bits-per-character number, your seq2seq model (L4.1) a BLEU score (L4.5), and the question is never “what is the number” but “is model A better than model B, or did I get a lucky test set”. A BLEU of 21.3 against 20.8 on 200 sentences means nothing until you know how much BLEU moves when you draw another 200 sentences. From here on every metric the course prints (L4.5, L6.7, the quantization verdicts of L8.5) carries a 95% interval, and the comparison tests of M07.5 build on the same machinery. This module gives you that interval three ways: from a formula for means, from a formula for proportions, and by resampling for everything else.
2. Principles
Section titled “2. Principles”| Symbol | Meaning | Type / shape |
|---|---|---|
| independent draws from one distribution (per-example scores) | reals | |
| the true mean and variance of that distribution | reals | |
| the sample mean | real | |
| the sample variance (Bessel’s ) | real | |
| the standard error: estimated sd of | real | |
| the miss rate; is the confidence level (0.95) | real in | |
| , | the standard normal CDF and its quantile | functions |
| , , | Student’s t with degrees of freedom, its CDF and quantile | functions |
| successes, trials, observed proportion | integers, real | |
| , | number of bootstrap resamples, the statistic on resample | integer, reals |
The law of large numbers. By linearity , and because independent variances add, (S-M07b q3 with all covariances 0). Chebyshev’s inequality then gives : the mean of more data is closer to the truth (you prove this in S-M07c q5). It also gives the rate: the typical error is , so four times the eval set halves the error, it does not quarter it.
The central limit theorem. Whatever the distribution of each (finite variance is all it needs), the standardized mean tends to the standard normal . So , and choosing (1.96 for 95%) turns that into an interval: covers in about of repeated samples. Per-example eval scores are never normal (0/1 accuracy, skewed losses), and the interval still works once is in the hundreds: that is the CLT.
Quantiles from first principles. , with math.erfc (it keeps full relative precision in the lower tail, where computed by subtraction would lose every digit). The quantile has no closed form, but is increasing, so bisection finds it: keep a bracket with and halve it until the midpoint equals an end, which is the last bit of a double. For use the symmetry ; is exact in float64 there.
Student’s t. The interval above needs , and you only have , computed from the same data. The ratio is wider-tailed than the normal (when happens to be small, is large), and for normal data it follows Student’s t with degrees of freedom exactly (Gosset, 1908). So the interval is
Two details carry it: the in (the deviations are measured from , which sits closer to the data than does, so dividing by underestimates ), and (the interval misses on both sides). For integer the CDF has a closed form (Abramowitz and Stegun 26.7.3 and 26.7.4). With and :
- even: ,
- odd: , just for ,
and . Each term is the previous one times a ratio ( or ), so a cumulative product computes the series in . The quantile is bisection again, on a bracket whose top doubles until ( needs at , so no fixed bound is safe). As grows, tends to the normal: against .
Proportions: the Wilson interval. An accuracy is a mean of 0/1 values, but the formula (the Wald interval) fails exactly where safety evals live: with 0 failures in 10 trials it says , certainty. Wilson (1927) inverts the test instead: the interval is every with . Squaring and solving the quadratic in gives
whose center is pulled toward and which is never empty. At its lower end is exactly 0 and at its upper end exactly 1; rounding can leave them a hair off, so the code pins them.
The bootstrap. For a statistic with no variance formula (BLEU, chrF, a median, pass@k), Efron’s idea (1979) is to let the data stand in for the population: draw indices with replacement, recompute the statistic on that resample, repeat times, and read the spread of . The percentile interval is their and quantiles. Two rules make it reproducible across languages (ag.12 re-implements it in Go): each index is from exactly one rng.uniform(), drawn in order (resample 0’s indices first); and the quantile is type 7: sort, , , value (numpy’s default). The point estimate stays on the original data.
3. Worked example by hand
Section titled “3. Worked example by hand”A t interval. , .
- Mean: .
- Deviations: ; squares sum to .
- , ; .
- , : (bisection on the odd series, which for is ).
- Half-width ; interval .
With the normal 1.959964 instead, the interval would be , a quarter narrower: at that is a large overstatement of certainty.
One value of the series. : , , . The odd series with stops after : , so .
A Wilson interval. , , , : , center , half-width . The interval is ; the Wald interval is .
These are test_hand_example_mean_ci and test_hand_example_wilson.
4. The interface
Section titled “4. The interface”def normal_cdf(z: float) -> floatdef normal_ppf(p: float) -> float # bisection, symmetric for p > 1/2def t_cdf(t: float, df: int) -> float # A&S 26.7.3 / 26.7.4def t_ppf(p: float, df: int) -> float # bracket doubled, then bisectiondef standard_error(x) -> float # s / sqrt(n), divisor n - 1def mean_ci(x, alpha: float = 0.05) -> tuple[float, float, float] # (mean, lo, hi)def wilson_interval(k: int, n: int, alpha: float = 0.05) -> tuple[float, float]def quantile(x, q: float) -> float # type 7def bootstrap_ci(x, stat, n_boot: int, alpha: float, rng) -> tuple[float, float, float] # (stat(x), lo, hi)What the tests check
Section titled “What the tests check”| Test | KIND | Checks | Why it matters downstream |
|---|---|---|---|
test_hand_example_mean_ci | unit | section 3: se , interval , | you and the test agree on Bessel, , and |
test_hand_example_wilson | unit | section 3: at | safety failure rates (ethics.04) |
test_normal_matches_scipy | golden | and to 1e-11 relative, tails to | every in the course |
test_t_matches_scipy | golden | and for 13 values of from 1 to 1000 | small eval sets get the right width |
test_mean_ci_matches_scipy | golden | standard_error and mean_ci equal scipy.stats.sem and t.interval | L6.7 reports |
test_wilson_matches_scipy | golden | scipy’s Wilson interval at the extremes and in the middle | accuracy CIs |
test_quantile_matches_numpy_type7 | golden | type 7 quantiles | Go’s bootstrap (ag.12) agrees |
test_bootstrap_matches_independent_resampling | golden | same seed, same interval as an independent implementation, mean and median | reproducible CIs for BLEU (L4.5) |
test_mean_ci_coverage_normal_small_n | statistical | 2000 samples of : coverage in | what “95%” promises |
test_mean_ci_coverage_skewed_data_by_clt | statistical | exponential data, : coverage in | the CLT on non-normal scores |
test_standard_error_shrinks_like_root_n | property | four copies of the data halve the standard error | sizing eval sets |
test_ppf_inverts_cdf_and_t_tends_to_normal | property | , symmetry, normal | the quantiles are consistent |
test_t_closed_forms | unit | (Cauchy) and closed forms | both series are right at their shortest |
test_bootstrap_draws_and_determinism | unit | uniforms; same seed same result; point is | cross-language reproducibility |
test_bootstrap_index_rule | boundary | , is index 2, is the last index | the floor rule |
test_bootstrap_shift_equivariance | property | moves the mean’s interval by | resampling does not look at values |
test_quantile_edges | boundary | , , one value, unsorted input | no index past the end |
test_wilson_stays_inside_zero_one | boundary | exact 0 and 1 at the ends; inside | intervals stay probabilities |
test_rejects_bad_arguments | boundary | , , , NaN, empty input | errors at the call site |
5. Pitfalls
Section titled “5. Pitfalls”| Pitfall | Symptom | Caught by |
|---|---|---|
| 1. the normal quantile with an estimated | at the “95%” interval covers 92% | test_hand_example_mean_ci, test_mean_ci_coverage_normal_small_n (mutant s01) |
| 2. dividing by instead of , or by instead of | intervals too narrow, by a little or by a lot | test_hand_example_mean_ci, test_standard_error_shrinks_like_root_n (mutants s02, s11) |
| 3. the quantile at instead of | a 90% interval labelled 95% | test_hand_example_mean_ci (mutant s03) |
| 4. the Wald interval for a proportion, or unpinned ends | after zero failures; ends a hair outside | test_hand_example_wilson, test_wilson_stays_inside_zero_one (mutants s04, s13) |
| 5. rounding to pick a bootstrap index | the last row drawn twice as often as the first; Go and Python disagree | test_bootstrap_index_rule (mutant s05) |
| 6. drawing one resample and reusing it | every replicate identical: a zero-width interval | test_bootstrap_matches_independent_resampling (mutant s06) |
6. Where it’s used next
Section titled “6. Where it’s used next”| Forward | L12.3 | Registered module relationship. |
| Forward | ethics.04 | Registered call site uses this module. |
| Direction | Module | How it uses this |
|---|---|---|
| Back | M07.1 | one uniform per draw, the same discipline as the bootstrap’s index rule |
| Back | M07.0 | expectation, variance, the normal distribution |
| Back | S-M07b | the variance of a sum, which is why the standard error is |
| Forward | S-M07c | the LLN proof, the CLT approximation, and these intervals by hand |
| Forward | L4.5 | BLEU and chrF with bootstrap intervals over sentences |
| Forward | L6.7 | every evaluation metric carries mean_ci or bootstrap_ci |
| Forward | L8.5 | quantization verdicts: is the perplexity change inside the noise |
| Forward | M07.5 | permutation tests and McNemar build on the same resampling |
| Forward | ag.12 | the Go eval harness re-implements bootstrap_ci from the same seed |
If you skip this module, ol check L6.7 stops with L6.7 needs M07.4: build it, or rerun with --ref-deps.
Going further
Section titled “Going further”| Your piece | Production equivalent | What it adds | Where to look |
|---|---|---|---|
t_ppf, normal_ppf | scipy’s stdtrit and ndtri (Cephes) | rational approximations plus Newton steps, for any | scipy/special/cephes/stdtr.c, ndtri.c |
bootstrap_ci (percentile) | scipy.stats.bootstrap (BCa) | bias-corrected and accelerated intervals, vectorized resampling | scipy/stats/_resampling.py |
wilson_interval | statsmodels.stats.proportion.proportion_confint | Agresti-Coull, Jeffreys, Clopper-Pearson side by side | statsmodels/stats/proportion.py |
| interval on eval scores | HELM and lm-evaluation-harness stderr | clustered standard errors when examples share a prompt | lm_eval/api/metrics.py (bootstrap_stderr) |