Solve set: LLN/CLT, confidence intervals
Overview
Section titled “Overview”| Module | S-M07c · solve · none · Pass 4 · 2 to 3 h |
| You build | answers in solve/S-M07c.toml (9 checked by SymPy) and 1 proof in solve/S-M07c/q5.md (self-graded against its rubric) |
| Contract | none: a pen and paper set |
| Tests | course/solve/S-M07c/key.toml (hidden): typed answers plus reject canaries; the problems are in course/solve/S-M07c/problems.md and in section 4 |
| Needs | S-M07b (the variance of a sum). Reading: M07.4 LLN, CLT, confidence intervals, bootstrap and the Probability and Statistics topic, limit theorems and estimation sections |
| Used by | no call site (a solve set). Do it alongside M07.4: q3, q6, q7, and q8 are its functions by hand, q5 is why its intervals shrink; part S-M07d follows with hypothesis tests |
| Milestone | MS-P4 (the Pass 4 gate runs ol check on every solve part of the pass) |
| Optional depth | Blitzstein and Hwang, Introduction to Probability (free), ch. 10 “Inequalities and limit theorems”; Wasserman, All of Statistics, ch. 5 to 8 |
Key Takeaways
Section titled “Key Takeaways”- The mean of independent draws has the same mean and of the variance, so its error shrinks like ; halving an error bar costs four times the data (q1, q4).
- Chebyshev’s inequality turns that variance into the weak law of large numbers in three lines (q2, q5).
- The central limit theorem lets a normal table answer questions about sums of non-normal variables, such as coin flips (q3).
- A confidence interval is a statement about the procedure, not about one interval: 95% of the intervals it would produce cover the true value (q9).
- Student’s t, Wilson, and the type 7 percentile rule are each a few lines of arithmetic, the same numbers
M07.4computes (q6, q7, q8).
How to work this chapter
Section titled “How to work this chapter”ol start S-M07c # writes solve/S-M07c.toml and the proof fileol check S-M07c # SymPy checks the answers, then asks the proof rubric (y/n)ol check S-M07c --regrade # ask the rubric again after you change the proof1. Why now
Section titled “1. Why now”M07.4 puts an interval on every number your models report from Pass 4 on: the bits per character of L3.6, the BLEU of L4.4, the accuracies of L6.7. Code computes those intervals without complaint whether or not you understand them, and a misread interval is worse than none: “95% sure model B is better” is not what a 95% interval says. This set makes the three ideas concrete by hand, so that the code is a transcription: why averaging works (the law of large numbers), why the error is normal and shrinks like (the central limit theorem), and what exactly an interval promises.
2. Principles
Section titled “2. Principles”| Symbol | Meaning | Type / shape |
|---|---|---|
independent draws, each with mean and variance (mu, s2) | reals | |
| the sample mean | real | |
(eps) | a tolerance, | real |
| the standard normal CDF | function | |
| sample standard deviation, divisor | real | |
| Student’s t quantile with degrees of freedom | real | |
| successes, trials, , the normal quantile | integers, reals | |
| sorted bootstrap replicates | reals |
2.1 Averages
Section titled “2.1 Averages”Linearity of expectation gives . Independence makes every covariance in vanish (S-M07b q3), so the variance of the sum is , and dividing by divides the variance by : . The standard deviation of the mean is .
2.2 Chebyshev and the law of large numbers
Section titled “2.2 Chebyshev and the law of large numbers”Markov’s inequality, for and , holds because . Applied to with it becomes Chebyshev’s inequality, . With the bound shrinks to 0 as grows: that is the weak law of large numbers.
2.3 The central limit theorem
Section titled “2.3 The central limit theorem”For large , is approximately standard normal whatever the distribution of each . Equivalently a sum is approximately normal with mean and variance , so .
2.4 Intervals
Section titled “2.4 Intervals”A 95% confidence interval is a procedure: data in, interval out, built so that over repeated samples 95% of the intervals contain the fixed, unknown . Once computed, a particular interval either contains or not; the 95% belongs to the procedure. The three procedures of M07.4:
- Student t: .
- Wilson for a proportion: .
- Percentile bootstrap: the type 7 quantiles at and of the replicates: , , value .
3. Worked example by hand
Section titled “3. Worked example by hand”These are siblings of q3, q6, and q8, not graded problems.
A CLT tail. 400 fair coin flips; . The count has mean and variance , sd 10, so and the probability is about . Twenty extra heads out of 400 is as surprising as ten out of 100: the sd grows like , not .
A t interval. , , , : the half-width is , so the interval is . In solve/ this would be answer = "[0.869, 5.131]".
A percentile interval. Five replicates and : , , value ; : , value . The 50% interval is .
4. The problem set
Section titled “4. The problem set”Write each answer in solve/S-M07c.toml; lettered parts are their own tables:
[q1.a]answer = "mu"[q3]answer = "0.02275"[q6]answer = "[9.1744, 10.8256]"[q9]answer = "c"[q5]proof = "S-M07c/q5.md"Write as s2, as eps, as mu. An interval is written [lo, hi]; exact fractions such as 2/7 are welcome.
The law of large numbers and the CLT
Section titled “The law of large numbers and the CLT”q1. are independent, each with mean and variance , and . (a) . [expr in mu] (b) . [expr in s2, n]
q2. For the same and any , Chebyshev’s inequality bounds . Give the bound. [expr in s2, n, eps]
q3. You flip a fair coin 100 times. Using the central limit theorem without a continuity correction, and for the standard normal CDF, approximate . [number, 3 significant digits]
q4. A 95% interval for an eval score is . By what factor must grow to halve its width (with unchanged)? [number, exact]
q5. Prove the weak law of large numbers for finite variance: if are independent with mean and variance , then for every , as . [proof]
Confidence intervals
Section titled “Confidence intervals”q6. An eval of prompts has mean score and sample standard deviation (divisor ). With , give the 95% Student t interval for the true mean. [interval, exact decimals]
q7. A safety check fails on 0 of prompts. Give the Wilson score interval for the failure rate with :
[interval, exact]
q8. Nine bootstrap replicates of a statistic, sorted, are . With the type 7 quantile (for sorted and level : , , value ), give the 80% percentile interval . [interval, exact]
q9. From one sample you compute the 95% confidence interval for a mean . Which statement is correct? [choice]
(a) , with random.
(b) 95% of the observations lie in .
(c) The procedure that produced it covers the true in 95% of repeated samples.
(d) A new sample’s mean lands in with probability 0.95.
5. Pitfalls
Section titled “5. Pitfalls”| Pitfall | Symptom | Caught by |
|---|---|---|
| Forgetting that is squared in a variance | the variance of a mean reported as or | q1 (canaries s2 and s2/n^2) |
| Confusing the variance with the standard error | where a variance is asked | q1 (canary sqrt(s2/n)) |
| Chebyshev on one draw instead of the mean | a bound that does not shrink with | q2 (canary s2/eps^2) |
| A two-sided tail for a one-sided question | 0.0455 instead of 0.02275 | q3 (canary) |
| Width proportional to | “double the data, halve the bar” | q4 (canary 2) |
| The normal 1.96 with an estimated at small , or without | intervals too narrow, or 5 times too wide | q6 (canaries) |
| The Wald interval at | : certainty from ten trials | q7 (canary) |
| Nearest-rank instead of interpolated quantiles | a bootstrap interval that disagrees with numpy and Go | q8 (canary [1, 9]) |
| Reading the 95% as a probability about this one interval | “95% chance the mean is in here” | q9 (canary a) |
6. Where it’s used next
Section titled “6. Where it’s used next”| Direction | Module | How it uses this |
|---|---|---|
| Back | S-M07b | the variance of a sum with covariances, q3 there |
| Back | M07.4 | mean_ci, wilson_interval, quantile, and bootstrap_ci are q6, q7, and q8 in code (reading) |
| Forward | S-M07d | hypothesis tests: the same sampling distributions, read as p-values |
| Forward | L6.7 | every metric with its interval, and what the interval does and does not claim |
| Forward | L4.5 | bootstrap intervals over sentences for BLEU |