Skip to content

Solve set: LLN/CLT, confidence intervals

ModuleS-M07c · solve · none · Pass 4 · 2 to 3 h
You buildanswers in solve/S-M07c.toml (9 checked by SymPy) and 1 proof in solve/S-M07c/q5.md (self-graded against its rubric)
Contractnone: a pen and paper set
Testscourse/solve/S-M07c/key.toml (hidden): typed answers plus reject canaries; the problems are in course/solve/S-M07c/problems.md and in section 4
NeedsS-M07b (the variance of a sum). Reading: M07.4 LLN, CLT, confidence intervals, bootstrap and the Probability and Statistics topic, limit theorems and estimation sections
Used byno call site (a solve set). Do it alongside M07.4: q3, q6, q7, and q8 are its functions by hand, q5 is why its intervals shrink; part S-M07d follows with hypothesis tests
MilestoneMS-P4 (the Pass 4 gate runs ol check on every solve part of the pass)
Optional depthBlitzstein and Hwang, Introduction to Probability (free), ch. 10 “Inequalities and limit theorems”; Wasserman, All of Statistics, ch. 5 to 8
  • The mean of nn independent draws has the same mean and 1/n1/n of the variance, so its error shrinks like 1/n1/\sqrt{n}; halving an error bar costs four times the data (q1, q4).
  • Chebyshev’s inequality turns that variance into the weak law of large numbers in three lines (q2, q5).
  • The central limit theorem lets a normal table answer questions about sums of non-normal variables, such as coin flips (q3).
  • A confidence interval is a statement about the procedure, not about one interval: 95% of the intervals it would produce cover the true value (q9).
  • Student’s t, Wilson, and the type 7 percentile rule are each a few lines of arithmetic, the same numbers M07.4 computes (q6, q7, q8).
Terminal window
ol start S-M07c # writes solve/S-M07c.toml and the proof file
ol check S-M07c # SymPy checks the answers, then asks the proof rubric (y/n)
ol check S-M07c --regrade # ask the rubric again after you change the proof

M07.4 puts an interval on every number your models report from Pass 4 on: the bits per character of L3.6, the BLEU of L4.4, the accuracies of L6.7. Code computes those intervals without complaint whether or not you understand them, and a misread interval is worse than none: “95% sure model B is better” is not what a 95% interval says. This set makes the three ideas concrete by hand, so that the code is a transcription: why averaging works (the law of large numbers), why the error is normal and shrinks like 1/n1/\sqrt{n} (the central limit theorem), and what exactly an interval promises.

SymbolMeaningType / shape
X1,…,XnX_1, \dots, X_nindependent draws, each with mean μ\mu and variance σ2\sigma^2 (mu, s2)reals
Xˉn\bar X_nthe sample mean (X1+⋯+Xn)/n(X_1 + \dots + X_n)/nreal
ε\varepsilon (eps)a tolerance, ε>0\varepsilon > 0real
Φ\Phithe standard normal CDFfunction
sssample standard deviation, divisor n−1n - 1real
t1−α/2, νt_{1-\alpha/2,\,\nu}Student’s t quantile with ν\nu degrees of freedomreal
k,n,p^,zk, n, \hat p, zsuccesses, trials, k/nk/n, the normal quantileintegers, reals
s0≤⋯≤sB−1s_0 \le \dots \le s_{B-1}sorted bootstrap replicatesreals

Linearity of expectation gives E[Xˉn]=μE[\bar X_n] = \mu. Independence makes every covariance in Var⁡(∑Xi)\operatorname{Var}(\sum X_i) vanish (S-M07b q3), so the variance of the sum is nσ2n\sigma^2, and dividing by nn divides the variance by n2n^2: Var⁡(Xˉn)=σ2/n\operatorname{Var}(\bar X_n) = \sigma^2/n. The standard deviation of the mean is σ/n\sigma/\sqrt{n}.

2.2 Chebyshev and the law of large numbers

Section titled “2.2 Chebyshev and the law of large numbers”

Markov’s inequality, P(Z≥a)≤E[Z]/aP(Z \ge a) \le E[Z]/a for Z≥0Z \ge 0 and a>0a > 0, holds because a 1[Z≥a]≤Za\,\mathbf{1}[Z \ge a] \le Z. Applied to Z=(Y−EY)2Z = (Y - EY)^2 with a=ε2a = \varepsilon^2 it becomes Chebyshev’s inequality, P(∣Y−EY∣≥ε)≤Var⁡(Y)/ε2P(|Y - EY| \ge \varepsilon) \le \operatorname{Var}(Y)/\varepsilon^2. With Y=XˉnY = \bar X_n the bound shrinks to 0 as nn grows: that is the weak law of large numbers.

For large nn, (Xˉn−μ)/(σ/n)(\bar X_n - \mu)/(\sigma/\sqrt{n}) is approximately standard normal whatever the distribution of each XiX_i. Equivalently a sum Sn=nXˉnS_n = n\bar X_n is approximately normal with mean nμn\mu and variance nσ2n\sigma^2, so P(Sn≥c)≈1−Φ((c−nμ)/nσ2)P(S_n \ge c) \approx 1 - \Phi\bigl((c - n\mu)/\sqrt{n\sigma^2}\bigr).

A 95% confidence interval is a procedure: data in, interval out, built so that over repeated samples 95% of the intervals contain the fixed, unknown μ\mu. Once computed, a particular interval either contains μ\mu or not; the 95% belongs to the procedure. The three procedures of M07.4:

  • Student t: xˉ±t1−α/2, n−1 s/n\bar x \pm t_{1-\alpha/2,\,n-1}\, s/\sqrt{n}.
  • Wilson for a proportion: (p^+z22n±zp^(1−p^)n+z24n2)/(1+z2n)\bigl(\hat p + \frac{z^2}{2n} \pm z\sqrt{\frac{\hat p(1-\hat p)}{n} + \frac{z^2}{4n^2}}\bigr) / (1 + \frac{z^2}{n}).
  • Percentile bootstrap: the type 7 quantiles at α/2\alpha/2 and 1−α/21 - \alpha/2 of the replicates: h=(B−1)qh = (B - 1)q, i=⌊h⌋i = \lfloor h \rfloor, value si+(h−i)(si+1−si)s_i + (h - i)(s_{i+1} - s_i).

These are siblings of q3, q6, and q8, not graded problems.

A CLT tail. 400 fair coin flips; P(at least 220 heads)P(\text{at least } 220 \text{ heads}). The count has mean 400⋅12=200400 \cdot \frac{1}{2} = 200 and variance 400⋅14=100400 \cdot \frac{1}{4} = 100, sd 10, so z=(220−200)/10=2z = (220 - 200)/10 = 2 and the probability is about 1−Φ(2)=1−0.97725=0.022751 - \Phi(2) = 1 - 0.97725 = 0.02275. Twenty extra heads out of 400 is as surprising as ten out of 100: the sd grows like n\sqrt{n}, not nn.

A t interval. n=16n = 16, xˉ=3\bar x = 3, s=4s = 4, t0.975, 15=2.131t_{0.975,\,15} = 2.131: the half-width is 2.131×4/16=2.1312.131 \times 4/\sqrt{16} = 2.131, so the interval is [0.869,5.131][0.869, 5.131]. In solve/ this would be answer = "[0.869, 5.131]".

A percentile interval. Five replicates 2,3,3,7,102, 3, 3, 7, 10 and q=0.25q = 0.25: h=4×0.25=1h = 4 \times 0.25 = 1, i=1i = 1, value s1=3s_1 = 3; q=0.75q = 0.75: h=3h = 3, value s3=7s_3 = 7. The 50% interval is [3,7][3, 7].

Write each answer in solve/S-M07c.toml; lettered parts are their own tables:

[q1.a]
answer = "mu"
[q3]
answer = "0.02275"
[q6]
answer = "[9.1744, 10.8256]"
[q9]
answer = "c"
[q5]
proof = "S-M07c/q5.md"

Write σ2\sigma^2 as s2, ε\varepsilon as eps, μ\mu as mu. An interval is written [lo, hi]; exact fractions such as 2/7 are welcome.

q1. X1,…,XnX_1, \dots, X_n are independent, each with mean μ\mu and variance σ2\sigma^2, and Xˉn=(X1+⋯+Xn)/n\bar X_n = (X_1 + \dots + X_n)/n. (a) E[Xˉn]E[\bar X_n]. [expr in mu] (b) Var⁡(Xˉn)\operatorname{Var}(\bar X_n). [expr in s2, n]

q2. For the same Xˉn\bar X_n and any ε>0\varepsilon > 0, Chebyshev’s inequality P(∣Y−EY∣≥ε)≤Var⁡(Y)/ε2P(|Y - E Y| \ge \varepsilon) \le \operatorname{Var}(Y)/\varepsilon^2 bounds P(∣Xˉn−μ∣≥ε)P(|\bar X_n - \mu| \ge \varepsilon). Give the bound. [expr in s2, n, eps]

q3. You flip a fair coin 100 times. Using the central limit theorem without a continuity correction, and Φ(2)=0.97725\Phi(2) = 0.97725 for the standard normal CDF, approximate P(at least 60 heads)P(\text{at least } 60 \text{ heads}). [number, 3 significant digits]

q4. A 95% interval for an eval score is xˉ±1.96 s/n\bar x \pm 1.96\, s/\sqrt{n}. By what factor must nn grow to halve its width (with ss unchanged)? [number, exact]

q5. Prove the weak law of large numbers for finite variance: if X1,X2,…X_1, X_2, \dots are independent with mean μ\mu and variance σ2<∞\sigma^2 < \infty, then for every ε>0\varepsilon > 0, P(∣Xˉn−μ∣≥ε)→0P(|\bar X_n - \mu| \ge \varepsilon) \to 0 as n→∞n \to \infty. [proof]

q6. An eval of n=25n = 25 prompts has mean score xˉ=10\bar x = 10 and sample standard deviation s=2s = 2 (divisor n−1n - 1). With t0.975, 24=2.064t_{0.975,\,24} = 2.064, give the 95% Student t interval for the true mean. [interval, exact decimals]

q7. A safety check fails on 0 of n=10n = 10 prompts. Give the Wilson score interval for the failure rate with z=2z = 2: p^+z22n±zp^(1−p^)n+z24n21+z2n.\frac{\hat p + \frac{z^2}{2n} \pm z\sqrt{\frac{\hat p(1 - \hat p)}{n} + \frac{z^2}{4n^2}}}{1 + \frac{z^2}{n}}. [interval, exact]

q8. Nine bootstrap replicates of a statistic, sorted, are 1,2,4,5,5,6,8,9,131, 2, 4, 5, 5, 6, 8, 9, 13. With the type 7 quantile (for sorted s0,…,sB−1s_0, \dots, s_{B-1} and level qq: h=(B−1)qh = (B - 1)q, i=⌊h⌋i = \lfloor h \rfloor, value si+(h−i)(si+1−si)s_i + (h - i)(s_{i+1} - s_i)), give the 80% percentile interval [quantile(0.1),quantile(0.9)][\text{quantile}(0.1), \text{quantile}(0.9)]. [interval, exact]

q9. From one sample you compute the 95% confidence interval [3.1,4.7][3.1, 4.7] for a mean μ\mu. Which statement is correct? [choice] (a) P(3.1≤μ≤4.7)=0.95P(3.1 \le \mu \le 4.7) = 0.95, with μ\mu random. (b) 95% of the observations lie in [3.1,4.7][3.1, 4.7]. (c) The procedure that produced it covers the true μ\mu in 95% of repeated samples. (d) A new sample’s mean lands in [3.1,4.7][3.1, 4.7] with probability 0.95.

PitfallSymptomCaught by
Forgetting that 1/n1/n is squared in a variancethe variance of a mean reported as σ2\sigma^2 or σ2/n2\sigma^2/n^2q1 (canaries s2 and s2/n^2)
Confusing the variance with the standard errorσ2/n\sqrt{\sigma^2/n} where a variance is askedq1 (canary sqrt(s2/n))
Chebyshev on one draw instead of the meana bound that does not shrink with nnq2 (canary s2/eps^2)
A two-sided tail for a one-sided question0.0455 instead of 0.02275q3 (canary)
Width proportional to 1/n1/n“double the data, halve the bar”q4 (canary 2)
The normal 1.96 with an estimated ss at small nn, or ss without n\sqrt{n}intervals too narrow, or 5 times too wideq6 (canaries)
The Wald interval at k=0k = 0[0,0][0, 0]: certainty from ten trialsq7 (canary)
Nearest-rank instead of interpolated quantilesa bootstrap interval that disagrees with numpy and Goq8 (canary [1, 9])
Reading the 95% as a probability about this one interval“95% chance the mean is in here”q9 (canary a)
DirectionModuleHow it uses this
BackS-M07bthe variance of a sum with covariances, q3 there
BackM07.4mean_ci, wilson_interval, quantile, and bootstrap_ci are q6, q7, and q8 in code (reading)
ForwardS-M07dhypothesis tests: the same sampling distributions, read as p-values
ForwardL6.7every metric with its interval, and what the interval does and does not claim
ForwardL4.5bootstrap intervals over sentences for BLEU