Skip to content

Logistic regression (IRLS), ROC-AUC, calibration (ECE)

ModuleM07.7 · build · Python · Pass 5 · 3 to 4 h
You buildpython/tinyllm/prob/metrics.py: logistic_regression_fit, logistic_predict_proba, roc_curve, roc_auc, reliability_bins, ece, fit_temperature
Contractcourse/contracts/py/tinyllm/prob/metrics.pyi · the head it fits: formats/linear-head.schema.json
Testscourse/tests/M07.7/test_metrics.py (what they check: section 4), golden values from scipy 1.17.1 in course/fixtures/M07.7/scipy_golden.json
NeedsM03.2 LU (lu, lu_solve: one solve per Newton step) · M01.4 trapezoid (the area under the ROC curve) · M01.2 Newton’s method (newton: the temperature’s root) · reading: M07.5 hypothesis tests
Used bylater: L6.5 fits the usage-policy linear head and reports its AUC and ECE (D33) · L3.5 scores probes · ethics.04 puts both numbers in the safety report
MilestoneMS-P5 (the Pass 5 gate)
Optional depthHastie, Tibshirani, Friedman, The Elements of Statistical Learning (free), section 4.4 (logistic regression and IRLS); Fawcett, “An introduction to ROC analysis” (2006); Guo, Pleiss, Sun, Weinberger, “On Calibration of Modern Neural Networks” (2017)
  • Logistic regression models P(y=1∣x)=σ(w⋅x)P(y = 1 \mid x) = \sigma(w \cdot x); its loss is convex, so the minimum is where the gradient X⊤(p−y)+λwX^\top(p - y) + \lambda w vanishes (test_fit_is_a_stationary_point).
  • Newton’s method on that loss is IRLS: each step is one linear solve with the Hessian X⊤SX+λIX^\top S X + \lambda I, S=diag⁡(p(1−p))S = \operatorname{diag}(p(1-p)), and the digits double per step (test_hand_example_newton_step, test_newton_converges_quadratically).
  • The ROC curve sweeps every threshold; its area, the AUC, is the probability a random positive outscores a random negative, ties counting half (test_hand_example_roc_auc, test_auc_is_mann_whitney).
  • Calibration asks whether “0.8” means right 80% of the time; the ECE is the count-weighted gap between confidence and accuracy over bins (test_hand_example_ece, test_ece_calibrated_vs_overconfident).
  • A ranking metric and a calibration metric answer different questions: sharpening probabilities leaves the AUC unchanged and wrecks the ECE.
Terminal window
ol start M07.7 # stubs metrics.py into your repo
ol tests M07.7 # read the test catalog first: rung R0, you write no tests here
ol check M07.7 # exit code is the verdict
ol check M07.7 --ref-deps # only if you skipped M03.2 or M01.4
ol diff M07.7 # after passing: your code against the reference

Pass 5 gives your system its first classifiers. L6.5 puts a classification head on BERT and ELECTRA, and the gateway’s usage policy (D33) is a linear head over your engine’s embeddings: a dot product and a sigmoid that Go evaluates on every request. That head has to be fitted, and gradient descent with a learning rate to tune is the wrong tool for a convex problem with a few hundred weights: Newton’s method solves it in five or six steps, each step a linear system you already know how to solve (M03.2). Once fitted, the head needs two numbers before anyone trusts it. Does it rank unsafe requests above safe ones (AUC)? And when it says 0.9, is it right 90% of the time, so a threshold of 0.9 means what the operator thinks (ECE)? This module builds the fitter and both metrics.

SymbolMeaningType / shape
XXfeature matrix, one row per examplefloat64[n, d]
XaX_aXX with a column of ones appended (the intercept column, last)float64[n, d+1]
yi∈{0,1}y_i \in \{0, 1\}labelsfloat64[n]
wwweights; with an intercept, wdw_{d} (the last) is the bias bbfloat64[d+1]
zi=xi⋅wz_i = x_i \cdot wthe logit (log-odds) of example iifloat64[n]
σ(z)=1/(1+e−z)\sigma(z) = 1/(1 + e^{-z})the sigmoid
pi=σ(zi)p_i = \sigma(z_i)predicted P(yi=1)P(y_i = 1)float64[n]
λ\lambda (l2)ridge penalty on the feature weightsfloat ≥0\ge 0
gg, HHgradient and Hessian of the loss[d+1], [d+1, d+1]
SSdiag⁡(pi(1−pi))\operatorname{diag}(p_i (1 - p_i))[n, n] (never formed)
TPR, FPRtrue and false positive rates at a threshold
BB (n_bins)number of confidence binsint

A linear score z=w⋅xz = w \cdot x can be any real number; a probability must lie in (0,1)(0, 1). The sigmoid maps one to the other, and its inverse is the log-odds: z=log⁡p1−pz = \log \frac{p}{1 - p}. So logistic regression says the log-odds are linear in the features; a unit step in feature jj multiplies the odds by ewje^{w_j}. Fitting maximizes the likelihood ∏ipiyi(1−pi)1−yi\prod_i p_i^{y_i}(1 - p_i)^{1 - y_i}, or equivalently minimizes

L(w)=−∑i[yilog⁡pi+(1−yi)log⁡(1−pi)]+λ2∥wfeat∥2.L(w) = -\sum_i \bigl[y_i \log p_i + (1 - y_i) \log(1 - p_i)\bigr] + \tfrac{\lambda}{2}\lVert w_{\text{feat}} \rVert^2 .

The derivative of the bracket with respect to ziz_i is just pi−yip_i - y_i (because σ′=σ(1−σ)\sigma' = \sigma(1 - \sigma); S-M07d q7 asks for it), so by the chain rule g=Xa⊤(p−y)+λwfeatg = X_a^\top (p - y) + \lambda w_{\text{feat}}. Differentiating once more, H=Xa⊤SXa+λIfeatH = X_a^\top S X_a + \lambda I_{\text{feat}}. For any vector vv, v⊤Hv=∑ipi(1−pi)(xi⋅v)2+λ∥vfeat∥2≥0v^\top H v = \sum_i p_i(1 - p_i)(x_i \cdot v)^2 + \lambda \lVert v_{\text{feat}} \rVert^2 \ge 0: the loss is convex, every stationary point is the global minimum (the proof is S-M07d q9). With an intercept, the intercept row of g=0g = 0 reads ∑i(pi−yi)=0\sum_i (p_i - y_i) = 0: the fitted probabilities average to the base rate.

M01.2 found a root of ff by x←x−f(x)/f′(x)x \leftarrow x - f(x)/f'(x). Minimizing LL is finding a root of its gradient, so the step is w←w−H−1gw \leftarrow w - H^{-1} g: solve Hδ=gH \delta = g (one lu and one lu_solve) and subtract. Expanding w−H−1gw - H^{-1}g shows each step is a weighted least-squares solve with weights pi(1−pi)p_i(1 - p_i) that change every iteration, hence iteratively reweighted least squares. Near the optimum Newton converges quadratically: the error is roughly squared each step, so 10−210^{-2} becomes 10−410^{-4}, then 10−810^{-8}, then rounding. Five or six steps from w=0w = 0 are typical; the reference stops once a step is below 10−1210^{-12} relative. Without a penalty and with perfectly separable data there is no finite minimum (the weights grow forever), which is one reason the head is always fitted with λ>0\lambda > 0.

The sigmoid needs care: e−ze^{-z} overflows for z<−709z < -709 and ez/(1+ez)e^{z}/(1 + e^{z}) is ∞/∞\infty/\infty for z>709z > 709. Computing e−∣z∣e^{-|z|}, which is at most 1, and choosing the formula by the sign of zz is exact everywhere.

The penalty λ2∥wfeat∥2\frac{\lambda}{2}\lVert w_{\text{feat}} \rVert^2 pulls feature weights toward 0; it adds λ\lambda to the Hessian diagonal, which also makes HH safely invertible. The intercept is not penalized: shrinking it toward 0 would bias every probability toward 0.5 instead of toward the base rate. With λ→∞\lambda \to \infty the features are switched off and σ(b)\sigma(b) becomes the fraction of positives. This is scikit-learn’s convention with C=1/λC = 1/\lambda.

A threshold tt turns scores into decisions: say 1 when score ≥t\ge t. Then TPR =#{positives≥t}/P= \#\{\text{positives} \ge t\}/P and FPR =#{negatives≥t}/N= \#\{\text{negatives} \ge t\}/N. Lowering tt from +∞+\infty to below the smallest score walks from (0,0)(0, 0) to (1,1)(1, 1): each positive passed is a step up, each negative a step right. That staircase is the ROC curve, one corner per distinct score. Several examples with the same score cannot be separated by any threshold, so they enter together: a diagonal step. Stepping through them one at a time would make the curve depend on the order of the input.

The area under the staircase, trapezoid(tpr, fpr) from M01.4, has a meaning: each up-step at horizontal position FPR contributes the fraction of negatives already below that positive. Summed, the AUC is the fraction of (positive, negative) pairs in which the positive scores higher, a tie counting 12\frac{1}{2} (the diagonal step’s triangle). That is the Mann-Whitney UU statistic divided by PNPN. Consequences: AUC ignores the scores’ scale (any increasing transform leaves it unchanged), flipping the labels gives 1−AUC1 - \text{AUC}, random scores give 0.5.

A classifier is calibrated if, among the predictions made with confidence cc, a fraction cc is correct. Group the predictions into BB equal-width bins of confidence, (b/B,(b+1)/B](b/B, (b + 1)/B], and compare each bin’s mean confidence with its accuracy: plotted, that is the reliability diagram. The expected calibration error is the count-weighted mean gap:

ECE=∑bnbn ∣accb−confb∣.\text{ECE} = \sum_{b} \frac{n_b}{n}\,\lvert \text{acc}_b - \text{conf}_b \rvert .

For a binary probability pp the prediction is 1 when p≥0.5p \ge 0.5 and the confidence is max⁡(p,1−p)\max(p, 1 - p); with class probabilities, the top class and its probability (the “top-label” ECE of Guo et al.). Even a perfectly calibrated classifier has an ECE of a few hundredths from sampling noise in each bin, which is why the test compares calibrated against overconfident rather than against 0.

A model that ranks well but is overconfident has one cheap fix (Guo et al. 2017): divide every logit row by one temperature T>0T > 0 fitted on held-out data. The argmax cannot change, so accuracy stays; only the confidences move. With β=1/T\beta = 1/T and p=softmax⁡(βzi)p = \operatorname{softmax}(\beta z_i), the mean NLL is

ℓ(β)=1n∑i[log⁡∑ceβzic−βziyi],ℓ′(β)=1n∑i(Ep[zi]−ziyi),ℓ′′(β)=1n∑iVar⁡p[zi]≥0.\ell(\beta) = \frac1n \sum_i \Bigl[\log \textstyle\sum_c e^{\beta z_{ic}} - \beta z_{i y_i}\Bigr], \qquad \ell'(\beta) = \frac1n \sum_i \bigl(\mathbb{E}_p[z_i] - z_{i y_i}\bigr), \qquad \ell''(\beta) = \frac1n \sum_i \operatorname{Var}_p[z_i] \ge 0 .

So ℓ\ell is convex in β\beta, and its minimum is the root of a scalar function whose derivative is known in closed form: exactly what M01.2’s newton(f, df, x0) solves. fit_temperature calls it with f=ℓ′f = \ell', df=ℓ′′df = \ell'', and x0=0x_0 = 0. The start matters: at β=0\beta = 0 the softmax is uniform and the curvature is largest, so the steps climb to the root; from β=1\beta = 1 a peaked softmax has almost no curvature, and for a model three times too confident the first step lands thousands of units away. A root at β≤0\beta \le 0 means the logits rank the labels no better than chance, and there is no temperature to report.

One Newton step. One feature, no intercept: x=(1,−1)x = (1, -1), y=(1,0)y = (1, 0), λ=1\lambda = 1. At w=0w = 0 both pi=12p_i = \frac{1}{2}.

QuantityComputationValue
gg(12−1)(1)+(12−0)(−1)+1⋅0(\tfrac12 - 1)(1) + (\tfrac12 - 0)(-1) + 1 \cdot 0−1-1
HH14⋅1+14⋅1+1\tfrac14 \cdot 1 + \tfrac14 \cdot 1 + 132\tfrac32
w1w_10−(−1)/320 - (-1)/\tfrac3223\tfrac23

The optimum satisfies g=0g = 0: 2(1−σ(w))=w2(1 - \sigma(w)) = w, so w⋆=0.6748…w^\star = 0.6748\ldots; one step already got within 0.008.

ROC and AUC. Scores (0.9,0.7,0.6,0.4,0.3)(0.9, 0.7, 0.6, 0.4, 0.3) with labels (1,0,1,1,0)(1, 0, 1, 1, 0): P=3P = 3, N=2N = 2.

threshold+∞+\infty0.90.70.60.40.3
(FPR, TPR)(0, 0)(0, 1/3)(1/2, 1/3)(1/2, 2/3)(1/2, 1)(1, 1)

The area is 12⋅13+12⋅1=23\frac12 \cdot \frac13 + \frac12 \cdot 1 = \frac23. Check by pairs: positive 0.9 beats both negatives, 0.6 beats only 0.3, 0.4 beats only 0.3: 4 of 6 pairs.

ECE. Probabilities (0.9,0.8,0.3,0.6)(0.9, 0.8, 0.3, 0.6), labels (1,0,0,1)(1, 0, 0, 1), B=5B = 5. Predictions (1,1,0,1)(1, 1, 0, 1), confidences (0.9,0.8,0.7,0.6)(0.9, 0.8, 0.7, 0.6), correct (yes,no,yes,yes)(\text{yes}, \text{no}, \text{yes}, \text{yes}). Bin indices ⌈5c⌉−1\lceil 5c \rceil - 1: 4,3,3,24, 3, 3, 2. Bin 4: conf 0.9, acc 1, gap 0.1; bin 3: conf 0.75, acc 0.5, gap 0.25; bin 2: conf 0.6, acc 1, gap 0.4. ECE =14(0.1)+24(0.25)+14(0.4)=0.25= \frac14 (0.1) + \frac24 (0.25) + \frac14 (0.4) = 0.25. These three examples are the first three tests.

Temperature. Three items with logits (0,1)(0, 1), labels (1,1,0)(1, 1, 0). Here Ep[z]=σ(β)\mathbb{E}_p[z] = \sigma(\beta), so ℓ′(β)=σ(β)−23\ell'(\beta) = \sigma(\beta) - \frac23: the calibrated confidence is the observed 23\frac23, β=ln⁡2\beta = \ln 2, and T=1/ln⁡2=1.4427T = 1/\ln 2 = 1.4427. With logits (0,2)(0, 2) the same confidence needs 2β=ln⁡22\beta = \ln 2, so TT doubles. This is test_hand_example_temperature.

def logistic_regression_fit(X: ArrayLike, y: ArrayLike, l2: float, iters: int, fit_intercept: bool = True) -> NDArray: ...
def logistic_predict_proba(X: ArrayLike, w: ArrayLike, fit_intercept: bool = True) -> NDArray: ...
def roc_curve(scores: ArrayLike, labels: ArrayLike) -> tuple[NDArray, NDArray, NDArray]: ... # fpr, tpr, thresholds
def roc_auc(scores: ArrayLike, labels: ArrayLike) -> float: ...
def reliability_bins(probs: ArrayLike, labels: ArrayLike, n_bins: int = 15) -> tuple[NDArray, NDArray, NDArray]: ...
def ece(probs: ArrayLike, labels: ArrayLike, n_bins: int = 15) -> float: ...
def fit_temperature(logits: ArrayLike, labels: ArrayLike, max_iter: int = 50) -> float: ... # T, by M01.2's newton

The intercept is the last weight, the layout L6.5 writes into the linear head’s b. roc_curve starts at (0,0)(0, 0) with threshold +∞+\infty and has one more point per distinct score; roc_auc is trapezoid(tpr, fpr). Bins are (b/B,(b+1)/B](b/B, (b+1)/B] with index clip⁡(⌈cB⌉−1,0,B−1)\operatorname{clip}(\lceil cB \rceil - 1, 0, B - 1).

TestKINDChecksWhy it matters downstream
test_hand_example_newton_stepunit, smokeone step gives 2/32/3; converged w=2(1−σ(w))w = 2(1 - \sigma(w))you and the tests agree on gg and HH
test_hand_example_roc_aucunit, smokethe six corners and 2/32/3L6.5’s reported AUC
test_hand_example_eceunit, smokebins (0,0,1,2,1)(0, 0, 1, 2, 1) and ECE 0.25the safety report’s calibration line
test_logistic_matches_scipygoldenIRLS equals scipy’s BFGS minimizer on four datasetsthe policy head is the right fit
test_fit_is_a_stationary_pointpropertygradient below 10−910^{-9}; mean pp = base ratesection 2.1
test_newton_converges_quadraticallypropertyerrors square per step; 8 steps equal 50section 2.2
test_l2_shrinks_feature_weightspropertyweight norm falls with λ\lambda; intercept freesection 2.3
test_sigmoid_never_overflowsboundaryz=±1000z = \pm 1000 gives exactly 1 and 0, never NaNembeddings with large norms
test_predict_proba_intercept_lastunitthe intercept is the last weightthe linear-head format
test_roc_matches_brute_forcegoldencorners and thresholds on continuous and tied scoressection 2.4
test_auc_is_mann_whitneygoldenAUC = U/(PN)U/(PN)section 2.5
test_auc_propertiespropertymonotone invariance, label flip 1−AUC1 - \text{AUC}, 1 and 0.5AUC is about ranking only
test_ties_take_a_diagonal_stepboundarytwo tied examples give 0.5 in either orderquantized scores tie often
test_reliability_and_ece_match_referencegoldenbins and ECE, binary and 3-classsection 2.6
test_ece_calibrated_vs_overconfidentstatisticalcalibrated under 0.04, sharpened above 0.1ECE measures the right gap
test_bin_edgesboundary0.5, 0.75, 1.0, and 0 land in the right binscross-language agreement
test_multiclass_top_labelunitfirst argmax, row max confidenceL6.5’s multi-class heads
test_rejects_bad_argumentsboundarybad labels, shapes, penalties, one-class ROC, bad probabilities raisebugs surface at the call
test_hand_example_temperatureunit, smokeT=1/ln⁡2T = 1/\ln 2 for gap 1, twice that for gap 2section 3
test_temperature_calibrates_an_overconfident_modelpropertylogits 3 times too sharp give TT near 3, a true NLL minimum, the same argmax, half the ECEsection 2.7
test_temperature_rejects_bad_inputboundary1-D logits, bad classes, NaN, worse than chance raiseno negative temperatures
PitfallSymptomCaught by
1. ez/(1+ez)e^{z}/(1 + e^{z}) or 1/(1+e−z)1/(1 + e^{-z}) without careNaN probabilities for large logitstest_sigmoid_never_overflows (mutant s01)
2. penalizing the interceptprobabilities pulled toward 0.5; base rate losttest_l2_shrinks_feature_weights (mutant s02)
3. the gradient of the likelihood instead of the lossNewton walks uphill and divergestest_hand_example_newton_step (mutant s03)
4. a Hessian without the weights p(1−p)p(1-p)linear, not quadratic, convergence; wrong after few stepstest_newton_converges_quadratically (mutant s04)
5. stepping through tied scores one at a timethe AUC depends on the input ordertest_ties_take_a_diagonal_step (mutant s05)
6. sorting scores ascendingAUC reported as 1−AUC1 - \text{AUC}test_hand_example_roc_auc (mutant s06)
7. a curve that does not start at (0,0)(0, 0)the first corner is missing from plots and thresholdstest_roc_matches_brute_force (mutant s07)
8. an unweighted mean over binsone stray prediction in an empty corner dominates the ECEtest_reliability_and_ece_match_reference (mutant s08)
9. binary confidence pp instead of max⁡(p,1−p)\max(p, 1 - p)confident negatives counted as unconfidenttest_hand_example_ece (mutant s09)
10. bins [lo,hi)[\text{lo}, \text{hi}) by flooredge values one bin up; another number than Go and the papertest_bin_edges (mutant s10)
11. the intercept column firstthe head’s b holds a feature weighttest_predict_proba_intercept_last (mutant s11)
12. returning β\beta instead of T=1/βT = 1/\betaan overconfident model gets sharper, not softertest_hand_example_temperature (mutant s12)
13. ℓ′\ell' without the label’s logitthe root no longer depends on the labelstest_hand_example_temperature (mutant s13)
DirectionModuleHow it uses this
BackM03.2lu and lu_solve solve Hδ=gH \delta = g each Newton step
BackM01.4trapezoid(tpr, fpr) is the AUC
BackM01.2newton(f, df, 0.0) finds the root of ℓ′(β)\ell'(\beta) for fit_temperature; IRLS is the same method on a gradient vector
ForwardL6.5fits the usage-policy head on embeddings and exports it as linear-head.schema.json with its AUC and ECE
ForwardL3.5probing classifiers over frozen representations
Forwardethics.04the safety report’s ROC and reliability diagram
ForwardS-M07dq5 to q9: regression, odds, the loss gradient, AUC by hand, and the convexity proof
Your pieceProduction equivalentWhat it addsWhere to look
logistic_regression_fitsklearn.linear_model.LogisticRegressionL-BFGS and Newton-CG solvers that never form HH, L1 penalties, multinomial fitssklearn/linear_model/_logistic.py
IRLSstatsmodels GLMthe same IRLS for every exponential family, with standard errors from H−1H^{-1}statsmodels/genmod/generalized_linear_model.py
roc_curve, roc_aucsklearn.metrics.roc_curve, roc_auc_scoredropping collinear corners, multiclass one-vs-rest averagingsklearn/metrics/_ranking.py
ecetemperature scalingone scalar fitted on held-out logits that fixes most overconfidenceGuo et al. (2017), section 4.2
the policy headLlama Guard, OpenAI moderationa classifier model instead of a linear head, with per-category thresholdsInan et al., “Llama Guard” (2023)