Logistic regression (IRLS), ROC-AUC, calibration (ECE)
Overview
Section titled “Overview”| Module | M07.7 · build · Python · Pass 5 · 3 to 4 h |
| You build | python/tinyllm/prob/metrics.py: logistic_regression_fit, logistic_predict_proba, roc_curve, roc_auc, reliability_bins, ece, fit_temperature |
| Contract | course/contracts/py/tinyllm/prob/metrics.pyi · the head it fits: formats/linear-head.schema.json |
| Tests | course/tests/M07.7/test_metrics.py (what they check: section 4), golden values from scipy 1.17.1 in course/fixtures/M07.7/scipy_golden.json |
| Needs | M03.2 LU (lu, lu_solve: one solve per Newton step) · M01.4 trapezoid (the area under the ROC curve) · M01.2 Newton’s method (newton: the temperature’s root) · reading: M07.5 hypothesis tests |
| Used by | later: L6.5 fits the usage-policy linear head and reports its AUC and ECE (D33) · L3.5 scores probes · ethics.04 puts both numbers in the safety report |
| Milestone | MS-P5 (the Pass 5 gate) |
| Optional depth | Hastie, Tibshirani, Friedman, The Elements of Statistical Learning (free), section 4.4 (logistic regression and IRLS); Fawcett, “An introduction to ROC analysis” (2006); Guo, Pleiss, Sun, Weinberger, “On Calibration of Modern Neural Networks” (2017) |
Key Takeaways
Section titled “Key Takeaways”- Logistic regression models ; its loss is convex, so the minimum is where the gradient vanishes (
test_fit_is_a_stationary_point). - Newton’s method on that loss is IRLS: each step is one linear solve with the Hessian , , and the digits double per step (
test_hand_example_newton_step,test_newton_converges_quadratically). - The ROC curve sweeps every threshold; its area, the AUC, is the probability a random positive outscores a random negative, ties counting half (
test_hand_example_roc_auc,test_auc_is_mann_whitney). - Calibration asks whether “0.8” means right 80% of the time; the ECE is the count-weighted gap between confidence and accuracy over bins (
test_hand_example_ece,test_ece_calibrated_vs_overconfident). - A ranking metric and a calibration metric answer different questions: sharpening probabilities leaves the AUC unchanged and wrecks the ECE.
How to work this chapter
Section titled “How to work this chapter”ol start M07.7 # stubs metrics.py into your repool tests M07.7 # read the test catalog first: rung R0, you write no tests hereol check M07.7 # exit code is the verdictol check M07.7 --ref-deps # only if you skipped M03.2 or M01.4ol diff M07.7 # after passing: your code against the reference1. Why now
Section titled “1. Why now”Pass 5 gives your system its first classifiers. L6.5 puts a classification head on BERT and ELECTRA, and the gateway’s usage policy (D33) is a linear head over your engine’s embeddings: a dot product and a sigmoid that Go evaluates on every request. That head has to be fitted, and gradient descent with a learning rate to tune is the wrong tool for a convex problem with a few hundred weights: Newton’s method solves it in five or six steps, each step a linear system you already know how to solve (M03.2). Once fitted, the head needs two numbers before anyone trusts it. Does it rank unsafe requests above safe ones (AUC)? And when it says 0.9, is it right 90% of the time, so a threshold of 0.9 means what the operator thinks (ECE)? This module builds the fitter and both metrics.
2. Principles
Section titled “2. Principles”| Symbol | Meaning | Type / shape |
|---|---|---|
| feature matrix, one row per example | float64[n, d] | |
| with a column of ones appended (the intercept column, last) | float64[n, d+1] | |
| labels | float64[n] | |
| weights; with an intercept, (the last) is the bias | float64[d+1] | |
| the logit (log-odds) of example | float64[n] | |
| the sigmoid | ||
| predicted | float64[n] | |
(l2) | ridge penalty on the feature weights | float |
| , | gradient and Hessian of the loss | [d+1], [d+1, d+1] |
[n, n] (never formed) | ||
| TPR, FPR | true and false positive rates at a threshold | |
(n_bins) | number of confidence bins | int |
2.1 The model and its loss
Section titled “2.1 The model and its loss”A linear score can be any real number; a probability must lie in . The sigmoid maps one to the other, and its inverse is the log-odds: . So logistic regression says the log-odds are linear in the features; a unit step in feature multiplies the odds by . Fitting maximizes the likelihood , or equivalently minimizes
The derivative of the bracket with respect to is just (because ; S-M07d q7 asks for it), so by the chain rule . Differentiating once more, . For any vector , : the loss is convex, every stationary point is the global minimum (the proof is S-M07d q9). With an intercept, the intercept row of reads : the fitted probabilities average to the base rate.
2.2 Fitting by Newton’s method
Section titled “2.2 Fitting by Newton’s method”M01.2 found a root of by . Minimizing is finding a root of its gradient, so the step is : solve (one lu and one lu_solve) and subtract. Expanding shows each step is a weighted least-squares solve with weights that change every iteration, hence iteratively reweighted least squares. Near the optimum Newton converges quadratically: the error is roughly squared each step, so becomes , then , then rounding. Five or six steps from are typical; the reference stops once a step is below relative. Without a penalty and with perfectly separable data there is no finite minimum (the weights grow forever), which is one reason the head is always fitted with .
The sigmoid needs care: overflows for and is for . Computing , which is at most 1, and choosing the formula by the sign of is exact everywhere.
2.3 Regularization
Section titled “2.3 Regularization”The penalty pulls feature weights toward 0; it adds to the Hessian diagonal, which also makes safely invertible. The intercept is not penalized: shrinking it toward 0 would bias every probability toward 0.5 instead of toward the base rate. With the features are switched off and becomes the fraction of positives. This is scikit-learn’s convention with .
2.4 The ROC curve
Section titled “2.4 The ROC curve”A threshold turns scores into decisions: say 1 when score . Then TPR and FPR . Lowering from to below the smallest score walks from to : each positive passed is a step up, each negative a step right. That staircase is the ROC curve, one corner per distinct score. Several examples with the same score cannot be separated by any threshold, so they enter together: a diagonal step. Stepping through them one at a time would make the curve depend on the order of the input.
2.5 AUC is a ranking probability
Section titled “2.5 AUC is a ranking probability”The area under the staircase, trapezoid(tpr, fpr) from M01.4, has a meaning: each up-step at horizontal position FPR contributes the fraction of negatives already below that positive. Summed, the AUC is the fraction of (positive, negative) pairs in which the positive scores higher, a tie counting (the diagonal step’s triangle). That is the Mann-Whitney statistic divided by . Consequences: AUC ignores the scores’ scale (any increasing transform leaves it unchanged), flipping the labels gives , random scores give 0.5.
2.6 Calibration and ECE
Section titled “2.6 Calibration and ECE”A classifier is calibrated if, among the predictions made with confidence , a fraction is correct. Group the predictions into equal-width bins of confidence, , and compare each bin’s mean confidence with its accuracy: plotted, that is the reliability diagram. The expected calibration error is the count-weighted mean gap:
For a binary probability the prediction is 1 when and the confidence is ; with class probabilities, the top class and its probability (the “top-label” ECE of Guo et al.). Even a perfectly calibrated classifier has an ECE of a few hundredths from sampling noise in each bin, which is why the test compares calibrated against overconfident rather than against 0.
2.7 Temperature scaling
Section titled “2.7 Temperature scaling”A model that ranks well but is overconfident has one cheap fix (Guo et al. 2017): divide every logit row by one temperature fitted on held-out data. The argmax cannot change, so accuracy stays; only the confidences move. With and , the mean NLL is
So is convex in , and its minimum is the root of a scalar function whose derivative is known in closed form: exactly what M01.2’s newton(f, df, x0) solves. fit_temperature calls it with , , and . The start matters: at the softmax is uniform and the curvature is largest, so the steps climb to the root; from a peaked softmax has almost no curvature, and for a model three times too confident the first step lands thousands of units away. A root at means the logits rank the labels no better than chance, and there is no temperature to report.
3. Worked example by hand
Section titled “3. Worked example by hand”One Newton step. One feature, no intercept: , , . At both .
| Quantity | Computation | Value |
|---|---|---|
The optimum satisfies : , so ; one step already got within 0.008.
ROC and AUC. Scores with labels : , .
| threshold | 0.9 | 0.7 | 0.6 | 0.4 | 0.3 | |
|---|---|---|---|---|---|---|
| (FPR, TPR) | (0, 0) | (0, 1/3) | (1/2, 1/3) | (1/2, 2/3) | (1/2, 1) | (1, 1) |
The area is . Check by pairs: positive 0.9 beats both negatives, 0.6 beats only 0.3, 0.4 beats only 0.3: 4 of 6 pairs.
ECE. Probabilities , labels , . Predictions , confidences , correct . Bin indices : . Bin 4: conf 0.9, acc 1, gap 0.1; bin 3: conf 0.75, acc 0.5, gap 0.25; bin 2: conf 0.6, acc 1, gap 0.4. ECE . These three examples are the first three tests.
Temperature. Three items with logits , labels . Here , so : the calibrated confidence is the observed , , and . With logits the same confidence needs , so doubles. This is test_hand_example_temperature.
4. The interface
Section titled “4. The interface”def logistic_regression_fit(X: ArrayLike, y: ArrayLike, l2: float, iters: int, fit_intercept: bool = True) -> NDArray: ...def logistic_predict_proba(X: ArrayLike, w: ArrayLike, fit_intercept: bool = True) -> NDArray: ...def roc_curve(scores: ArrayLike, labels: ArrayLike) -> tuple[NDArray, NDArray, NDArray]: ... # fpr, tpr, thresholdsdef roc_auc(scores: ArrayLike, labels: ArrayLike) -> float: ...def reliability_bins(probs: ArrayLike, labels: ArrayLike, n_bins: int = 15) -> tuple[NDArray, NDArray, NDArray]: ...def ece(probs: ArrayLike, labels: ArrayLike, n_bins: int = 15) -> float: ...def fit_temperature(logits: ArrayLike, labels: ArrayLike, max_iter: int = 50) -> float: ... # T, by M01.2's newtonThe intercept is the last weight, the layout L6.5 writes into the linear head’s b. roc_curve starts at with threshold and has one more point per distinct score; roc_auc is trapezoid(tpr, fpr). Bins are with index .
What the tests check
Section titled “What the tests check”| Test | KIND | Checks | Why it matters downstream |
|---|---|---|---|
test_hand_example_newton_step | unit, smoke | one step gives ; converged | you and the tests agree on and |
test_hand_example_roc_auc | unit, smoke | the six corners and | L6.5’s reported AUC |
test_hand_example_ece | unit, smoke | bins and ECE 0.25 | the safety report’s calibration line |
test_logistic_matches_scipy | golden | IRLS equals scipy’s BFGS minimizer on four datasets | the policy head is the right fit |
test_fit_is_a_stationary_point | property | gradient below ; mean = base rate | section 2.1 |
test_newton_converges_quadratically | property | errors square per step; 8 steps equal 50 | section 2.2 |
test_l2_shrinks_feature_weights | property | weight norm falls with ; intercept free | section 2.3 |
test_sigmoid_never_overflows | boundary | gives exactly 1 and 0, never NaN | embeddings with large norms |
test_predict_proba_intercept_last | unit | the intercept is the last weight | the linear-head format |
test_roc_matches_brute_force | golden | corners and thresholds on continuous and tied scores | section 2.4 |
test_auc_is_mann_whitney | golden | AUC = | section 2.5 |
test_auc_properties | property | monotone invariance, label flip , 1 and 0.5 | AUC is about ranking only |
test_ties_take_a_diagonal_step | boundary | two tied examples give 0.5 in either order | quantized scores tie often |
test_reliability_and_ece_match_reference | golden | bins and ECE, binary and 3-class | section 2.6 |
test_ece_calibrated_vs_overconfident | statistical | calibrated under 0.04, sharpened above 0.1 | ECE measures the right gap |
test_bin_edges | boundary | 0.5, 0.75, 1.0, and 0 land in the right bins | cross-language agreement |
test_multiclass_top_label | unit | first argmax, row max confidence | L6.5’s multi-class heads |
test_rejects_bad_arguments | boundary | bad labels, shapes, penalties, one-class ROC, bad probabilities raise | bugs surface at the call |
test_hand_example_temperature | unit, smoke | for gap 1, twice that for gap 2 | section 3 |
test_temperature_calibrates_an_overconfident_model | property | logits 3 times too sharp give near 3, a true NLL minimum, the same argmax, half the ECE | section 2.7 |
test_temperature_rejects_bad_input | boundary | 1-D logits, bad classes, NaN, worse than chance raise | no negative temperatures |
5. Pitfalls
Section titled “5. Pitfalls”| Pitfall | Symptom | Caught by |
|---|---|---|
| 1. or without care | NaN probabilities for large logits | test_sigmoid_never_overflows (mutant s01) |
| 2. penalizing the intercept | probabilities pulled toward 0.5; base rate lost | test_l2_shrinks_feature_weights (mutant s02) |
| 3. the gradient of the likelihood instead of the loss | Newton walks uphill and diverges | test_hand_example_newton_step (mutant s03) |
| 4. a Hessian without the weights | linear, not quadratic, convergence; wrong after few steps | test_newton_converges_quadratically (mutant s04) |
| 5. stepping through tied scores one at a time | the AUC depends on the input order | test_ties_take_a_diagonal_step (mutant s05) |
| 6. sorting scores ascending | AUC reported as | test_hand_example_roc_auc (mutant s06) |
| 7. a curve that does not start at | the first corner is missing from plots and thresholds | test_roc_matches_brute_force (mutant s07) |
| 8. an unweighted mean over bins | one stray prediction in an empty corner dominates the ECE | test_reliability_and_ece_match_reference (mutant s08) |
| 9. binary confidence instead of | confident negatives counted as unconfident | test_hand_example_ece (mutant s09) |
| 10. bins by floor | edge values one bin up; another number than Go and the paper | test_bin_edges (mutant s10) |
| 11. the intercept column first | the head’s b holds a feature weight | test_predict_proba_intercept_last (mutant s11) |
| 12. returning instead of | an overconfident model gets sharper, not softer | test_hand_example_temperature (mutant s12) |
| 13. without the label’s logit | the root no longer depends on the labels | test_hand_example_temperature (mutant s13) |
6. Where it’s used next
Section titled “6. Where it’s used next”| Direction | Module | How it uses this |
|---|---|---|
| Back | M03.2 | lu and lu_solve solve each Newton step |
| Back | M01.4 | trapezoid(tpr, fpr) is the AUC |
| Back | M01.2 | newton(f, df, 0.0) finds the root of for fit_temperature; IRLS is the same method on a gradient vector |
| Forward | L6.5 | fits the usage-policy head on embeddings and exports it as linear-head.schema.json with its AUC and ECE |
| Forward | L3.5 | probing classifiers over frozen representations |
| Forward | ethics.04 | the safety report’s ROC and reliability diagram |
| Forward | S-M07d | q5 to q9: regression, odds, the loss gradient, AUC by hand, and the convexity proof |
Going further
Section titled “Going further”| Your piece | Production equivalent | What it adds | Where to look |
|---|---|---|---|
logistic_regression_fit | sklearn.linear_model.LogisticRegression | L-BFGS and Newton-CG solvers that never form , L1 penalties, multinomial fits | sklearn/linear_model/_logistic.py |
| IRLS | statsmodels GLM | the same IRLS for every exponential family, with standard errors from | statsmodels/genmod/generalized_linear_model.py |
roc_curve, roc_auc | sklearn.metrics.roc_curve, roc_auc_score | dropping collinear corners, multiclass one-vs-rest averaging | sklearn/metrics/_ranking.py |
ece | temperature scaling | one scalar fitted on held-out logits that fixes most overconfidence | Guo et al. (2017), section 4.2 |
| the policy head | Llama Guard, OpenAI moderation | a classifier model instead of a linear head, with per-category thresholds | Inan et al., “Llama Guard” (2023) |