Skip to content

Run comparison and regression gate

Moduleload.02 · side · Go · Pass 7 · 3 to 5 h (optional)
You buildgo/loadgen/compare/compare.go: metric and budget parsing, histogram reports expanded into samples, nearest-rank statistics by selection, a seeded one-sided two-sample permutation test, and the gate that fails a head run only when it is worse beyond the budget AND significantly so
Contractthe input: formats/loadgen-report.schema.json · the role {loadgen} compare: spec/cli-roles.md
Testscourse/tests/go/load_02/, 10 tests (what they check: section 4)
Needsload.01 your report, histograms, and rng.Stream (chapter) · reading: M07.5 permutation tests (chapter), re-implemented here in Go
Used byno module: your {loadgen} compare entry point runs it, in dep.05’s optional perf-gate job and in drill ops.07 (git bisect run)
Milestonenone (optional; not part of MS-L10)
Optional depthGood, Permutation, Parametric, and Bootstrap Tests of Hypotheses, ch. 3; Kalibera and Jones, Rigorous Benchmarking in Reasonable Time (free); Mytkowicz et al., Producing Wrong Data Without Doing Anything Obviously Wrong
  • A regression is a head run worse than the base by more than the budget AND unlikely under noise; either condition alone cries wolf (TestGateNeedsBothBudgetAndSignificance).
  • The permutation test asks how often a random relabeling of the pooled samples is at least as bad as what was observed; it needs no distributional assumption about latencies (TestCompareHandExample).
  • One-sided: a faster head is never a regression, so its p-value is near 1 (TestPermutationTestIsSeededAndOneSided).
  • p=(1+count)/(1+N)p = (1 + \text{count})/(1 + N) is never 0, and a seeded shuffle gives the same verdict twice on the same files (TestPValueIsNeverZero).
  • Over many simulated pairs, identical runs pass at least 95% of the time and a +10% shift is caught at least 90% of the time (TestIdenticalRunsPassAndShiftsAreCaught).
Terminal window
ol start load.02 # stubs go/loadgen/compare/compare.go
ol tests load.02
ol check load.02 # exit code is the verdict
ol check load.02 --ref-deps # only if load.01 is not passing yet
ol diff load.02

Your {loadgen} compare <base.json> <head.json> --metric ttft_p95 --max-regress 5% [--alpha 0.05] [--permutations 1000] [--seed 0] reads the two reports, calls compare.Compare, prints the verdict as its final JSON line, and exits 1 when Regressed.


load.01 turns a run into numbers; nothing yet turns two runs into a decision. CI (dep.05) must fail a pull request that makes TTFT worse, and drill ops.07 hands you twenty commits and asks which one did it, which git bisect run can answer only if one command exits 0 for “fine” and 1 for “regressed”. Latency is noisy: two runs of the same build differ by several percent, so “head p95 > base p95” fails half of all pull requests, and “head p95 > 1.05 base p95” still fails a few percent of identical builds and misses real regressions hidden in noise. This module writes the gate: an effect-size budget plus a significance test. It is optional: only your {loadgen} compare entry point calls it, so no milestone requires it, and dep.05 checks its perf-gate job only when your workflow has one.

SymbolMeaningType
a1..anaa_1..a_{n_a}, b1..bnbb_1..b_{n_b}base and head samples of one metric (ms, or 0/1 errors)[]float64
S(⋅)S(\cdot)the statistic: nearest-rank percentile or meanfunction
Δ=(S(b)−S(a))/S(a)\Delta = (S(b) - S(a))/S(a)relative change, worse when positivefloat64
β\betathe budget (--max-regress 5% is 0.05)float64
NNrandom relabelings; α\alpha the significance level (0.05)

A report keeps each latency as histogram buckets [upper_ms, count] (load.01): a bucket becomes count copies of its upper bound (within 1/128 of the true values). error_rate becomes one 1 per failed request and one 0 per other request, so the same test covers it.

If base and head come from the same distribution, the labels are arbitrary: every way of choosing which nan_a of the na+nbn_a + n_b pooled samples are “base” is equally likely. Observe d=S(b)−S(a)d = S(b) - S(a). Draw NN random relabelings (Fisher-Yates with rng.Stream(seed, shuffle), load.01), compute d∗d^* for each, and count those with d∗≥dd^* \ge d. The one-sided p-value is p=(1+#{d∗≥d})/(1+N)p = (1 + \#\{d^* \ge d\})/(1 + N): the observed labeling is one of the possible ones, so p≥1/(1+N)p \ge 1/(1 + N). Only “worse” counts: a head that got faster has d<0d < 0 and pp near 1.

Regressed   ⟺  Δ>β\iff \Delta > \beta and p<αp < \alpha. The budget keeps tiny real changes (a 2% slowdown on a huge sample is significant but allowed); the test keeps big apparent changes on tiny samples (one request each, ten times slower, p≈0.5p \approx 0.5) from failing CI. With a base statistic of 0 (no errors), any rise is Δ=+∞\Delta = +\infty.

The nearest-rank percentile is the ⌈qn−10−9⌉\lceil qn - 10^{-9} \rceil-th smallest value, the same rule as load.01’s histogram. Each relabeling needs it on a fresh split, so the gate uses quickselect (expected O(n)O(n)) on a copy, never sorting or reordering the caller’s slice.

Base e2e samples {10,20,30}\{10, 20, 30\} ms, head {40,50,60}\{40, 50, 60\} ms, metric e2e_mean, budget 5% (test TestCompareHandExample).

  1. S(a)=20S(a) = 20, S(b)=50S(b) = 50, Δ=(50−20)/20=1.5>0.05\Delta = (50 - 20)/20 = 1.5 > 0.05.
  2. d=30d = 30. There are (63)=20\binom{6}{3} = 20 ways to pick three of the six values as “head”; only {40,50,60}\{40, 50, 60\} gives a difference of 30 or more (every other choice moves a smaller value into the head mean). So the exact p=1/20=0.05p = 1/20 = 0.05, and 20000 random relabelings estimate it within 0.005.
  3. p=0.05p = 0.05 is not below α=0.05\alpha = 0.05: three samples cannot establish a regression, however large the change. Not regressed.

With 400 samples per run instead, a 10% shift gives pp far below 0.05 and Δ≈0.10>0.05\Delta \approx 0.10 > 0.05: regressed.

go/loadgen/compare/compare.go
type Stat struct { Name string; Q float64; Mean bool }
func (s Stat) Value(xs []float64) float64 // nearest rank or mean; xs unchanged; NaN when empty
type Options struct { Metric string; MaxRegress, Alpha float64; Permutations int; Seed uint64 } // Alpha 0.05, Permutations 1000 by default
type Verdict struct { Metric string; Base, Head, Delta, PValue float64; Regressed bool }
func ParseRegress(s string) (float64, error) // "5%" or "0.05"
func ParseMetric(m string) (string, Stat, error) // "ttft_p95" -> ("ttft_ms", p95); "error_rate"
func Samples(r *loadgen.Report, metric string) ([]float64, error)
func PermutationTest(base, head []float64, s Stat, n int, r *rng.PCG32) float64
func Gate(base, head []float64, s Stat, o Options) (Verdict, error)
func Compare(base, head *loadgen.Report, o Options) (Verdict, error)

Samples come from math/rand/v2 with fixed seeds (lognormal latencies); the statistical test bounds its counts binomially at about p=10−3p = 10^{-3}.

TestKINDChecksWhy it matters downstream
TestCompareHandExampleunitsection 3: statistics, Δ=1.5\Delta = 1.5, p≈0.05p \approx 0.05, not regressedthe worked example
TestParseMetricAndBudgetunit5% is 0.05; metric names map to report keys; bad input refusedthe CLI’s flags mean what they say
TestValueIsNearestRankAndLeavesInputAloneproperty300 random slices with ties: equal to the sorted rank; input unchangedthe statistic and the split stay correct
TestSamplesExpandTheHistogramunitbuckets to samples; error_rate to 0/1; a missing metric is an errorreports become test data
TestPermutationTestIsSeededAndOneSidedunitsame seed, same p; faster head p>0.9p > 0.9; 30% slower p<0.01p < 0.01CI verdicts are reproducible
TestPValueIsNeverZeroboundarycompletely separated samples: p=1/1000p = 1/1000an honest bound, not certainty
TestGateNeedsBothBudgetAndSignificanceunit1 sample each: noise; 2% on 4000 samples: within budget; 10%: regressedneither condition alone decides
TestZeroBaselineboundary0 to 0 errors is no change; 0 to 30 of 170 is +∞+\infty and regressederror-rate gates
TestCompareReportsEndToEndunittwo load.01 reports, default alpha and permutations: 30% slower fails, self passesthe CI path
TestIdenticalRunsPassAndShiftsAreCaughtstatistical100 pairs each: at most 12 false alarms, at least 82 catchesthe gate’s two promises
PitfallSymptomCaught by
A two-sided testa faster build fails CITestPermutationTestIsSeededAndOneSided (mutant s01)
p=count/Np = \text{count}/Np = 0 claims certainty from a finite sampleTestPValueIsNeverZero (mutant s02)
The budget alone decidesidentical builds fail CI several percent of the timeTestGateNeedsBothBudgetAndSignificance (mutant s03)
Significance alone decidesa real but allowed 2% change fails CITestGateNeedsBothBudgetAndSignificance (mutant s04)
5% read as 5a head six times slower passesTestParseMetricAndBudget (mutant s05)
One sample per bucketcounts ignored: the test sees a handful of valuesTestSamplesExpandTheHistogram (mutant s06)
Never shufflingevery relabeling equals the observed one: p = 1, nothing is ever caughtTestPermutationTestIsSeededAndOneSided (mutant s07)
Rank by floorthe statistic differs from the report’s percentileTestValueIsNearestRankAndLeavesInputAlone (mutant s08)
DirectionModuleHow it uses this
Backload.01Report, its histograms, and rng.Stream for the shuffles

dep.05’s optional perf-gate job runs {loadgen} compare on main against the last good run; drill ops.07 bisects a perf regression with it.

Your pieceProduction equivalentWhat it addsWhere to look
a permutation test per metricBencher, Conbenchcontinuous benchmarking with history and change-point detectionBencher, Conbench
one comparisonbenchstatMann-Whitney U across repeated runs, with confidence intervalsbenchstat
a fixed budgetchange-point detection (E-divisive)finds when a series shifted, not just whether two runs differMongoDB’s change point detection