Skip to content

Paired A/B experiments

Moduleag.12 · build · Go · Pass 10 · 3 to 4 h
You buildpaired subject runs, paired bootstrap confidence intervals, and an experiment report
Testscourse/tests/go/ag_12/: TestPairedDeltaHandExample, TestFailedPairsExcluded, TestBootstrapPairs, TestKnownEffect, TestNullCoverage, TestEvalSuiteSpecCarriesAB
Needsag.09 runner, ag.10 scorers, ag.11 judge, load.01, dur.11; reading: M07.4, M07.5
Used bycraft.23; optional MS-C2
MilestoneMS-agent

Two agent versions should answer the same cases so the comparison is not dominated by an easier or harder sample. The experiment runner pairs by case id, computes each case’s score difference, and estimates uncertainty over those differences.

For case i, define d_i = score_exp_i - score_base_i. The reported effect is mean(d). A paired bootstrap samples case indices with replacement and keeps each base/experiment pair together. Its percentile interval estimates uncertainty in the effect. The same seed and inputs must reproduce the same result.

Cases missing a score on either side are excluded from the paired metric and counted. They must never be converted to zero. Reject duplicate ids and mismatched scorer names before computing a comparison.

Let base scores be [0, 1, 0] and experiment scores be [1, 1, 0]. The paired deltas are [1, 0, 0], so the effect is 1/3. An unpaired comparison would throw away the fact that the second and third cases were shared, widening or shifting the result unnecessarily.

Implement RunExperiment in go/agent/eval/experiment. Return each metric’s mean delta, confidence bounds, p-value, and pair count. Upgrade the durable EvalSuite workflow so the same run can produce ordinary eval results and an optional A/B block. TestPairedDeltaHandExample checks the hand calculation; TestFailedPairsExcluded ensures failed pairs are counted but omitted; TestBootstrapPairs checks deterministic resampling; TestKnownEffect and TestNullCoverage check interval behavior; TestEvalSuiteSpecCarriesAB checks workflow serialization.

TestWhat it checks
TestPairedDeltaHandExampleworked paired-delta calculation
TestFailedPairsExcludedexcludes incomplete pairs
TestBootstrapPairsresamples paired rows deterministically
TestKnownEffectdetects a synthetic positive effect
TestNullCoverageinterval coverage under a null effect
TestEvalSuiteSpecCarriesABforwards experiment configuration
Terminal window
ol start ag.12
ol tests ag.12
ol check ag.12
  • Resample pairs, never the two arms independently.
  • Preserve deterministic case ordering before sampling.
  • Use the same scorer definition and seed on both arms.
  • A confidence interval containing zero is uncertainty, not proof of equivalence.
DirectionModuleHow it uses this
Backag.09runs each subject and summarizes scores
Backag.10supplies deterministic scorers
Backag.11supplies optional judge scores
Backload.01provides deterministic observation loading
Backdur.11provides the durable EvalSuite workflow being upgraded
Forwardcraft.23uses paired evaluations as regression tests
ForwardC2optional post-training comparison

MS-agent reports a known synthetic effect, while optional MS-C2 compares the post-trained model against its base.

Your pieceProduction equivalentWhat it adds
paired bootstraprandomized controlled experimentuncertainty while controlling case difficulty
EvalSuite workflowrelease evaluation pipelinedurable, resumable experiment runs
delta and intervalsequential monitoringevidence gathered over repeated releases