Skip to content

Eval runner

Moduleag.09 · build · Go · Pass 10 · 4 to 5 h
You buildgo/agent/eval/: eval.go (Run, observations, scores, aggregation), stats.go (mean, percentile bootstrap), report.go (load suites, write results.jsonl and summary.json), subject.go (provider, agent, and durable-run subjects)
Contractsuites in course/contracts/formats/eval-case.schema.json, reports in course/contracts/formats/eval-result.schema.json; the Go API in section 4
Testscourse/tests/go/ag_09/ (what they check: section 4)
Needsag.01 provider and deltas, ag.02 tools (the tests give an agent a tool), ag.03 the agent loop, ag.05 durable agent runs, load.01 the Go PCG32; reading: M07.4 (confidence intervals and the bootstrap, re-implemented here in Go), case study 05
Used byag.10 scorers · ag.11 LLM judge · ag.12 A/B experiments · craft.23 eval tests; your {ctl} eval verb drives it
MilestoneMS-agent
Optional depthEfron and Tibshirani, An Introduction to the Bootstrap, ch. 6 and 13; Miller, Adding Error Bars to Evals (free)
  • A scorer that fails is a missing measurement: it is excluded from the mean and counted, never averaged in as 0.
  • Outcome, evidence, and grade are separate: a subject that crashed has no output to grade, and the report says so instead of scoring an empty string.
  • Every number in a report carries a confidence interval; with samples, the bootstrap resamples cases, not rows.
  • The result is the same whatever order concurrent work finishes in, and the same seed gives the same interval in any language.
  • Eval reports are a file format (eval-result.schema.json) that the release gate and the A/B runner read.
Terminal window
ol start ag.09 # stubs go/agent/eval/*.go into your repo
ol tests ag.09 # read the test catalog first
ol check ag.09 # exit code is the verdict
ol check ag.09 --ref-deps # only if you skipped ag.01 to ag.05 or load.01
ol diff ag.09 # after passing: your code against the reference

Then add {ctl} eval run --suite <file> --base-url <url> --model <name> to your umbrella CLI (learner territory): load the suite, build a ProviderSubject (or an AgentSubject around your agent), run it with the scorers of ag.10, and write evals/<suite>/<run_id>/. MS-agent runs it over {fixture:ag.09/suite.jsonl} against your gateway.


Your agent answers questions over the course docs, calls query_usage, and runs durably. Is it any good? Change the system prompt, and is it better or worse? Today you would ask it three questions, read the answers, and decide by feel. That fails twice: three answers say almost nothing (section 3 shows how little), and “by feel” cannot be rerun after the next change. dur.11’s EvalSuite already runs the model zoo’s suites for the language models of Part 6; this module is the same discipline for agents: a runner that sends every case of a suite through a subject, scores every output with code, and reports means with intervals in a file the release gate can read. The scorers come in ag.10, the LLM judge in ag.11, and the A/B comparison in ag.12; all of them plug into what you build here.

SymbolMeaningType
nncases in the suiteinteger
SSsamples per case (WithSamples)integer, default 1
xi,sx_{i,s}a score of case ii, sample ss, when the scorer succeededfloat64
xˉ\bar{x}the mean of every xi,sx_{i,s} that existsfloat64
BBbootstrap resamples (n_boot)integer, default 2000
α\alphaone minus the confidence levelfloat64, default 0.05
xˉ(b)∗\bar{x}^{*}_{(b)}the bb-th smallest resample mean, b=0,…,B−1b = 0, \ldots, B-1float64

A suite is a JSON-lines file of cases (eval-case.schema.json): case_id, input (a prompt string or {"messages": [...]} in the OpenAI shape), optional ground_truth, tags, and scorer_args. Loading it is the first place an eval goes wrong silently, so LoadSuite rejects a missing required key, an unknown key (a typo such as groundtruth would otherwise drop every answer key), and a repeated case_id (two cases with one id overwrite each other in every report and A/B pairing), naming the line. Each case becomes an observation; after the subject runs, the observation holds the output, its timing, and annotations its scorers need.

Case study 05’s lesson, applied to every row:

FactWhere it livesA failure means
outcomeRow.SubjectErrorthe subject never produced an output: the provider was down, the durable run is waiting for approval
measurementScore.Errorthe scorer could not score this output: it returned an error, panicked, produced NaN or infinity, or set Error (an unparsable judge reply, ag.11)
gradeScore.Valuethe output was scored

When the subject fails, every score of that row is an error that names the subject’s failure, and no scorer is shown an empty output (which an exact-match scorer would happily grade 0). When a scorer fails, that value is excluded from its mean and counted in errored. Averaging a failure in as 0 blames the model for a broken scorer; a run where the judge’s endpoint was down would report the model as useless. The counts stay in the report beside the mean, because excluding failures also hides them: 3 of 500 errored is noise, 300 of 500 means the number is about the other 200.

Cases are independent, so Run works on up to WithConcurrency(n) of them at once. Completion order is then random, and anything that depends on it changes between runs: the row order of the report, and with it any statistic that walks rows in order. The fix is structural: each (case, sample) has a slot decided before anything runs, case index * S + sample, and its worker writes only there. Each worker also gets its own copy of the case’s annotation map; subjects write into it (tool_calls), and two workers writing one shared map is a data race that go test -race reports and production silently corrupts.

xˉ\bar{x} alone hides how much it could move. With n=3n = 3 the mean can only be 0,1/3,2/3,10, 1/3, 2/3, 1 for pass/fail scores, and section 3 shows that its 95% interval covers almost everything. The percentile bootstrap (M07.4) estimates the spread without a formula: pretend the cases you have are the population, draw nn of them with replacement, compute the mean of the draw, repeat BB times, and sort the BB means. The interval is

[xˉ(⌊Bα/2⌋)∗, xˉ(⌈B(1−α/2)⌉−1)∗],\left[\bar{x}^{*}_{(\lfloor B\alpha/2 \rfloor)},\ \bar{x}^{*}_{(\lceil B(1-\alpha/2) \rceil - 1)}\right],

indices 50 and 1949 for B=2000B = 2000, α=0.05\alpha = 0.05. Off by one at either end is the classic bug: ceil(...) without the - 1 reads one past the bound.

Resample cases, not rows. With S>1S > 1, the SS samples of a case are correlated: the same question is hard for every sample. Resampling the nSn S rows as if independent treats them as nSnS cases and gives an interval that is too narrow. This is the cluster bootstrap: draw case indices, and each drawn case brings all its samples.

Determinism. The draws come from PCG32 (load.01) on the sample sub-stream, rng.Stream(seed, rng.PurposeSample); each resample draws nn indices with Below(n), in order. With the same seed the interval is identical in Go and in the course’s Python oracle, which is how the tests can compare intervals exactly.

eval-result.schema.json fixes two files under evals/<suite>/<run_id>/:

  • results.jsonl: one row per (case, sample): suite, case_id, subject, sample, input_sha (SHA-256 of the input’s bytes, so a changed case is visible), output, scores (a number, or null for a failed score), errors (the messages of failed scores), latency_ms, ttft_ms (null when no token arrived; 0 would claim an instant answer), tokens, trace_id (null when not traced).
  • summary.json: suite, run_id, subjects, metrics (subject, then score name, then mean, ci_low, ci_high, n, errored), n_boot, seed, and in ag.12 the ab block.

Both files are written to a temporary name and renamed, so a crash never leaves half a report for the release gate to read.

A subject fills an observation from its input. Three come with the runner:

SubjectRunsRecords
ProviderSubject(p, now, opts...)one streamed chat completion through any ag.01 providertext; TTFT (first text delta), every text delta’s arrival (TokenTimes), total latency, completion tokens
AgentSubject(a, now)one run of your ag.03 agentthe same timing over the run’s streamed text, completion tokens summed over its model calls, and a tool_calls annotation: name, arguments, gate verdict, error flag, result
DurableSubject(cfg, prefix, now)one ag.05 durable run per case and sample, id RunID(prefix, case, sample)the answer; a run that stops awaiting approval or indeterminate fails the case

The durable subject is what makes a 500-case agent eval affordable: if the process dies at case 300, rerunning the suite replays the 300 finished runs from their journals without calling the model. Its run ids must be valid file names (separators replaced, long ids hashed), stable for a case and sample, and different per sample, or sample 1 silently replays sample 0. now is the clock (nil means time.Now); tests pass a fake one so timings are exact.

Aggregation. Four cases scored by an exact-match scorer that cannot score case c:

CaseScore
a1
b0
cerror
d1

xˉ=(1+0+1)/3=2/3\bar{x} = (1 + 0 + 1)/3 = 2/3 with n=3n = 3, errored =1= 1. Averaging the error in as 0 would give 2/4=0.52/4 = 0.5. This is TestHandExample.

How wide is three? Take the three values that exist, {1,0,1}\{1, 0, 1\}, and enumerate the bootstrap exactly instead of sampling: each of the 3 draws picks a 1 with probability 2/32/3, so the number of ones KK in a resample is binomial:

KKresample meanprobability
00(1/3)3=1/27=0.037(1/3)^3 = 1/27 = 0.037
11/31/33⋅(2/3)(1/3)2=6/27=0.2223 \cdot (2/3)(1/3)^2 = 6/27 = 0.222
22/32/33⋅(2/3)2(1/3)=12/27=0.4443 \cdot (2/3)^2 (1/3) = 12/27 = 0.444
31(2/3)3=8/27=0.296(2/3)^3 = 8/27 = 0.296

The 2.5th percentile falls in the first row (0.037 > 0.025), so the lower bound is 0; the 97.5th falls in the last row, so the upper bound is 1. The 95% interval is [0,1][0, 1]: three cases cannot tell a perfect agent from a useless one. With the 20 binary cases of the fixture (14 ones), the sampled interval is [0.5,0.9][0.5, 0.9]: still wide, now informative.

package eval // import "tinyllm/agent/eval"
type Timing struct { Start time.Time; TTFT, Total time.Duration; TokenTimes []time.Duration }
type Observation struct {
ID string; Sample int
Input, Output, GroundTruth json.RawMessage
Annotations map[string]json.RawMessage
Timing Timing; Tokens int; TraceID string
}
type Score struct { Name string; Value float64; Reason, Error string }
type Scorer interface { Name() string; Score(ctx context.Context, o Observation) (Score, error) }
type Subject func(ctx context.Context, o *Observation) error
type Row struct { Observation; Subject string; Scores []Score; SubjectError string }
func (r Row) Errored(i int) bool
type Stat struct { Mean, CILow, CIHigh float64; N, Errored int }
type SuiteResult struct {
Suite, Subject string; Rows []Row; Scorers []string; Metrics map[string]Stat
SubjectErrors, NBoot int; Alpha float64; Seed uint64
}
func Run(ctx context.Context, name string, obs []Observation, s []Scorer, opts ...Option) (*SuiteResult, error)
func WithSubject(name string, s Subject) Option
func WithConcurrency(n int) Option // default 4
func WithSamples(n int) Option // default 1
func WithBootstrap(nBoot int, alpha float64) Option // default 2000, 0.05
func WithSeed(seed uint64) Option
func WithClock(now func() time.Time) Option
var ErrSuite error
func Mean(xs []float64) float64
func BootstrapRNG(seed uint64) *rng.PCG32
func PercentileBounds(nBoot int, alpha float64) (lo, hi int)
func BootstrapCI(groups [][]float64, nBoot int, alpha float64, r *rng.PCG32) (lo, hi float64)
func LoadSuite(r io.Reader) ([]Observation, error)
func InputSHA(input json.RawMessage) string
type ResultRow struct { /* one results.jsonl line */ }
type ABStat struct { Delta, CILow, CIHigh, PValue float64; NPairs int }
type AB struct { Base, Exp string; Metrics map[string]ABStat }
type Summary struct { Suite, RunID string; Subjects []string; Metrics map[string]map[string]Stat; AB *AB; NBoot int; Seed uint64 }
func (r *SuiteResult) ResultRows() []ResultRow
func (r *SuiteResult) Summary(runID string) Summary
func WriteResults(dir string, rows []ResultRow, sum Summary) error
var ErrInput error
func Messages(input json.RawMessage) ([]types.Message, error)
type ToolCallRecord struct { Name string; Args json.RawMessage; Verdict string; IsError bool; Result string }
func ProviderSubject(p types.Provider, now func() time.Time, opts ...types.CallOption) Subject
func AgentSubject(a *loop.Agent, now func() time.Time) Subject
func RunID(prefix, caseID string, sample int) string
func DurableSubject(cfg durableagent.Config, prefix string, now func() time.Time) Subject

Without WithSubject, Run scores the observations as given (their outputs already filled), which is how ag.10’s tests score recorded answers. Cancelling the context ends the run with the context’s error and no result. Use only the standard library, your ag.01, ag.03, ag.05 packages, and tinyllm/ds/rng.

TestKINDChecksWhy it matters downstream
TestHandExampleunitsection 3: mean 2/3, n 3, errored 1; the CI contains the meanyou and the tests agree on exclusion
TestScorerFailuresExcludedfaultan error, a panic, NaN, infinity, an Error field: each excluded and counted; other scorers unaffectedthe judge (ag.11) returns unparsable replies as errors
TestDeterministicOrderUnderConcurrencyproperty40 cases, 2 samples, 8 workers finishing out of order: rows in case then sample order, numbers equal to 1 worker; no shared annotation map (-race)reports diff cleanly between runs
TestBootstrapMatchesOracleconformancethree fixtures (binary, grouped, 90%) equal the Python oracle’s boundsthe interval is reproducible across languages
TestPercentileBoundsboundaryindices for BB = 2000, 1000, 500, 10, 1no off-by-one at either end
TestRunResamplesCasesconformance10 cases x 3 samples through Run equal the oracle’s cluster bootstrapsamples never inflate confidence
TestSubjectErrorFailsRowfaulta failed subject’s row: every score errored with its message, counted in SubjectErrorsoutcome is not confused with grade
TestRejectsBadSuitesboundaryrepeated or empty case ids, repeated scorer names: ErrSuiteno case or metric overwrites another
TestLoadSuiteunitthe fixture suite loads; missing, unknown, repeated, and malformed lines are errors naming the linehand-written suites fail loudly
TestReportSchemaconformanceresults.jsonl and summary.json keys, input_sha, null scores with errors, null TTFT and trace, the summary’s statistics; no temp file leftdur.12 and ag.12 read your reports
TestProviderSubjectTimingunitwith a 10 ms clock: TTFT 10 ms, token times 10 and 20 ms, latency 30 ms, 2 tokens, "Hello"; an error delta fails the caseag.10’s TTFT, TTLT, and ITL scorers
TestAgentSubjectRecordsToolCallsunittwo calls in order with verdicts, results, and error flags; tokens summed over model callsag.10’s tool-success scorer
TestDurableSubjectReplaysfaulta second run of the suite makes no model calls; samples are separate runsa crashed eval resumes for free
TestRunIDboundaryseparators, .., 300-byte ids: valid, short, stable, distinct per samplerun ids are file names in ag.05’s store
TestMessagesunitstring and chat inputs; numbers, empty chats, tool roles: ErrInputno case runs with an empty prompt
TestCancelStopsRunfaultcancelling early or with every case in flight returns context.Canceled and no resulta cut-short eval never looks complete
PitfallSymptomCaught by
1. averaging a failed score in as 0a dead judge endpoint reports the model as uselessTestHandExample, TestScorerFailuresExcluded (mutant s01)
2. trusting scorers: no recover, NaN accepted, the Error field ignoredone bad case ends the eval, or NaN poisons the meanTestScorerFailuresExcluded (mutants s02, s03, s04)
3. appending rows as workers finishtwo runs of one suite produce different reportsTestDeterministicOrderUnderConcurrency (mutant s05)
4. sharing the case’s annotation map between workersa data race; annotations of one row appear in anotherTestDeterministicOrderUnderConcurrency (mutant s06)
5. resampling rows when there are samplesintervals too narrow by about S\sqrt{S} for correlated samplesTestRunResamplesCases (mutant s07)
6. a percentile index off by one, or another random streamintervals that match no other implementationTestBootstrapMatchesOracle, TestPercentileBounds (mutants s08, s09)
7. scoring a failed subject’s empty outputcrashes counted as wrong answersTestSubjectErrorFailsRow (mutant s10)
8. a lax suite loadera typo in a key silently drops every answer key; two cases share an idTestLoadSuite (mutants s11, s12)
9. 0 for unknown values in the report, or hashing the decoded inputa failed score reads as a wrong answer; input_sha hides changed casesTestReportSchema (mutants s13, s14, s15)
10. TTFT taken at the last tokenlatency scorers report the whole answer time as time to first tokenTestProviderSubjectTiming (mutant s16)
11. not matching tool results to their callsthe tool-success scorer sees no failuresTestAgentSubjectRecordsToolCalls (mutant s17)
12. one durable run per case instead of per sample, or ids with separatorssample 1 replays sample 0; a case id with / cannot be storedTestDurableSubjectReplays, TestRunID (mutants s18, s21)
13. accepting two scorers with one nameone metric silently replaces the otherTestRejectsBadSuites (mutant s19)
14. returning partial results after a cancelan interrupted eval looks like a smaller, complete oneTestCancelStopsRun (mutant s20)
15. accepting any input shapea case with a typo runs with an empty promptTestMessages (mutant s22)
DirectionModuleHow it uses this
Backag.01Provider, the deltas, and Accumulator inside ProviderSubject
Backag.02the tool registry of the agent under test
Backag.03loop.Agent and its events inside AgentSubject
Backag.05durableagent.Run inside DurableSubject
Backload.01rng.Stream and Below behind every bootstrap
Forwardag.10scorers implement Scorer and read Timing, GroundTruth, and the tool_calls annotation
Forwardag.11the judge is a Scorer; its unparsable replies are Score.Error, excluded and counted
Forwardag.12RunExperiment runs two subjects through this runner and fills Summary.AB
Forwardcraft.23eval results become regression-test evidence

If you skip this module, ol check ag.10 reports needs ag.09: build it, or pass --ref-deps.

Your pieceProduction equivalentWhat it addsWhere to look
Runsaige eval/harness, OpenAI Evals, Inspect AItask registries, sandboxes, model-graded and human-graded mixes, eval logs viewerssaige (free), Inspect (free)
BootstrapCIscipy.stats.bootstrapBCa intervals that correct bias and skewSciPy docs (free)
DurableSubjectTemporal-backed eval pipelinesper-case retries, rate limits, and resumable runs across machinesdur.*, Temporal docs (free)