Skip to content

LLM judge with position control

Moduleag.11 · build · Go · Pass 10 · 3 h
You buildrubric-based score parsing, pairwise comparisons with position swaps, and sampled scoring
Testscourse/tests/go/ag_11/: TestJudgeParsesEvidence, TestJudgeRejectsMalformedOutput, TestPairwisePositionSwap, TestSampledTolerance, TestKappaAgainstLabels
Needsag.01 provider and ag.09 runner; reading: M07.4
Used byag.12 experiment runner, craft.23
MilestoneMS-agent

Code scorers in ag.10 are reliable when the target is explicit, such as an exact answer or a state delta. They cannot judge whether an explanation is useful or whether a response follows a nuanced rubric. A language model can help, but its output must be treated as a noisy measurement.

Ask for a constrained JSON response with a numeric score and short evidence. Reject malformed output, out-of-range values, and missing evidence. Never turn a judge failure into score zero. For pairwise comparisons, ask both A/B and B/A, map the second answer back to the original labels, and resolve disagreement using a fixed rule.

When repeated human labels are available, Cohen’s kappa is κ = (p_o - p_e)/(1 - p_e), where p_o is observed agreement and p_e is agreement expected from the label marginals. Sampling reduces judge cost; report both the number sampled and tolerance used.

For a pair of responses, the first prompt says A then B and returns A. The swapped prompt says B then A and returns A, which maps to original B. The judge disagrees across positions, so the comparison is marked position-sensitive and cannot silently count as a clean win. TestPairwisePositionSwap catches a judge that prefers the first option regardless of quality.

Build NewJudgeScorer, pairwise scoring, kappa reporting, and Sampled. The generator interface should accept a prompt and return raw JSON text, which lets tests provide deterministic responses without network access. TestJudgeParsesEvidence and TestJudgeRejectsMalformedOutput check response validation; TestPairwisePositionSwap catches position bias; TestSampledTolerance checks sample accounting; TestKappaAgainstLabels checks agreement against frozen labels.

TestWhat it checks
TestPairwisePositionSwaphand-worked position swap and winner mapping
TestJudgeParsesEvidencevalid rubric response parsing
TestJudgeRejectsMalformedOutputmalformed and out-of-range responses
TestSampledTolerancedeterministic sample selection
TestKappaAgainstLabelsagreement against frozen labels
Terminal window
ol start ag.11
ol tests ag.11
ol check ag.11
  • Parse and validate the full response; do not extract the first digit from prose.
  • Swap candidate positions and invert the second result before combining.
  • Keep judge errors distinct from low scores.
  • A sampled scorer’s uncertainty is not the same as its inner scorer’s value spread.
DirectionModuleHow it uses this
Backag.01provides the provider used by the judge generator
Backag.09stores judge results and scorer failures
Forwardag.12compares versions with code and judge scorers
Forwardcraft.23uses judge scores as evaluation evidence

MS-agent reports the judge’s sample count, tolerance, and agreement with the frozen labels fixture.

Your pieceProduction equivalentWhat it adds
pairwise position swapLMSYS Chatbot Arenamany human comparisons aggregated into model rankings
sampled judgeeval budget policycost controls with a measured quality tolerance
rubrichuman annotation guidecalibrated labels for judge validation