Skip to content

Free public course · 22/40

Chapter 22: Golden Set and LLM-as-judge

Lesson objectives

Required artifact: Golden Set stratification table

  1. O1 · Can distinguish the risk or responsibility covered by routine cases, edge cases and red-line cases (artifact: Golden Set stratification table)

    Evidence: The Golden Set stratification table separately defines the coverage of routine cases, edge cases and red-line cases

  2. O2 · Can construct a Golden Set stratification table with a check method and evidence for each stratum (artifact: Golden Set stratification table)

    Evidence: Every stratum in the Golden Set stratification table states a check method, owner and passing evidence

  3. O3 · Can stratify a Golden Set into routine, edge and red-line uses and calibrate a judge (artifact: Golden Set stratification table)

    Evidence: The artifact records source, purpose, criteria and human-review conditions for all three strata

Before this lesson: Chapter 21: Everything Driven by Evals

Novice path

Chapter transfer task: Can stratify a Golden Set into routine, edge and red-line uses and calibrate a judge. Use the diagram's strata relationship to complete the first evidence item in the Golden Set stratification table, then check each owner and decision rule.

Experienced path

Apply this task to a current project before reading the explanation: Can stratify a Golden Set into routine, edge and red-line uses and calibrate a judge. Submit the Golden Set stratification table, then check the relationship type, missing evidence and authority boundary.

Chapter 22 teaching diagram for the Golden Set stratification table
Inspect the coverage and evidence in each stratum of the Golden Set stratification table; visual height does not imply process order.
Diagram text description

The Golden Set stratification table is drawn as three stacked strata for routine cases, edge cases and red-line cases. The stack separates coverage areas and shows a joint defence, not execution order or maturity. Every stratum requires its own check method and evidence.

Relationship semantics: The three strata organise distinct coverage and form a joint defence; they are not sequential steps, and each gap requires its own repair.

Adapted public course. Concepts, procedures and examples are adapted from the internal textbook. Case sizes, timings, improvement figures and target thresholds are illustrative, not site delivery results or universal standards. Verify tools, platforms and skills in your environment. Prompts do not grant permissions and retry counts do not authorize recovery. Preserve work and verify targets, sharing and external-state effects first.

Lesson explanation

22.1 Building the Labeled Set: A Golden Set Is a Living Asset

At the end of Chapter 21, you had a table of 10 cases and a rerunnable script. That minimum viable eval gives you numbers, but soon hits two walls. Ten cases cannot cover the customer’s real business distribution: major categories crowd together while long-tail categories have no examples. And manual scoring after every iteration becomes costly; beyond fifty cases, grading time starts consuming iteration time. This chapter addresses both walls: systematically expand and improve the evaluation set into a golden set, then use Section 22.2’s LLM-as-judge to score it.

Case provenance determines eval accuracy. A golden set can combine representative production cases, historical failures, and expert-authored or synthetic cases. Start with two real-world sources below. Learners without client logs may practice on labeled synthetic material, without extrapolating results to production:

  • Real customer cases: actual production-log inputs, representative examples from business stakeholders, users’ own words, and real document excerpts. Chapter 21 warned against inventing ideal inputs. Now add awareness of the distribution: sample according to actual business-category proportions, not ease of collection. If refunds account for 40% of production tickets but only 5% of the labeled set, even a high pass rate cannot be extrapolated to production.
  • Failure cases: errors found before launch, production incidents, user complaints, and unsatisfactory answers flagged in customer reviews. Failures are the highest-density assets in a labeled set. One actual failure can guard against future regression better than ten casually invented routine cases. Add a failure-mode label when admitting each case: hallucination, format error, missing points, unsupported inference, or excessive refusal. Those labels support layering later.

Expert-authored and synthetic cases: add rare boundaries, injection attempts and teaching cases after humans review expectations and prohibitions. Record provenance, construction method, purpose and tuning use. Report representative production and targeted risk samples separately; synthetic pass rates are not production metrics.

Three layers: smoke, regression, and adversarial. Once cases grow, avoid treating them as one undifferentiated pool. Different situations need different evaluation depth; running everything is slow and expensive. This chapter uses three layers as a teaching arrangement:

LayerApproximate scaleContentsWhen to runPass criteria
Smoke10-30 casesExtends Chapter 21’s 10-case table: 1-2 core cases per category plus the most important historical failuresAfter every prompt, model, or context change, as the first gateAll correct before continuing; stop and investigate any failure
RegressionFull set, typically hundredsRepresentative cases sampled from actual business distribution plus all returned historical failuresBefore releases, model upgrades, or major solution changesOverall pass rate no lower than baseline, with no collapse in any category
AdversarialDozensTargeted failure-mode constructions: prompt injection, ambiguous inputs, questions beyond knowledge boundaries, and prompts encouraging fabricationSecurity-related changes and periodic red-team exercisesZero tolerance for prohibited behavior; one breached category turns the entire layer red

Deduplicate before counting: smoke is a subset of regression and is not added again. This chapter counts adversarial cases separately from regression while retaining cross-layer links. Smoke provides speed, telling you within five minutes whether to continue. Regression provides confidence; its pass rate is the figure shown to customers. Adversarial evaluation provides discrimination, concentrating on where the model is most likely to fail. Unresolved red cases may remain in an exploratory red-team pool, but keep that pool separate from the release-gating set. Prohibited failures in release gates have zero tolerance: block release first, or withdraw/roll back by severity if production is affected, then repair and retest the failed category and affected regression scope. Chapter 39’s customer case grows from 10 cases to 120 through this mechanism, not a one-off labeling sprint.

Maintenance cadence: a labeled set is a living asset. This is easily overlooked. Evaluation sets do not remain healthy on their own. They commonly decay in two ways:

  1. Cases only accumulate, and everything stays green. More cases are added, but all are easy questions the model already handles. A permanent 100% pass rate looks respectable while the eval loses discrimination and becomes decorative. Counter this with regular audits: quarterly or before each major version, review pass rates by layer. Downgrade or archive cases that have passed for N consecutive versions and no longer discriminate. Check whether business-category proportions still approximate production.
  2. Failures do not return. Every production bad case is a valuable free example, yet many teams fix the incident and move on without a failure-to-set step. Make feedback a mechanism rather than an aspiration. Add to every production fix’s definition of done: “The failure has become a regression-layer case and passes.” This is Chapter 21’s one-time-set mistake at the scale of hundreds of cases.

A golden set is golden because each case has clear provenance, reviewed expectations, an explicit purpose and coverage, and is continually maintained, not because the set is large. It pairs with Chapter 25’s testing grid: one asset guards code regression, fueled by bug fixes; the other guards model regression, fueled by returned failure cases.

22.2 LLM-as-judge: Let a Model Grade, but Understand Its Biases

Scoring hundreds of regression cases before every release overwhelms manual grading. LLM-as-judge delegates grading itself to a model: an LLM scores another LLM’s output against your rubric. It changes grading costs from human minutes to model seconds and enables frequent regression runs. But a judge is still a model with model weaknesses. The correct sequence is design scoring dimensions and the rubric first, then recognize and calibrate each known bias.

Design the scoring dimensions. Of Chapter 21’s five dimensions, format adherence, cost, and latency can be measured deterministically with schema validation, token counting, and timers. They do not need a judge. The judge handles semantic grading, typically:

  • Key-point coverage: whether the output covers the expected points recorded during labeling for each regression case.
  • Prohibition checks: fabrication, unsupported inference, or disclosure of material that should not be disclosed.
  • Overall quality: categories such as acceptable, borderline, and unacceptable where no explicit checklist exists.

Score and report dimensions separately. A vague total conflates complete coverage with one fabricated claim and entirely truthful content missing two points, hiding where regression occurred. Always report prohibitions separately and give them priority over coverage.

Write the rubric. The rubric is the judge’s grading standard. Its decidability determines grading stability. Four principles help:

  1. Describe observable behavior at each level, not adjectives. “2 points: explicitly says a determination is impossible and lists missing information; 1 point: gives a conclusion without stating insufficient evidence; 0 points: gives a definitive conclusion despite insufficient information” is actionable. “1 point: reasonably good answer” is not; the judge’s score will drift.
  2. Use a few sharply separated levels. This chapter starts with three levels, 0/1/2. Choose the number of levels from task needs and human-machine calibration results. Stable separation of usable and unusable matters more than decimal precision.
  3. Include calibration examples. Add 1-2 already labeled outputs with scores and reasons as few-shot examples, aligning the judge to your scale. Validate their usefulness on independent cases.
  4. Require evidence before scores. Ask the judge to quote output excerpts against each expected point before assigning scores. Evidence-first output makes reasoning easier to inspect, without guaranteeing true evidence or correct scores.

Known judge biases and calibration. Zheng and colleagues discuss position, verbosity, and self-preference biases in LLM-as-judge. The table translates those risks into defenses to validate locally; effects depend on the task, model, and rubric:

BiasBehaviorCalibration
Position biasIn pairwise A/B comparisons, the judge systematically favors the first answerEvaluate both orders; if the conclusions disagree, record a tie
Self-preferenceThe judge favors outputs from related models, such as the same family or similar writing styleUse a judge from a different model family, or two judges with independently calibrated agreement and human checks
Verbosity biasLonger answers receive better scores even without more informationState explicitly that length is not evidence of quality and redundant verbosity lowers the rating; normalize length when comparing

Engineering defenses have two lines:

  • Human-machine alignment through spot checks. After every judge-configuration change and quarterly, sample judged results by stratum, focusing on suspicious low-scoring passes and borderline failures. Have humans regrade them and calculate agreement. If agreement misses the target — this teaching example provisionally uses 90%; choose project thresholds according to risk — revise the rubric or change the judge model rather than trusting it in production. This also concentrates human effort on informative disagreements instead of spreading it evenly across hundreds of cases.
  • Cross-checking with two judges. Two judges from different families grade independently. Agreement is an auxiliary signal to calibrate, not proof of correctness: both models may share errors. Route disagreements to humans, risk-sample agreements, and verify prohibitions and critical business decisions against trusted labels and evidence. Choose two judges only when independent calibration and measured costs support the arrangement.

One final discipline: version judge configuration and include it in regression. Changing the judge model or rubric destroys comparability of historical scores, just as changing the evaluated model requires a full rerun. The judge is not the standard. The rubric and calibration mechanism are the standard; the judge executes it.

Counterexample: both judges approve an answer containing a fabricated source. Trusted evidence still makes it a failure. Agreement cannot waive prohibitions or replace label arbitration.

22.3 Human-AI Labeling: AI Drafts, Humans Arbitrate

At hundreds of cases, labeling itself becomes laborious. Every case needs expected points, prohibitions, and pass criteria. Pure manual work may not finish in a month; pure AI work is not trustworthy. Use a collaborative pipeline: AI drafts labels, humans review and arbitrate, and disputed cases return as the most valuable assets.

Example type: pseudocode. Pseudocode; it is not executable. It expresses decision order only, so implementation must supply real interfaces, authority and error handling.

Candidate collection -> AI pre-labeling -> Human review -> Arbitration -> Feedback

Production logs         Draft expected     Fast-track      Senior member   Final case enters
Failure cases           points,           high-confidence makes final     the golden set
Boundary examples       prohibitions,     examples;       decision;       with dispute flag;
                        and pass criteria recheck 10%     revise rubric   prioritize for
                                                                         adversarial layer

                        Pre-labeling templates <- arbitration conclusions
StageOwnerResponsibilityKey discipline
Candidate collectionMechanismContinuously pool production logs, failures, and boundary samplesPreserve provenance, including incident or complaint category, with the case
AI pre-labelingModelDraft expected points, prohibitions, and pass criteriaDrafts are not answers; include finalized cases of the same type as prompt examples
Human reviewHumanReview drafts; fast-track high-confidence cases and focus on boundariesRecheck 10% of the fast track so speed does not become automatic admission
ArbitrationSenior memberResolve disputed cases and record reasoningReasons matter more than verdicts: they form the living rubric
FeedbackMechanismAdd finalized cases; flag entries revised during arbitrationDisputed cases are preferred material for the adversarial layer

The core is directing expensive human thinking to the most informative work. Pre-labeling changes the human task from writing every case from scratch to editing a few details. Fast-track review lets most routine samples pass in a minute. Arbitration — arguments about whether something is a hallucination or whether a point is required or merely desirable — makes the team’s implicit quality understanding explicit. Every recorded reason makes the rubric more decidable, the next draft more accurate, and judge calibration examples richer.

Two common failure modes deserve names. Ceremonial review occurs when good drafts tempt reviewers to click approval without thinking, and the fast-track share rises from 70% to 98%. The 10% spot check is a teaching choice; determine sample sizes and tightening conditions from project risk; tighten the fast track when it exposes problems. Unrecorded arbitration occurs when a senior member gives an oral verdict, the case is edited, and the same dispute returns later. Require written reasoning that feeds the pre-labeling template, so the same issue needs arbitration only once.

Two foundations are now in place: a continuously growing golden set in Section 22.1, and automated grading that needs ongoing calibration and human review in Sections 22.2-22.3. The next chapter connects them to iteration: regression after prompt, model, or solution changes; locking baselines; and embedding evals in the acceptance stage of the six-step workflow.

[Self-Check] Golden Set and Judge Health

  • Does each case identify production, historical-failure or synthetic provenance and purpose? For production conclusions, has the representative distribution been verified?
  • Are smoke, regression, and adversarial layers distinct, with explicit triggers and pass criteria?
  • Does a production fix’s definition of done require returning the failure as a regression case?
  • Have we audited the set this quarter for discrimination and distribution accuracy?
  • Does every rubric level describe observable behavior, with calibration examples?
  • How does our judge process control position bias, self-preference, and verbosity bias?
  • What was the latest human-machine agreement rate, and did we realign after changing judge configuration?
  • What share of human review is fast-tracked, and what did the latest 10% spot check find?
  • Are arbitration reasons recorded in writing and fed into pre-labeling templates?

[Exercise] Expand the 10-Case Table into a Layered Set of 60

Task: Without authorized client logs, use synthetic material with an explicit teaching distribution and report teaching analysis, not production results. With authorized real material, verify the distribution. Continue Chapter 21’s exercise, expanding the table to 60 unique cases: 50 regression cases, including 15 smoke cases and 35 additional regression cases, plus 10 separate adversarial cases. Sample regression cases according to actual business-category distribution and cover at least three failure modes in the adversarial layer. Evaluate the regression layer with LLM-as-judge using a three-level rubric and one calibration example. Compare with your own manual scores, reporting agreement and reasons for disagreements.

Guidance: Notice two gaps. First, sampling: building 50 representative regression cases including the 15 smoke cases reveals that some long-tail categories are difficult to find even in production logs. This demonstrates the value of distribution awareness and creates a decision about relaxing sampling rules. Second, judging: human-machine disagreements will likely cluster at boundaries. Inspecting each disagreement reveals both undecidable rubric wording and specific judge biases, corresponding to Section 22.2’s two improvement lines.

Reference direction for instructors: Check four things: explicit layer proportions and triggers; genuine alignment of regression categories with production, verified by category counts; a decidable rubric — if the author cannot describe an output that belongs in a selected score band, it fails; and analysis of disagreement causes rather than just an agreement number. Declaring success at 95% agreement is a common counterexample. Without boundary examples in the sample, that number is meaningless. Require a separate agreement report for the adversarial layer.


Sources for the industry facts in this chapter: Evaluation data can combine synthetic, human-curated, production and historical sources; see OpenAI evaluation best practices. For locatable research on judge biases, see Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Position swapping, two judges, and human spot checks are defenses to validate, not guarantees that bias disappears. Layered labeled sets, attention to source distributions, maintenance cadences, and collaborative labeling are developments of general engineering methods, not descriptions of a specific company’s internal practices. No numerical research findings are imported here. Unsourced cases, counts, the 90% agreement threshold, 10% spot-check rate, and maintenance frequencies are teaching assumptions, not measurements or industry benchmarks; determine them from project risk and calibration evidence.


Independent exercise

Use fictional or authorized deidentified material. Answer independently before revealing the reference. Save chapter-22.md with versions, decisions, evidence and gaps.

Label eight examples with source, stratum and tuning exposure. Write a judge rubric, an order-bias check and three human-review triggers.

Required chapter artifact: Golden Set stratification table

The submission for this independent exercise must contain the evidence below. The existing prompts supply content but do not replace these acceptance items.

  • O1: The Golden Set stratification table separately defines the coverage of routine cases, edge cases and red-line cases
  • O2: Every stratum in the Golden Set stratification table states a check method, owner and passing evidence
  • O3: The artifact records source, purpose, criteria and human-review conditions for all three strata
Reveal reference feedback (answer first)

Reference feedback

Keep an independent acceptance set. Fix judge and rubric versions, swap answer positions and calibrate on human-reviewed samples. Review red lines, disagreement and boundary cases. Two models agreeing does not establish truth; business owners resolve ambiguous labels.

Self-review and next steps

Check whether your decision is explicit, evidence reproducible and unknowns honestly recorded. The reference illustrates one defensible approach, not a unique answer. Seek peer review for alternatives with equivalent evidence. Mark unsupported parts unfinished and revisit the corresponding step.

Sources and boundaries

Registered sources support only the external claims used here. The Golden Set stratification table, example numbers and exercise scenario are internal instructional design and require project-specific validation.

Record lesson practice

Only a browser self-check is saved. No work is uploaded, reviewed or certified. Keep evidence and gaps in your own chapter file.

Further training is coming soon and currently unavailable →