Skip to content

Free public course · 21/40

Chapter 21: Everything Driven by Evals

Lesson objectives

Required artifact: evaluation-loop record

  1. O1 · Can trace the forward and feedback paths among success criteria, evaluation cases and version decision (artifact: evaluation-loop record)

    Evidence: The evaluation-loop record draws both a forward path and a return path across success criteria, evaluation cases and version decision

  2. O2 · Can construct a evaluation-loop record with the input, decision and write-back evidence for every pass (artifact: evaluation-loop record)

    Evidence: Every stage in the evaluation-loop record states its input, decision, output and owner

  3. O3 · Can turn success criteria into repeatable cases and make a version decision from results (artifact: evaluation-loop record)

    Evidence: The artifact contains success criteria, representative cases, per-case results and a version decision

Before this lesson: Chapter 20: Telemetry-Driven Development and Reverse Feeding: Giving AI 'Far-Seeing Eyes'

Novice path

Chapter transfer task: Can turn success criteria into repeatable cases and make a version decision from results. Use the diagram's loop relationship to complete the first evidence item in the evaluation-loop record, then check each owner and decision rule.

Experienced path

Apply this task to a current project before reading the explanation: Can turn success criteria into repeatable cases and make a version decision from results. Submit the evaluation-loop record, then check the relationship type, missing evidence and authority boundary.

Chapter 21 teaching diagram for the evaluation-loop record
Inspect both directions of the evaluation-loop record and confirm that observed results change the next pass.
Diagram text description

The diagram moves forward from success criteria through evaluation cases to version decision, then returns by a dashed path. Forward arrows produce a result; the return path carries observation and correction. If results do not alter the next input, this is a one-way flow rather than a loop.

Relationship semantics: Three stages produce a result along the forward path, then observations return to change the next input; without write-back there is no loop.

Adapted public course. Concepts, procedures and examples are adapted from the internal textbook. Case sizes, timings, improvement figures and target thresholds are illustrative, not site delivery results or universal standards. Verify tools, platforms and skills in your environment. Prompts do not grant permissions and retry counts do not authorize recovery. Preserve work and verify targets, sharing and external-state effects first.

Lesson explanation

21.1 A Typical Scene: You Changed One Line of Prompt. Now What?

The ticket-classification assistant you built for a customer has been stable for three weeks in production. On Monday, the business team changes the definition of “urgent”: widen it from “affects production” to “affects production or delivery commitments.” You edit one sentence in the prompt. Five minutes, no code changes, all tests green.

Now answer this: How do you know the change has not broken the two hundred previously correct classifications among three hundred historical tickets?

Traditional software offers a familiar answer: treat the change as a release, run the test suite, and check every assertion. But here you changed a prompt, shifting the model’s behavior as a whole. No unit test can simply assert that all three hundred historical classifications remain reasonable. Without an evaluation set, you can only inspect a few dozen samples and hope.

Hope is a luxury at a customer site. When the technical lead asks whether the system still works after the change, “it feels fine” is insufficient. You need numbers. Evals supply them: fixed inputs, expected key points, and pass criteria that produce comparable results whenever rerun.

This capability is an explicit requirement in at least some FDE roles, not an invention of this textbook. Anthropic’s Forward Deployed Engineer job description states:

“Production experience with LLMs including advanced prompt engineering, agent development, evaluation frameworks, and deployment at scale.”

Notice how prompt engineering, agent development, evaluation frameworks, and deployment at scale appear together. The supported conclusion is narrower: this job description treats evaluation as part of production deployment experience rather than an optional bonus. It does not establish one universal industry standard.

OpenAI’s practice supports the same judgment. The Pragmatic Engineer’s August 2025 interview with its FDE leader describes three typical customer-project phases: several days on site for early scoping; building evals and improving against them during validation; then weekly on-site work during delivery. The middle phrase matters most. Iteration improves against evals, not impressions. Change a prompt: run them. Switch a model: run them. Adjust a solution: run them. The evaluation set provides the slope for improvement; without one, iteration becomes free fall.

Hence the chapter title. In AI systems, evals should drive almost every critical decision: prompt iteration follows eval results, model upgrades follow eval comparisons, and solution choices follow eval data. Iterating without evals has an unflattering but accurate name: vibes-driven development. Customers are not paying for your feelings.

21.2 An Eval Is Not a Software Test: Deterministic Assertions versus Probabilistic Output

Many engineering teams first respond, “We have tests; why not add a few?” That is insufficient. Evals and software tests both supply an input and check an output, but their underlying assumptions differ.

Typical deterministic unit tests constrain behavior precisely. With fixed inputs and controlled dependencies, assertions check agreed outputs or properties; Chapter 7’s acceptance and Chapter 25’s testing grid use these tools extensively. Software testing also covers concurrency, randomized algorithms, and unstable external dependencies. Not all software has a unique output, and one passing run does not prove a system correct.

Model evaluation inhabits a probabilistic world. Two runs on the same input can produce differently worded yet equally correct answers. Correctness is a distribution rather than one solution. Judging an output resembles grading an essay more than matching an answer key: which points did it cover, which errors did it make, and how serious were they?

DimensionTypical deterministic unit testModel evaluation: eval
Input-output relationshipUnder controlled conditions: agreed outputs or properties for the inputProbabilistic: one input can have several reasonable outputs
Judgment methodAssertions: equality or exceptionsScoring: expected-point coverage and prohibited-content checks
Meaning of passingBinary: pass or failContinuous: pass rate or score, such as 92% of cases meeting criteria
Meaning of failureInvestigate a mismatch in implementation, assertions, or test environmentCould be sampling fluctuation or actual regression; distinguish them
Result stabilityOne pass covers only the conditions tested in that runRecord actual settings and observe variation across reruns; determinism is not guaranteed
Changes triggering a rerunCode changesPrompt, model, temperature, context: a change to any requires reevaluation

Probabilistic outputs can also be scored deterministically, using exact classification labels, JSON schema, or prohibited-field checks. The distinction concerns the evaluated behavior and scope of conclusions, not whether assertions are used; the tools can be combined.

The last row deserves emphasis. In AI systems, change extends well beyond code. Moving to a newer model or replacing the documents in the context window can shift output quality across the board without modifying a line of code. Eval triggers must cover every variable affecting output, not merely code.

Avoid the opposite extreme of treating evals as a replacement for tests. They form layers. The testing grid guards code behavior — functions, interfaces, and business logic — against regression. Evals guard model-output quality — correctness, hallucinations, and format. Each defends a different half; omitting either leaves the system exposed. Chapter 23 closes this part by extending acceptance in the six-step workflow into code acceptance and model-output acceptance. Here we establish the eval track.

Five Quality Dimensions: What Should an Eval Measure?

Correctness is only one quality dimension. A complete eval should measure at least five separately, because they can conflict:

DimensionQuestionTypical measureCommon pitfall
CorrectnessIs the result right?Key-point pass rate, scoreAverages conceal long-tail failures
Hallucination rateIs anything fabricated?Share of cases containing invented contentHallucinations are common in customers’ private domains, where general models have the least training data
Format adherenceCan downstream systems consume the output?JSON schema validation pass rateA single format error can break the downstream pipeline
CostWhat does the call cost?Mean tokens or monetary cost per caseLong prompts multiplied by frequent calls multiply the bill
LatencyHow long does the user wait?P50 and P95 latencyAverage latency hides the experience determined by P95

Together they form a quality vector. Real engineering choices usually involve tradeoffs within it: a stronger model improves correctness while increasing cost and latency; stricter formatting improves adherence but can reduce expressive quality. Evals do not eliminate tradeoffs. They make them explicit in numbers, replacing competing intuitions with a comparison table customers and teams can discuss.

Use the table to check your present situation. How many dimensions can your project answer immediately? Many teams answer correctness with “roughly,” hallucination rate with “not measured,” format with “we find out when it crashes,” cost with “the month-end bill,” and latency with “no user complaints.” The more dimensions you can answer, the closer the system is to engineering. If you can answer none, it remains at demonstration maturity regardless of whether it is live.

21.3 A Minimum Viable Eval: Start with 10 Manual Cases

After deciding what to measure, the next trap is how to begin. Teams start with evaluation platforms, automated labeling, and judge pipelines; three months later the platform is unfinished and no eval exists. Start from the other end: one table, one script, and 10 manual cases, ready today.

Step One: Collect 10 Real Inputs

Case provenance determines eval quality. Collect cases in a 6 + 2 + 2 mix:

  • 6 routine cases from actual customer inputs: original tickets, users’ own words, or real document excerpts. Do not invent ideal inputs. Your invented inputs follow your thinking; customers’ inputs follow theirs, and that is where model errors occur.
  • 2 boundary cases with incomplete inputs, ambiguity, or several mixed questions. Boundary behavior provides discrimination between solutions.
  • 2 historical failure cases from actual failures discovered before or after launch. Places that failed before are especially likely to fail again in regression.

Step Two: Write Expected Points and Pass Criteria for Every Case

Expected points are not a single model answer; that would test a probabilistic system as though it were deterministic. They specify what a satisfactory answer must cover and what must never appear. The following template uses Chapter 1’s teaching reconciliation scenario. Situations and amounts are teaching assumptions; replace them with authorized data and traceable provenance when using it:

#Case nameInput: supply a real case and provenanceRequired pointsProhibited contentPass criteria
01Locate a routine discrepancyMarch voucher records with an RMB 12,400 discrepancy (teaching assumption; supply actual data provenance)1. Identify amount and account; 2. Locate voucher ID; 3. Stay within the evidence, without inferring the causeInvented voucher IDs3/3 points
02Missing dataA month with two days of transactions missing from the upstream system1. Identify missing data; 2. Explicitly state that a determination cannot be completed; 3. Specify the missing scopeGuessing a cause while data is missingAny prohibited item is an automatic failure
03Ambiguous wording“This amount doesn’t reconcile,” without a period or account1. Request clarification or list possible interpretations; 2. Do not arbitrarily select oneUnsupported assertions2/2 points
04Historical failureA cross-currency voucher falsely flagged during the first week after launch1. Handle currency conversion correctly; 2. Match the reference discrepancy amountRepeating the old error: adding values directly2/2 points
………………

Two design details matter. Prohibitions take priority over covered points: fabricating a voucher ID is much more serious than missing a point, hence the automatic failure. Declare pass criteria per case: tolerances differ; boundary cases may allow flexibility, while core cases require every point.

Step Three: Make It Rerunnable

Turn the case table into a data file and add the simplest script: call the model in a loop, score coverage, and output a result table. Manual scoring is entirely acceptable at first. The technical work is straightforward; the discipline is repeatability. Anyone running the same command at any time should obtain numbers using comparable definitions. The first complete run becomes the baseline. Every later prompt change, model switch, or solution adjustment compares with that baseline rather than memory.

Ten cases do not establish statistical confidence. They do achieve the crucial first transition from judging every change by feeling to having numbers after every change. Cases grow naturally with the project as new failures return as new examples. When manual scoring becomes unsustainable, Chapter 22’s layered golden set and LLM-as-judge take over. When frequent changes require a locked baseline and protection from accidental regression, Chapter 23 supplies versioning and the regression process.

Four common starting mistakes belong at the top of the case table:

  1. Including only easy cases: the model answers all ten correctly and the eval is permanently green. Discriminating between outputs is the goal, not green lights.
  2. Writing expected points as a unique answer: demanding word-for-word reproduction of your reference turns grading into answer matching.
  3. Writing the eval set once and never updating it: every new field failure should add a case. Without feedback, the eval decays into a decoration that always passes.
  4. Measuring only correctness: ignoring hallucination rate, format, cost, and latency until one of them blows up in production.

[Self-Check] Minimum Viable Eval

  • Can I explain in one sentence what this eval measures and who will read it?
  • Do inputs come from real customer data rather than invented ideal inputs?
  • Do the 10 cases include boundary cases and historical failures?
  • Does each case specify key-point coverage rather than a unique correct answer?
  • Do cases prohibiting fabrication fail automatically when that prohibition is violated?
  • Is there a baseline result for comparison after the next change?
  • Beyond correctness, which of the five quality dimensions do I measure? For those I do not measure, would I immediately notice a problem?

[Exercise] Write Your Project’s First Eval Case Table

Task: Choose an AI feature you are building or recently built: summarization, classification, extraction, or question answering. Collect 10 real inputs using Section 21.3’s 6 + 2 + 2 mix. For each, write expected points, prohibited content, and pass criteria, then produce the first baseline result table.

Guidance: Notice two gaps. First, real inputs: copying customers’ or users’ actual words often reveals more incompleteness and ambiguity than expected. That is exactly the eval’s value. Second, expected points: by the third or fourth case, you may struggle to decide what must be covered versus what would merely be desirable. This struggle makes quality standards explicit and deserves time.

Reference direction for instructors: Review three things: authentic case provenance, since idealized invented inputs are recognizable; prohibitions that capture the scenario’s most damaging failure, such as fabricating content absent from a summary’s source or forcing an extraction value under uncertainty; and layered pass criteria, strict for core cases and tolerant for boundaries. Ten identical-looking tables are a common counterexample: their author filled a template without thinking through the scenario’s failure modes.


Sources for the industry facts in this chapter: The quotation on production LLM experience is verbatim from Anthropic’s Forward Deployed Engineer job description. OpenAI’s three-phase customer-project process and building evals and improving against them during validation come from The Pragmatic Engineer’s August 2025 interview with its FDE leader. The distinction between evals and software tests, the five quality dimensions, and the minimum viable eval and case-table template are this book’s methodological development, not claims about a particular company’s internal practices. No additional industry statistics are cited. Unsourced cases, amounts, proportions, case counts, and frequencies are teaching assumptions, not customer measurements or industry benchmarks.


Independent exercise

Use fictional or authorized deidentified material. Answer independently before revealing the reference. Save chapter-21.md with versions, decisions, evidence and gaps.

Design two common, boundary, historical-error and adversarial examples each, with expectations, prohibitions, grading and versions. Separate rule tests from real model evaluation.

Required chapter artifact: evaluation-loop record

The submission for this independent exercise must contain the evidence below. The existing prompts supply content but do not replace these acceptance items.

  • O1: The evaluation-loop record draws both a forward path and a return path across success criteria, evaluation cases and version decision
  • O2: Every stage in the evaluation-loop record states its input, decision, output and owner
  • O3: The artifact contains success criteria, representative cases, per-case results and a version decision
Reveal reference feedback (answer first)

Reference feedback

Categories may use exact matching, explanations need a rubric, and secret exposure is a separate red line. Record source permission, prompt, model, parameters, data and per-case outputs. Rule exercises validate a pipeline, not model quality; exercise size is not production sufficiency.

Self-review and next steps

Check whether your decision is explicit, evidence reproducible and unknowns honestly recorded. The reference illustrates one defensible approach, not a unique answer. Seek peer review for alternatives with equivalent evidence. Mark unsupported parts unfinished and revisit the corresponding step.

Sources and boundaries

Registered sources support only the external claims used here. The evaluation-loop record, example numbers and exercise scenario are internal instructional design and require project-specific validation.

  • Define success criteria and build evaluations

    Anthropic · 2026-09-17 · Evaluation starts from observable success criteria, test cases and versioned evidence.

  • NIST AI Risk Management Framework

    National Institute of Standards and Technology · 2026-09-17 · AI risk governance requires ongoing identification, measurement, management and documentation across design, development, deployment and use.

  • FDE origins, role distinctions and OpenAI work phases

    The Pragmatic Engineer · 2026-09-17 · This interview reports FDE origins, OpenAI SA/FDE distinctions, three customer-project phases and team knowledge-sharing cadences; it is secondary interview evidence, not an industry standard.

  • Anthropic Forward Deployed Engineer role

    Anthropic · 2026-09-28 · This Anthropic role lists production LLM experience alongside prompt engineering, agent development, evaluation frameworks and deployment at scale; requirements vary across FDE roles.

Record lesson practice

Only a browser self-check is saved. No work is uploaded, reviewed or certified. Keep evidence and gaps in your own chapter file.

Further training is coming soon and currently unavailable →