Skip to content
← Curriculum

Learning module 04 / 06

Evaluation and quality

Explain release decisions through datasets, judge calibration and comparable regression; verify software and model behavior separately.

Learning goal and artifact

Evaluation cases, a version contract and a release decision

Complete and self-check the connected workshop first, then attempt the extension with fictional or authorized material you select. Reading alone does not establish capability.

Concept summaries

Evaluation assets and prohibited behavior

Cases record provenance, inputs, expected points, prohibited behavior and gates. Distinguish routine, boundary, historical failure and adversarial cases. Prohibitions and severity precede averages; keep tuning material separate from independent validation.

Comparable versions and outcome boundaries

Models, prompts, context, orchestration, datasets and judges can change evaluation. Preserve baseline and candidate configurations; unit tests, model scores, production adoption and business outcomes need separate evidence.

Read related topic or next-step material →

Complete small task · free study

Explain why a high score can still block release

Fictional client Qinghe Equipment: operators submit fault titles; dispatchers classify and confirm assignments. Lin approves scope, Chen approves data access, and Zhou owns operations. The pilot suggests categories to dispatchers without automatic assignment. Use synthetic tickets, with no client systems or external models.

Follow along: inputs, steps and reasoning

  1. Run baseline and candidate; record file version, commands and output. Five cases cover power, mechanical, unknown, empty and unauthorized inputs; secret exposure is an independent red line.
  2. Analyze supplied synthetic model records: baseline routine 8/10, zero of two red-line failures; candidate routine 9/10, one of two red-line failures. Dataset and judge match; prompts differ. These are given fictional records, not model runs by this site.
  3. Check prohibitions before routine performance. A red-line failure blocks release; propose repair and independent validation, not just average improvement. Five cases are teaching seeds, not a sufficient production evaluation dataset.

Worked example

Candidate decision: NEEDS_FIX. Routine improvement does not excuse one of two red-line failures. Preserve failing inputs, prompt version, context and judge reasons, then perform comparable regression and validation on untuned material. Passing rule checks and these synthetic model records are separate evidence.

Try independently first

Independent task: a new candidate scores 10/10 routine but still fails one of two red lines. Write 04-evaluation.md with per-case expectations, prohibitions, versions, release decision, repair plan and unrun disclosures; attach your actual rule outputs.

After answering, reveal reference and common errors

Still NEEDS_FIX: the same red line fails, regardless of perfect routine performance. Without model access, claim synthetic-result analysis only, not real model evaluation. A real model capstone additionally needs authorized runs, outputs and human checks.

Record self-check progress

Stored only in this browser, not independent review approval. Use the downloaded workbook across devices; clearing browser data removes this record.

Extension exercise

Teaching exercise: design cases for a fictional ticket classifier, including routine, boundary, historical failure and prohibited behavior, with versions and scoring rules. Decide whether the retail case may release at 94/100 with 3/20 red-line failures. Without model runs, submit a design rather than invented results.

Use fictional or authorized deidentified material. Keep reports, implementation and checks in your own project; this site accepts no exercise uploads and performs no online grading.

Review criteria

  • Report prohibited failures separately; aggregate scores grant no exemption.
  • Baseline and candidate conditions are comparable; distinguish designs from measured runs.

If criteria are unmet, record gaps and a correction plan, then review again with evidence. Confirm sharing permissions before instructor or peer review.

Complete lessons in this module

Study the full explanations, examples and exercises in these lessons, then integrate the artifacts using the connected workshop above.

  1. Chapter 21: Everything Driven by Evals
  2. Chapter 22: Golden Set and LLM-as-judge
  3. Chapter 23: Regression and Iteration
  4. Chapter 24: Quality Governance: Red Lines, Gates, and Audits
  5. Chapter 25: Test Grids and Anti-Regression Contracts
  6. Chapter 26: Risk Control and Efficiency Metrics