Explain release decisions through datasets, judge calibration and comparable regression; verify software and model behavior separately.
Learning goal and artifact
Evaluation cases, a version contract and a release decision
Complete and self-check the connected workshop first, then attempt the extension with fictional or authorized material you select. Reading alone does not establish capability.
Concept summaries
Evaluation assets and prohibited behavior
Cases record provenance, inputs, expected points, prohibited behavior and gates. Distinguish routine, boundary, historical failure and adversarial cases. Prohibitions and severity precede averages; keep tuning material separate from independent validation.
Comparable versions and outcome boundaries
Models, prompts, context, orchestration, datasets and judges can change evaluation. Preserve baseline and candidate configurations; unit tests, model scores, production adoption and business outcomes need separate evidence.
Fictional client Qinghe Equipment: operators submit fault titles; dispatchers classify and confirm assignments. Lin approves scope, Chen approves data access, and Zhou owns operations. The pilot suggests categories to dispatchers without automatic assignment. Use synthetic tickets, with no client systems or external models.
Run baseline and candidate; record file version, commands and output. Five cases cover power, mechanical, unknown, empty and unauthorized inputs; secret exposure is an independent red line.
Analyze supplied synthetic model records: baseline routine 8/10, zero of two red-line failures; candidate routine 9/10, one of two red-line failures. Dataset and judge match; prompts differ. These are given fictional records, not model runs by this site.
Check prohibitions before routine performance. A red-line failure blocks release; propose repair and independent validation, not just average improvement. Five cases are teaching seeds, not a sufficient production evaluation dataset.
Worked example
Candidate decision: NEEDS_FIX. Routine improvement does not excuse one of two red-line failures. Preserve failing inputs, prompt version, context and judge reasons, then perform comparable regression and validation on untuned material. Passing rule checks and these synthetic model records are separate evidence.
Try independently first
Independent task: a new candidate scores 10/10 routine but still fails one of two red lines. Write 04-evaluation.md with per-case expectations, prohibitions, versions, release decision, repair plan and unrun disclosures; attach your actual rule outputs.
After answering, reveal reference and common errors
Still NEEDS_FIX: the same red line fails, regardless of perfect routine performance. Without model access, claim synthetic-result analysis only, not real model evaluation. A real model capstone additionally needs authorized runs, outputs and human checks.
Teaching exercise: design cases for a fictional ticket classifier, including routine, boundary, historical failure and prohibited behavior, with versions and scoring rules. Decide whether the retail case may release at 94/100 with 3/20 red-line failures. Without model runs, submit a design rather than invented results.
Use fictional or authorized deidentified material. Keep reports, implementation and checks in your own project; this site accepts no exercise uploads and performs no online grading.
Review criteria
Report prohibited failures separately; aggregate scores grant no exemption.
Baseline and candidate conditions are comparable; distinguish designs from measured runs.
If criteria are unmet, record gaps and a correction plan, then review again with evidence. Confirm sharing permissions before instructor or peer review.
Complete lessons in this module
Study the full explanations, examples and exercises in these lessons, then integrate the artifacts using the connected workshop above.