# Chapter 23: Regression and Iteration

## Lesson objectives

- **O1** Can trace the forward and feedback paths among baseline version, candidate version and affected regression (artifact: version-regression comparison)
  - Evidence: The version-regression comparison draws both a forward path and a return path across baseline version, candidate version and affected regression
- **O2** Can construct a version-regression comparison with the input, decision and write-back evidence for every pass (artifact: version-regression comparison)
  - Evidence: Every stage in the version-regression comparison states its input, decision, output and owner
- **O3** Can control version variables and compare affected regressions between baseline and candidate (artifact: version-regression comparison)
  - Evidence: The artifact lists model, prompt, context, tool, dataset and judge versions with observed differences

## Learning paths

- **Novice**: Chapter transfer task: Can control version variables and compare affected regressions between baseline and candidate. Use the diagram's loop relationship to complete the first evidence item in the version-regression comparison, then check each owner and decision rule.
- **Experienced**: Apply this task to a current project before reading the explanation: Can control version variables and compare affected regressions between baseline and candidate. Submit the version-regression comparison, then check the relationship type, missing evidence and authority boundary.

## Teaching diagram

- **Required artifact**: version-regression comparison
- **Diagram kind**: regression-loop
- **Relationship semantics**: Three stages produce a result along the forward path, then observations return to change the next input; without write-back there is no loop.
- **Core concepts**: baseline version · candidate version · affected regression

![Chapter 23 teaching diagram for the version-regression comparison](/learning/diagrams/en/chapter-23.svg)

Inspect both directions of the version-regression comparison and confirm that observed results change the next pass.

The diagram moves forward from baseline version through candidate version to affected regression, then returns by a dashed path. Forward arrows produce a result; the return path carries observation and correction. If results do not alter the next input, this is a one-way flow rather than a loop.

> **Adapted public course.** Concepts, procedures and examples are adapted from the internal textbook. Case sizes, timings, improvement figures and target thresholds are illustrative, not site delivery results or universal standards. Verify tools, platforms and skills in your environment. Prompts do not grant permissions and retry counts do not authorize recovery. Preserve work and verify targets, sharing and external-state effects first.

## Lesson explanation

### 23.1 Change-Triggered Regression: Match the Layer to the Change

At the end of Chapter 22, your ticket-classification assistant had a layered set of 60 cases and a calibrated judge. Two months later, a new Monday brings another test: the model provider releases a new generation at a 40% lower unit price, and the customer's technical lead asks, "Why aren't we switching?" This is larger than a one-sentence prompt edit. Switching models moves the entire output distribution, requiring renewed measurement of correctness, hallucination rate, format, cost, and latency.

This chapter explains how to reevaluate systematically after a change and compare against a baseline, turning the decision to switch or edit into a regression assessment supported by numbers.

The principle is simple: match the evaluation layer to what changed. Change type determines the minimum regression scope. A wording edit need not always run the full set; a model replacement cannot rely on ten cases.

| Change type | Typical scenario | Impact | Minimum run | Pass threshold |
|---|---|---|---|---|
| Small prompt edit | Wording or minor interpretation adjustment | Usually local, but global drift is possible | Smoke | All correct to continue; stop and investigate failures, then confirm with regression |
| Structural prompt rewrite | Replace role instructions or output framework | Global | Full regression | Overall pass rate no lower than baseline, with no category collapse |
| Model switch or upgrade | New version or cost-reduction migration | Global; rebalance all five dimensions | Regression and adversarial | Produce a five-dimensional comparison of pass rate, hallucination rate, format adherence, cost, and latency; decide dimension by dimension |
| Sampling parameters | Adjust temperature or top-p | Moves the output distribution | Smoke and sampled regression | Stable repeated runs under fixed definitions, without systematic drift |
| Context engineering | Change retrieval strategy, tool definitions, or document batch | Global | Full regression | Same as structural prompt rewrite |
| Judge configuration | Change judge model or rubric | Changes the ruler itself | Human-machine agreement spot check first, then full rerun to rebuild baseline | Historical scores are reference only, not comparable, until the new baseline takes effect |
| Code change with no model-behavior impact | Pure display styling confirmed not to alter inputs, call paths, or user-visible model behavior | Code behavior | Chapter 25's software testing grid | All tests green and impact analysis supports eval exemption |

Two easily missed disciplines matter. Judge changes are a special row: changing the judge invalidates historical score comparisons, so recalibrate the ruler and rebuild the baseline before interpreting scores. Eval exemption depends on behavioral impact, not whether the model changed: orchestration code can alter retrieval, tool calls, retries, input assembly, or context order. Run evals appropriate to that impact as well as software tests. Only confirmed absence of effects on model inputs, call paths, and user-visible model behavior permits a testing-grid-only path. A sustainable regression discipline also specifies what need not run.

Fix comparison definitions when running, or the numbers mean nothing:

1. Record the actual run configuration: use sampling settings consistent with the intended release and declare temperature, top-p, and model version. Record seed and system_fingerprint where supported. A fixed seed may reduce variation but does not guarantee determinism; do not secretly use zero temperature when production uses another setting.
2. Compare matching versions: golden set and judge versions must match the baseline, as Section 23.2 explains.
3. Distinguish fluctuation from regression: report per-category sample and failure counts and repeated-run results; decide using a preagreed rerun count and threshold. A category decline in a small batch is an investigation signal, not automatic proof of regression. Prohibited failures trigger gate action first; do not delay blocking while awaiting statistical confirmation.

Record every rerun in a ledger: change, golden set version, judge version, and results by layer. A table is enough. Both baseline locking in Section 23.2 and the feedback loop in Section 23.3 use these records.

### 23.2 Versioning and Baseline Locking: An Anti-Regression Contract for Model Output

Regression evaluation requires a definite comparison target. Version all three components:

- Golden set: treat it as a code-level asset. Put it in git, review case changes like code changes, and explain additions and removals: which incident contributed a new case, which nondiscriminating cases were archived. Silent case changes without version changes make every movement in pass rate inexplicable.
- Judge configuration: bundle judge model, rubric text, and calibration examples into one version. Increment it if any component changes, fulfilling Chapter 22's final discipline.
- Result baseline: whenever a change passes regression acceptance and is released, lock its full result snapshot as the baseline. Future comparisons use that snapshot, not a memory that the last score was perhaps 91%.

The comparability rule follows: two eval results are comparable only with the same golden set version, judge version, and run definitions. Miss one and the comparison fails. A typical incident is a team changing its judge model, watching the same system rise from 87% to 92%, and celebrating better quality. The system never changed; only the ruler did. After a judge upgrade, first rerun the complete baseline results with the new judge to rebuild a comparable baseline, then discuss differences.

This is Chapter 25's anti-regression contract projected into two worlds: code behavior there, model output here.

| | Code anti-regression: Chapter 25 | Model-output anti-regression: this chapter |
|---|---|---|
| Protects | Code behavior | Model-output quality |
| Locks | Commit hash | Baseline snapshot: golden set version, judge version, run definitions, results |
| Proves | New code at least matches baseline: no fewer tests, no lower coverage | New configuration at least matches baseline: no lower pass rate, no category collapse |
| Asset growth | Bug fixes become tests | Failures return as golden set cases, following Chapter 22's definition of done |
| Recovery from regression | Roll back to baseline commit | Roll back to baseline prompt/model configuration |

The final row matters: prompts, model selections, and sampling parameters must be as reversible as code. Version the configuration and make restoring the baseline a single operation. Production code configuration would not live only in a chat window; neither should model configuration.

When should the baseline be promoted? When a change passes regression acceptance and is released: release locks the new baseline. Otherwise, hidden drift accumulates. Repeated comparisons with an ever older baseline obscure intermediate changes until a gap of more than ten points appears and nobody can identify where the trajectory went wrong.

### 23.3 The Eval-Driven Feedback Loop: From Proving No Harm to Deciding What Comes Next

The first two sections provide defense: establish that a change has not damaged quality. Evaluation also informs decisions. OpenAI's San Francisco FDE posting measures success through adoption, workflow impact and evaluation feedback that shapes product and model plans. It also asks engineers to share observations about model performance with Product and Research. This paraphrases one company's role, not a universal FDE standard. In this course, evaluation should produce evidence for subsequent decisions, beyond a pass-rate report.

There are two loops. The outer loop carries field evidence to model providers and product/research teams, influencing model and product evolution; Chapter 29's "Outer Knowledge Loop" covers it. The inner loop informs your own solution-iteration decisions. Start it here by assessing severity and prohibitions before all four exits below. Prohibited failures first block release or trigger withdrawal/rollback; every exit is subject to that prerequisite:

| Regression result | Decision | Action |
|---|---|---|
| At least baseline across the board | Adopt | Release, promote baseline, record in ledger |
| Overall unchanged, local category collapse | Classify severity before targeted repair | Block prohibited failures before release; withdraw or roll back if released. Otherwise repair within approved bounds and retest the failed category and affected regression scope |
| Clearly below baseline | Roll back | Restore baseline configuration; reject the change but retain its conclusion |
| Still below requirements after several iterations | Escalate | The solution may be wrong, not just the prompt: change model, architecture, or commitment scope; discuss openly with customer and team using eval data |

Classify severity and prohibitions before every exit: an unchanged overall pass rate cannot permit unauthorized actions or sensitive-data leakage. Chapter 39 rejects configuration B on this same rule. Block unreleased changes; withdraw or roll back released ones according to risk, then arrange bounded repairs. Pass rates guide subsequent decisions but cannot override hard gates or turn failure into PASS by lowering acceptance standards.

The third row is often wasted: rolled-back changes are assets too. "The new model doubles hallucinations on our ticket data and is unsuitable" is negative knowledge. Record it so a colleague considering the same model three months later can avoid repeating the work. Teams that discard negative conclusions pay repeatedly for the same dead end.

The inner loop's real discipline is citing numbers in iteration decisions. "I think the new model is better" is inadequate in a solution review. A valid contribution is: "Regression pass rate is 94% versus the 91% baseline, but hallucinations are 3.1% versus 1.8% and P95 latency adds 800 milliseconds. Use it only for classification and retain the original model for generation." These numerical decisions are what Chapter 21 means by improving against evals.

Another loop reaches the customer. Eval trends are strong communication evidence: a weekly chart of pass rate and hallucination rate, annotated with changes, is worth more than repeated assurances that the system runs well. This answers the technical lead's opening question. Deliver a five-dimensional comparison and a conditional recommendation on switching models.

### 23.4 Connecting Evals to the Six-Step Workflow: Two Acceptance Tracks

Chapter 7's six-step workflow -- decompose, issue instructions, code, accept, branch on the result, update the blueprint -- was designed for code delivery. Connecting evals does not require another process. Expand the fourth step's object: code acceptance and model-output acceptance in parallel.

| | Code acceptance track | Model-output acceptance track |
|---|---|---|
| Object | AI-generated code | Model outputs affected by the change |
| Tools | Five-dimensional checklist: function, code, boundaries, security, blueprint; plus testing grid | Smoke/regression evals compared with baseline: Chapters 21-23 |
| Outcomes | PASS / NEEDS_FIX / REBUILD | Three corresponding outcomes: at least baseline / local collapse / clearly below baseline |
| Trigger | Every milestone | Only milestones affecting model behavior: prompts, sampling, context, model changes |

Add a small connection in each of the first four steps:

1. Decompose: mark whether each milestone affects model behavior. Decide during decomposition, not belatedly at acceptance.
2. Issue instructions: include eval pass criteria for milestones affecting model behavior. Chapter 7's principle that acceptance criteria make the best instructions also holds here. "After the change, all 15 smoke cases must pass and regression categories must show no collapse" constrains iteration from the beginning.
3. Code: unchanged. Observe rather than interrupt while AI edits prompts or code.
4. Accept: retain the code track; on the model track, apply Section 23.1's trigger matrix: smoke for small changes, regression for structural ones.

Step five checks severity and prohibitions first: block or withdraw/roll back prohibited failures. Only after hard gates and all agreed conditions pass can at least baseline map to PASS. Noncritical local issues may enter NEEDS_FIX within approved bounds; after repair, rerun the failed category, affected regression scope, and smoke, not smoke alone. Clearly below baseline means REBUILD and restoration of the baseline configuration. Step six also expands: write baseline promotions and negative rollback conclusions into the blueprint. Like architectural decisions, they are project memory for the next session.

Equally important is what to omit. Milestones that do not affect model behavior, such as pure display styling confirmed by impact analysis, use only the code track. Do not run evals for ceremony. Two-track acceptance spends validation effort on actual risks.

The three pieces of Part Seven now fit. Chapter 21 makes model output measurable; Chapter 22 makes the measurement trustworthy; this chapter makes it regression-testable and integrates it into daily work. Evaluation protects the half of AI-system quality that lies outside code.

[Self-Check] Regression and Iteration

- [ ] Is there a written change-trigger matrix specifying the minimum layer and threshold for each change?
- [ ] After judge changes, do we rebuild the baseline before comparing, or compare incompatible scores?
- [ ] Are release-consistent settings and supported seed/fingerprint recorded, including the limitation that reduced variation does not guarantee determinism?
- [ ] Do rerun rules distinguish isolated fluctuations from category-wide regression?
- [ ] Are golden set, judge configuration, and result baseline versioned, with matching comparison versions?
- [ ] Are prompts and model configuration in version control, with one-step baseline restoration?
- [ ] Was the latest rollback conclusion recorded rather than forgotten?
- [ ] Which numbers supported the latest model-switch decision in a solution review?
- [ ] Are model-affecting milestones marked during decomposition and given eval criteria in their instructions?

### [Exercise] Complete a Model-Migration Regression Assessment

Task: Continue with Chapter 22's 60-case layered set. Assume the provider releases a new model and the customer requests a migration assessment. Follow Section 23.1's matrix: fix run definitions, evaluate old and new models on regression and adversarial layers using the pre-agreed repeat count, and produce a five-dimensional table of pass rate, hallucination rate, format adherence, cost, and latency. Select one of Section 23.3's four exits and write a recommendation for the customer's technical lead in no more than 200 words (the Chinese edition uses 200 Chinese characters); both editions require evidence, conditions, and a next step. Whether adopting or rolling back, add a ledger entry.

Guidance: Notice two gaps. First, comparison: the new model will rarely dominate every dimension. Better pass rates with unacceptable latency force the question of which dimension is nonnegotiable for the customer; evaluation becomes a delivery decision. Second, measurement definitions: after the first pair of runs, change one case and rerun. Watch the case-set change contaminate the pass-rate comparison. This makes Section 23.2's comparability rule concrete.

Reference direction for instructors: Check four things: declared fixed definitions, including temperature, seed, and set version; separate numbers for all five dimensions, with a single composite score unacceptable; severity and prohibition checks before selecting an exit, since even local failure can require blocking or rollback, followed by retesting the failed category and affected regression scope after bounded repair; and a conditional, bounded customer recommendation. Both "switch everything" and "do not switch" require justification. A polished comparison table without a case-set version is a common counterexample that cannot withstand scrutiny.

---

> Sources for the industry facts in this chapter: The discussion of adoption, workflow impact and evaluation feedback paraphrases the OpenAI San Francisco FDE posting checked on 2026-09-28; it is not a direct quotation. Improving against evals comes from The Pragmatic Engineer's August 2025 interview with OpenAI's FDE leader, cited in Chapter 21. The trigger matrix, versioning and baseline locking, four feedback exits, and two acceptance tracks are this book's methodological development, not descriptions of a particular company's internal practices. No additional industry statistics are cited. Unsourced cases, numbers, rerun counts, and cadences are teaching assumptions to replace with project evidence and preagreed thresholds. For the technical limitation that seed offers best-effort reproducibility without guaranteed determinism, see the [OpenAI Cookbook](https://developers.openai.com/cookbook/examples/reproducible_outputs_with_the_seed_parameter).

---

> Part Seven is complete. You can answer its three opening questions. How do you know a prompt change did no harm? A change-trigger matrix and baseline comparison. Should you switch to a new model? A five-dimensional comparison and four decision exits. How do you answer a customer asking whether the system works? Rerunnable numbers and trends. Evaluation protects model-output quality; Chapters 24-26 protect the other half through quality and risk governance: red lines, gates, the testing grid, and risk measurement. Chapter 25's testing grid and this chapter's eval regression are paired defenses of the same anti-regression contract in the code and model worlds.

---

## Independent exercise

Use fictional or authorized deidentified material. Answer independently before revealing the reference. Save `chapter-23.md` with versions, decisions, evidence and gaps.

Record code, prompt, model, data and judge versions for baseline and candidate. Plan a comparable rerun after a judge upgrade improves scores.

<!-- chapter-artifact-requirement -->
### Required chapter artifact: version-regression comparison

The submission for this independent exercise must contain the evidence below. The existing prompts supply content but do not replace these acceptance items.

- O1: The version-regression comparison draws both a forward path and a return path across baseline version, candidate version and affected regression
- O2: Every stage in the version-regression comparison states its input, decision, output and owner
- O3: The artifact lists model, prompt, context, tool, dataset and judge versions with observed differences

<details>
<summary>Reveal reference feedback (answer first)</summary>

## Reference feedback

Regrade both systems' saved outputs with the same new judge; distinguish regrading from regenerating outputs. Inspect per-case changes, red lines and cost, not just means. Changing data or models requires a comparable baseline; a new ruler is not product improvement.

### Self-review and next steps

Check whether your decision is explicit, evidence reproducible and unknowns honestly recorded. The reference illustrates one defensible approach, not a unique answer. Seek peer review for alternatives with equivalent evidence. Mark unsupported parts unfinished and revisit the corresponding step.

</details>

## Sources and boundaries

Registered sources support only the external claims used here. The version-regression comparison, example numbers and exercise scenario are internal instructional design and require project-specific validation.

- [Define success criteria and build evaluations](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests) — Anthropic, 2026-09-17. Evaluation starts from observable success criteria, test cases and versioned evidence.
- [GitHub Actions documentation](https://docs.github.com/en/actions) — GitHub, 2026-09-17. Repository workflows can automate build, test and continuous-integration tasks.
- [Reproducible outputs with the seed parameter](https://developers.openai.com/cookbook/examples/reproducible_outputs_with_the_seed_parameter) — OpenAI, 2026-09-17. Seed and system fingerprint can improve reproducibility but do not guarantee complete determinism.
- [FDE origins, role distinctions and OpenAI work phases](https://newsletter.pragmaticengineer.com/p/forward-deployed-engineers) — The Pragmatic Engineer, 2026-09-17. This interview reports FDE origins, OpenAI SA/FDE distinctions, three customer-project phases and team knowledge-sharing cadences; it is secondary interview evidence, not an industry standard.
- [Forward Deployed Engineer role — San Francisco](https://openai.com/careers/forward-deployed-engineer-(fde)-sf-san-francisco/) — OpenAI, 2026-09-28. This OpenAI role covers discovery, scoping, design, build, rollout, adoption and field feedback; it is not an industry-wide standard.
