Skip to content

Free public course · 9/40

Chapter 9: Inspection-Driven Development: Define Standards Before Letting AI Write Code

Lesson objectives

Required artifact: acceptance branch-decision table

  1. O1 · Can define the condition for moving from acceptance criteria into observed evidence (artifact: acceptance branch-decision table)

    Evidence: The acceptance branch-decision table states the entry condition from acceptance criteria to observed evidence

  2. O2 · Can construct a acceptance branch-decision table with branches, vetoes and accountable owners (artifact: acceptance branch-decision table)

    Evidence: The acceptance branch-decision table contains at least one proceed branch, one stop branch and their owners

  3. O3 · Can choose pass, repair or rebuild from observed evidence (artifact: acceptance branch-decision table)

    Evidence: The artifact records criteria, observations, branch choice and unsupported inferences for one case

Before this lesson: Chapter 8: The Three Disciplines and One Supplementary Principle

Novice path

Chapter transfer task: Can choose pass, repair or rebuild from observed evidence. Use the diagram's decision relationship to complete the first evidence item in the acceptance branch-decision table, then check each owner and decision rule.

Experienced path

Apply this task to a current project before reading the explanation: Can choose pass, repair or rebuild from observed evidence. Submit the acceptance branch-decision table, then check the relationship type, missing evidence and authority boundary.

Chapter 9 teaching diagram for the acceptance branch-decision table
Follow the conditions in the acceptance branch-decision table and make a traceable stop-or-proceed decision at the gate.
Diagram text description

The diagram starts with acceptance criteria, branches through the conditions in observed evidence, and reaches a gate controlled by PASS / fix / rebuild. Every case must end at stop or proceed; missing evidence and red-line failures cannot be offset by other strengths.

Relationship semantics: Inputs follow conditional branches into an explicit gate; stop and proceed are exclusive, and red-line failures cannot be offset by other scores.

Adapted public course. Concepts, procedures and examples are adapted from the internal textbook. Case sizes, timings, improvement figures and target thresholds are illustrative, not site delivery results or universal standards. Verify tools, platforms and skills in your environment. Prompts do not grant permissions and retry counts do not authorize recovery. Preserve work and verify targets, sharing and external-state effects first.

Lesson explanation

9.1 Why Acceptance Standards Are a “Hard Constraint”

In traditional development, TDD (test-driven development) writes test cases first, then you write code to pass the tests. In the AI era, this logic has been pushed to its ultimate evolution: define the “acceptance standards” first, then let AI write the code.

If you only say, “implement a user login feature,” AI will produce countless wildly different implementations, and there is a very high chance it misses the edge cases you care about. The correct “acceptance-driven” instruction should look like this:

“Next, we will implement user login. Your code must satisfy the following acceptance standards:

  1. Accept the account and password;
  2. The password must be verified against a Bcrypt hash;
  3. On successful verification, issue a JWT containing the User ID, with an expiration of only 2 hours;
  4. Any failure must uniformly throw a 401 AppError. If the boundaries above are clear, start writing the code.”

Why does the order matter so much? Because acceptance standards are a “constraint” for AI. When AI knows “my code is not done until it passes these 5 tests,” its generation strategy changes from “write code that looks correct” to “write code that can pass these 5 tests.” The latter is far more reliable than the former.

Let us look at a comparison:

  • Vague instruction: “Implement the user login endpoint.” AI may ignore password encryption, token expiration, and error handling — it writes a login endpoint that “works,” but you cannot be sure whether it is “secure to use.”
  • Instruction with acceptance standards: “Implement the user login endpoint. Acceptance standards: 1. Passwords must be compared using a bcrypt hash; 2. On success, return a JWT containing user_id, valid for 2 hours; 3. On failure, return a uniform 401 AppError; 4. After 5 consecutive failures, lock the account for 30 minutes.” The code AI generates will precisely cover these 4 standards, because acceptance standards are a “hard constraint” — AI knows these items will be checked one by one.

Acceptance-driven development has a hidden benefit as well: it forces you to think through “what counts as done” before coding. In many cases, while writing acceptance standards you discover the ambiguities in the requirements — what status code should a wrong password return? How does an account get unlocked after being locked? The answers to these questions must be decided before coding — if you are not sure, AI will decide for you, and its choice may not be what you want.

There is also an iron rule that must be followed: code that has not fully passed acceptance is poor-quality material that can shatter at any moment — no matter how perfect it looks, it must never be merged into the project, and it must never become the contextual dependency for the implementation of the next phase.

9.2 Three Defense Lines: Functional Acceptance, Architecture Acceptance, Security Acceptance

The core of the acceptance system is the design of “three defense lines.” Why three? Because errors in AI-generated code have three levels, and you need three levels of detection methods.

Level 1: The Functional Defense Line — detects “whether the code runs.”

This is the most intuitive check. You write test cases, run the tests, and look at the results. If you ask AI to implement “user registration,” it writes code, and you run the tests — registration succeeds, the password is stored in the database, and login works. The functional defense line passes.

But the functional defense line has a blind spot: it only checks “whether the code runs as expected,” not “whether the code runs the right way.” Your registration feature runs, but the password is stored in plaintext — functional tests cannot discover this problem. Because the input of a functional test is “username + password” and the output is “registration succeeded,” the test cases will not inspect what is stored in the database.

Level 2: The Architecture Defense Line — detects “whether the code follows the blueprint.”

This is a defense line unique to AI coding. AI can easily “casually” change things it should not change while implementing a feature — for example, to fix a bug it directly modifies the database table structure. Functional tests cannot see this, because the feature is still correct, but the architecture has drifted.

The detection method for the architecture defense line is simple: compare the blueprint file (CONTEXT.md) with the actual code. If the blueprint states “passwords must be encrypted with bcrypt,” the architecture defense line checks whether the actual code calls bcrypt; if the blueprint states “errors must uniformly throw AppException,” the architecture defense line checks whether the actual code uses AppException.

Level 3: The Security Defense Line — detects “whether the code introduces security risks.”

This is the defense line most easily overlooked. AI-generated code is often “functionally correct but security-fragile.” The reason is straightforward: AI’s training data contains a large amount of code that “works but is not secure” — it learned how to “write features,” but not how to “write secure code.”

Let us look at an example of how the three defense lines work together. Suppose the same bug — the login endpoint returns a 500 error when the password is wrong:

  • The functional defense line sees: the returned status code is wrong and needs fixing;
  • The architecture defense line sees: the error-handling logic is not within the unified exception-handling layer and needs refactoring;
  • The security defense line sees: the error message directly leaks database connection information and needs to be fixed.

The same bug, and the three defense lines see three different problems at three different levels. Fixing only the first level leaves the problems of the second and third levels unresolved. This is why three defense lines are needed — each level covers the blind spots that the previous level cannot reach. Novice acceptance usually focuses only on the first level; experienced developers check all three.

9.3 The Acceptance Decision Tree: PASS / NEEDS_FIX / REBUILD

After acceptance is complete, you need to make a decision based on the conclusion. That is the acceptance decision tree — it is “the most counterintuitive yet most important” step of the entire methodology.

Example type: pseudocode. Pseudocode; it is not executable. It expresses decision order only, so implementation must supply real interfaces, authority and error handling.

Acceptance result
├── PASS ──────→ git commit to lock in, proceed to the next milestone
├── NEEDS_FIX ─→ Have AI fix it (at most twice), then re-accept
│                  └─ Still failing after two fixes → escalate to REBUILD
└── REBUILD ───→ use Section 10.4's safe recovery path to return to an accepted foundation, then reissue a more precise instruction

PASS is easy to understand — acceptance passes, commit the code, and proceed to the next milestone.

NEEDS_FIX is also easy to understand — acceptance finds problems, AI fixes them automatically, and acceptance is rerun. But there is an important limit here: NEEDS_FIX can be executed at most twice. If the problem persists after two fixes, it is automatically escalated to REBUILD. Why? Because AI in “fix mode” easily falls into confirmation bias — it keeps patching on the wrong foundation and the more it patches, the messier things get. If you let it keep trying on the same problem, you may end up with code that “passes acceptance but has introduced more problems.”

REBUILD is the most counterintuitive yet most important branch. In traditional development, when code is wrong, you fix it; but in AI coding, the cost of “fixing” may exceed the cost of “rewriting.” The decision logic of REBUILD is: when AI has already shown signs of “chaos” (still failing after two fixes, a fix introducing new problems, or code size growing abnormally), abandon the current work immediately, roll back, and rebuild. Do not try to “fix it one more time.”

9.4 The Three Signals of Architecture Drift: Foundation Tampering, Over-Engineering, Uncontrolled Growth

Architecture drift is the most hidden and destructive problem in AI coding — while implementing a feature, AI may “casually” modify places it should not have modified. Recognizing the three signals of architecture drift is the most important capability in acceptance.

Signal 1: Foundation tampering.

If AI modifies the following types of code, you should REBUILD immediately: database connection configuration, authentication and authorization logic, global middleware, core utility functions, and shared data model definitions.

  • Why it happens: AI finds that implementing the current feature “requires” modifying foundation code. But usually that is not a real need — it is AI taking a shortcut.
  • Damage: this behavior of “robbing Peter to pay Paul” directly causes other modules that depend on the underlying model to collapse completely at some point in the future.

Signal 2: Over-engineering.

AI introduces the following unnecessary complexity: adding abstraction layers for simple scenarios (interfaces, factories, the strategy pattern), introducing third-party libraries the project does not need, and adding configuration options that the current feature does not need.

  • Why it happens: AI tends toward “just in case” design rather than “just enough” design.
  • Damage: merely to handle an edge case with an extremely low probability, AI suddenly brings in an extremely complex third-party library, or forcibly applies heavyweight design patterns such as an event bus or reflection — digging a huge crater just to cover a small pothole.

Signal 3: Uncontrolled growth.

AI piles up too much code in a single file: a component file exceeding 300 lines, a utility file containing multiple unrelated functions, or a single API route handling multiple unrelated requests.

  • Why it happens: when “appending code,” AI does not proactively refactor. It directly adds new features on top of the existing file, causing the file to bloat.
  • Damage: large numbers of if-else nesting appear, and logic that should have been abstracted and reused is mechanically copy-pasted — the “entropy increase” of the code has reached a critical point.

Detection methods:

  1. Use git diff to review the list of changed files — unexpected file modifications may indicate foundation tampering;
  2. Check the length of newly added files — a new file exceeding 300 lines may indicate uncontrolled growth;
  3. Check newly added dependencies — an unexpected dependency in package.json may indicate over-engineering;
  4. Compare against the blueprint’s directory structure — code placed where the blueprint did not agree on may indicate architecture drift.

[Template] Acceptance Checklist and Acceptance Report

Acceptance checklist (adjust it to the actual situation of your project):

Functional acceptance

  • □ All requirement points have been implemented
  • □ The main flow runs correctly
  • □ Edge cases are handled (empty data, abnormal values, extreme cases)
  • □ UI interactions match expectations (loading states, error prompts, empty states)

Code acceptance

  • □ Code style is consistent with the project (indentation, naming, comments)
  • □ No obvious code quality issues (duplicate code, overly long functions, unreasonable naming)
  • □ No dead code (uncalled functions, unused variables)
  • □ Error handling is reasonable (never swallow exceptions, never expose sensitive information)

Architecture acceptance

  • □ No core code that should not have been modified was modified
  • □ No unnecessary dependencies or abstractions were introduced
  • □ No single file has grown excessively (suggested cap: 300 lines)
  • □ New code is consistent with the project’s directory structure
  • □ API design follows the project’s conventions

Security acceptance

  • □ User input is validated or escaped
  • □ Sensitive endpoints have access control
  • □ No hardcoded secrets or credentials
  • □ Database queries use parameterized queries or an ORM
  • □ No data fields that should not be exposed are returned

Acceptance report template:

Example type: reference. Reference fragment; it is not guaranteed to run alone. Adapt it to the lesson context, project versions and real interfaces, then validate with observed output.

## Acceptance report: Milestone X

### Conclusion: PASS / NEEDS_FIX / REBUILD

### Functional acceptance
- [x] All requirements implemented
- [x] Main flow works
- [ ] Edge case: empty search keyword not handled

### Code acceptance
- [x] Code style consistent
- [x] No duplicate code
- [x] Error handling reasonable

### Architecture acceptance
- [x] No core code tampered with
- [x] No unnecessary dependencies introduced
- [x] Directory structure compliant

### Security acceptance
- [x] Input validation complete
- [x] Access control in place
- [x] No hardcoded credentials

### Fix suggestions (only if NEEDS_FIX)
1. Add an empty-string check in the search function
2. Return the full list when the search keyword is empty

Independent exercise

Use fictional or authorized deidentified material. Answer independently before revealing the reference. Save chapter-09.md with versions, decisions, evidence and gaps.

Write two functional, architectural and security checks for the ticket API. Decide whether a functionally correct but unsafe candidate can advance.

Required chapter artifact: acceptance branch-decision table

The submission for this independent exercise must contain the evidence below. The existing prompts supply content but do not replace these acceptance items.

  • O1: The acceptance branch-decision table states the entry condition from acceptance criteria to observed evidence
  • O2: The acceptance branch-decision table contains at least one proceed branch, one stop branch and their owners
  • O3: The artifact records criteria, observations, branch choice and unsupported inferences for one case
Reveal reference feedback (answer first)

Reference feedback

Check valid and empty inputs, classifier independence from UI and stable API contracts, and redacted logs and access boundaries. Correct classification with a leaked key is NEEDS_FIX and blocks release. REBUILD requires structural evidence and an approved recovery plan.

Self-review and next steps

Check whether your decision is explicit, evidence reproducible and unknowns honestly recorded. The reference illustrates one defensible approach, not a unique answer. Seek peer review for alternatives with equivalent evidence. Mark unsupported parts unfinished and revisit the corresponding step.

Sources and boundaries

Registered sources support only the external claims used here. The acceptance branch-decision table, example numbers and exercise scenario are internal instructional design and require project-specific validation.

Record lesson practice

Only a browser self-check is saved. No work is uploaded, reviewed or certified. Keep evidence and gaps in your own chapter file.

Further training is coming soon and currently unavailable →