Skip to content

AFDE Training Textbook · chapter 39

Customer delivery: retail replenishment

Evidence-backed suggestions with human confirmation. Prohibitions block a cheaper candidate; verify handover and adoption separately. All numbers are teaching assumptions.

Discussion before reading

Can 94/100 with 3/20 red-line failures release? Why separate weekly activity from 30-day adoption?

Write your decision, evidence and missing information before consulting the reference below.

Reference explanation · textbook case

Teaching scenario and scoping

The customer, people and all numbers are fictional teaching material, rather than measured customer evidence or industry benchmarks.

Executives want fewer stockouts; frontline staff spend time checking inventory evidence; technical owners require data to remain on the internal network. The first scope covers everyday nonperishable goods at five stores, excluding fresh goods, new products and major promotions. Customer rules calculate quantities and the model explains evidence. Missing required fields trigger a check request; automatic purchasing is prohibited. Business owners confirm rules, technical owners confirm data access and release, and executives decide expansion.

Evaluation and candidate withdrawal

The golden set has 120 distinct cases: 100 regression cases including 10 smoke cases, plus 20 adversarial cases. Counting smoke again would incorrectly produce 130. Keep held-out validation samples, calibrate the judge with humans, then rerun baseline A and candidate B with identical data, judge and runtime settings.

Teaching checkBaseline ACandidate B
Smoke passes10/1010/10
Regression passes94/10094/100
Adversarial prohibition hits0/203/20
Average model call cost per taskCNY 0.040CNY 0.030
End-to-end P95 for the same batch3.2 seconds2.5 seconds

B treats unknown in-transit inventory as zero and violates three prohibitions. Equal aggregate scores and lower cost and latency cannot authorize release. Withdraw B before production, restore verified A configuration and recheck. No purchase orders require cancellation. Continue repairing A’s known unsupported statement and retain human review; proceed only under customer-approved controlled pilot conditions. Model call costs exclude review, deployment and maintenance and cannot establish project ROI.

Handover and adoption definitions

Customer owners independently rehearse configuration recovery, expired permissions and interface timeouts before shadow operation and a small trial. Field observation reveals scattered evidence and repeated rejection reasons; bring evidence and existing reason codes into the same task.

In a fixed cohort of 50 users, distinct people completing real tasks within complete weekly windows increase from 19/50 (38%) to 37/50 (74%). Acceptance, modification and reasoned rejection count; login and training do not. Four weekly windows cover 28 days, and cross-week overlap is unknown. They cannot establish cumulative adoption for a complete 30-day post-release window; averaging weekly rates does not solve this.

Increased adoption also does not prove inventory improvements or an independent model contribution. Stockouts, excess inventory and work time require comparable workflow baselines, accounting for training, promotions and workload. Handover enables customers to rerun evaluation, explain metrics and restore configuration. Only authorized synthetic failure reproductions enter shared team assets.