# Chapter 20: Telemetry-Driven Development and Reverse Feeding: Giving AI 'Far-Seeing Eyes'

## Lesson objectives

- **O1** Can trace the forward and feedback paths among runtime signal, causal diagnosis and evidence write-back (artifact: telemetry diagnosis loop)
  - Evidence: The telemetry diagnosis loop draws both a forward path and a return path across runtime signal, causal diagnosis and evidence write-back
- **O2** Can construct a telemetry diagnosis loop with the input, decision and write-back evidence for every pass (artifact: telemetry diagnosis loop)
  - Evidence: Every stage in the telemetry diagnosis loop states its input, decision, output and owner
- **O3** Can turn an operating signal into hypotheses, causal checks and evidence write-back (artifact: telemetry diagnosis loop)
  - Evidence: The artifact records one signal, competing hypotheses, validation action and write-back location

## Learning paths

- **Novice**: Chapter transfer task: Can turn an operating signal into hypotheses, causal checks and evidence write-back. Use the diagram's loop relationship to complete the first evidence item in the telemetry diagnosis loop, then check each owner and decision rule.
- **Experienced**: Apply this task to a current project before reading the explanation: Can turn an operating signal into hypotheses, causal checks and evidence write-back. Submit the telemetry diagnosis loop, then check the relationship type, missing evidence and authority boundary.

## Teaching diagram

- **Required artifact**: telemetry diagnosis loop
- **Diagram kind**: telemetry-loop
- **Relationship semantics**: Three stages produce a result along the forward path, then observations return to change the next input; without write-back there is no loop.
- **Core concepts**: runtime signal · causal diagnosis · evidence write-back

![Chapter 20 teaching diagram for the telemetry diagnosis loop](/learning/diagrams/en/chapter-20.svg)

Inspect both directions of the telemetry diagnosis loop and confirm that observed results change the next pass.

The diagram moves forward from runtime signal through causal diagnosis to evidence write-back, then returns by a dashed path. Forward arrows produce a result; the return path carries observation and correction. If results do not alter the next input, this is a one-way flow rather than a loop.

> **Adapted public course.** Concepts, procedures and examples are adapted from the internal textbook. Case sizes, timings, improvement figures and target thresholds are illustrative, not site delivery results or universal standards. Verify tools, platforms and skills in your environment. Prompts do not grant permissions and retry counts do not authorize recovery. Preserve work and verify targets, sharing and external-state effects first.

## Lesson explanation

### 20.1 Test code, configuration and environment hypotheses with runtime evidence

Failures may come from logic, configuration, data, concurrency or external services. Without project statistics, do not assert that most failures belong to one category. A model can analyze supplied code and logs but cannot know live state without authorized evidence. Telemetry tests hypotheses rather than presuming a cause.

Where bugs hide: five dimensions AI cannot see.

1. The Temporal Dimension. The execution order of code is full of uncertainty in multithreaded and asynchronous programming. A difference of a few milliseconds in the sequence of two events can lead to completely different outcomes -- this is the so-called "Race Condition."

- AI's blind spot: AI reads code linearly and statically by default. It cannot "see" that, in actual runtime, a network callback might suddenly interleave with a half-finished user-click handler and corrupt shared state.
- Typical bug: a frontend app fires two API requests on page load -- fetch user info (A) and fetch the shopping cart (B), where the cart depends on the user info. Due to network jitter, request B returns before A, and the code crashes because it has no user info. This bug is non-reproducible in the vast majority of cases.

2. The State Dimension. A long-running application accumulates and evolves internal state over time. A bug may only trigger after the system has run for 72 hours, processed one million requests, and memory usage hits a certain threshold.

- AI's blind spot: AI always sees the "initial state" or "ideal state" of the code.
- Typical bug: a counter service, under high concurrency, drifts further and further off because it does not use atomic operations. This problem can never be reproduced in a low-concurrency dev environment.

3. The Dependency Dimension. Your code is just one node in a vast software ecosystem, depending on the operating system, third-party libraries, external APIs, databases, cache services... A problem in any one of these propagates to your code.

- AI's blind spot: AI does not know that the payment API you call returned 500 from 3 to 4 PM because the provider's data center lost power.
- Typical bug: a Python app runs fine on the developer's macOS, but after being deployed to a Linux server, the file-path handling function crashes because it did not account for the difference between `\` and `/`.

4. The Configuration Dimension. The same codebase can behave wildly differently under different configurations (environment variables, feature flags, A/B test groups).

- AI's blind spot: AI cannot see that `FEATURE_FLAG_NEW_CHECKOUT` is set to true in production, thereby activating a brand-new, insufficiently tested checkout flow.
- Typical bug: the code runs perfectly in the test environment, but the moment it goes live, a flood of order failures erupts. Hours of investigation reveal that the production database connection pool was sized too small and was rapidly exhausted under high traffic.

5. The User Input Dimension. You can never predict how users will "break" your software -- typing Emoji, ultra-long strings, even malicious scripts.

- AI's blind spot: when AI generates code it tends to handle the "happy path," assuming users enter well-formed emails and passwords, and it will not proactively defend against a user inputting `'; DROP TABLE users; --`.
- Typical bug: an image-upload feature, when processing a filename containing a special Unicode character (such as U+202E, the right-to-left override), causes the entire filesystem API call to fail.

Conclusion: facing these "environment-dependent" bugs, it is far from enough to just throw the error log and a code snippet at AI -- this is like showing a doctor only an X-ray without telling them the patient's age, medical history, and lifestyle. Our job, through "telemetry," is to provide AI with a three-dimensional, holographic "medical record" covering time, state, dependencies, configuration, and input.

### 20.2 The Strategy for Embedding Structured Logs (Tiered Logging, Context Injection, Event-Driven)

Forget `console.log("here")`. This kind of "unstructured" log is barely readable for humans and, for AI, is almost unparseable "garbage text." They are like a pile of scribbled sticky notes rather than a well-organized report.

Structured Logging means recording log information in a consistent, machine-readable format (usually JSON). Each log entry is no longer a simple string but an object rich in "metadata":

<!-- code-example:chapter-20-E1 mode:reference -->
> **Example type: reference.** Reference fragment; it is not guaranteed to run alone. Adapt it to the lesson context, project versions and real interfaces, then validate with observed output.
```text
❌ Unstructured log:
[2023-11-01 10:30:15] ERROR: User login failed for user test@example.com from IP 192.168.1.10.

✅ Structured log (JSON format):
{
  "timestamp": "2023-11-01T10:30:15.123Z",
  "level": "ERROR",
  "message": "User login failed",
  "context": {
    "event_type": "USER_LOGIN_ATTEMPT",
    "username": "test@example.com",
    "source_ip": "192.168.1.10",
    "reason": "INVALID_PASSWORD"
  }
}
```

To AI, structured logs are data it can directly "understand" and "analyze." The embedding strategy is to design your logs the way you design APIs:

Strategy one: tiered logging, clear intent.

Do not log everything as INFO. Use standard log levels to express the importance of each entry:
- DEBUG: extremely detailed information used only for diagnosis in dev environments;
- INFO: records key "milestone" events in the application lifecycle;
- WARN: an expected, recoverable "anomaly" occurred, but the application can still keep running;
- ERROR: a serious error occurred that caused the current operation to fail; core functionality may be impaired;
- FATAL/CRITICAL: a catastrophic error occurred that crashed the entire application instance.

How AI uses log levels: when you give AI a batch of logs, it first focuses on ERROR and FATAL events to quickly locate the problem's focal point; then it uses INFO/WARN to understand the operational context.

Strategy two: context injection, rich detail.

Every log entry should carry as much "context" relevant to the current operation as possible:
- Who? (user): `user_id`, `session_id`, `tenant_id`;
- Where? (location): `service_name`, `module_name`, `function_name`;
- What? (related entity): `order_id`, `product_id`, `request_id`;
- How? (parameters): key function parameters, important fields in the request body (mind redaction).

How AI uses context: contextual information is the key to AI's "correlation analysis." When AI sees a login-failure log for `user_id: 9527` and an anomalous-order log for `user_id: 9527` in the same time window, it can establish a causal link.

Strategy three: event-driven, standardized naming.

Rather than logging a vague "did something," treat each log as a discrete "event" with a clear name:

<!-- code-example:chapter-20-E2 mode:reference -->
> **Example type: reference.** Reference fragment; it is not guaranteed to run alone. Adapt it to the lesson context, project versions and real interfaces, then validate with observed output.
```text
❌ Bad message:  "Saving user"
✅ Good event_type: "USER_PROFILE_UPDATE_STARTED"

❌ Bad message:  "User saved"
✅ Good event_type: "USER_PROFILE_UPDATE_SUCCEEDED"

❌ Bad message:  "Error saving user"
✅ Good event_type: "USER_PROFILE_UPDATE_FAILED"
```

Using a `NOUN_VERB_STATE` format is a good practice. Standardized event names let AI more easily understand the complete lifecycle of a business process.

Strategy four: logs as telemetry, not debugging.

The purpose of embedding logs is not so that you "might" need to investigate later, but to let the system "tell" its own story while it runs. This means you should proactively and generously log at the code's "critical paths" and "decision points": the entry and exit of every API request; the start, success, and failure of every important business flow; before and after every interaction with an external system. The _started / _succeeded / _failed event triplet is the most basic and powerful paradigm of structured logging.

### 20.3 Distributed Tracing: Catching Race Conditions with trace_id

Structured logging solves the richness problem of "single-point" information. But to capture bugs related to "time" and "order" (such as race conditions), we need Distributed Tracing.

The core idea is simple: assign a single unique identifier, `trace_id`, to a complete user operation (such as one API request) and carry that identifier across all services and calls.

[A simplified example]:

1. The user clicks the "Buy" button, and the frontend generates a `trace_id: "abc-123"`;
2. The frontend sends the API request `POST /orders`, carriess `X-Trace-Id: abc-123` in the HTTP header, and logs `ORDER_SUBMIT_STARTED`;
3. The backend order service receives the request, parses the `trace_id` from the HTTP header, and logs `ORDER_RECEIVED`;
4. The order service calls the payment service, continues passing the `trace_id` in the RPC request, and logs `PAYMENT_PROCESSED`;
5. The payment service processes the request and returns, logging `trace_id: abc-123` throughout.

The power of distributed tracing: when you want to investigate what actually happened in the operation with `trace_id: "abc-123"`, you get a complete operation flow, precisely ordered by time, spanning the frontend, the order service, and the payment service.

How AI uses this event stream to catch race conditions: suppose you hit the bug "a user's account was charged twice." Without `trace_id`, you would see two nearly identical payment records in the logs and could not tell whether they were one operation or two independent ones. With `trace_id`:

<!-- code-example:chapter-20-E3 mode:reference -->
> **Example type: reference.** Reference fragment; it is not guaranteed to run alone. Adapt it to the lesson context, project versions and real interfaces, then validate with observed output.
```text
10:30:15.100Z | INFO | ORDER_SUBMIT_STARTED | trace_id: a-1, ...
10:30:15.150Z | INFO | ORDER_SUBMIT_STARTED | trace_id: b-2, ...
10:30:15.200Z | INFO | ORDER_RECEIVED     | trace_id: a-1, ...
10:30:15.250Z | INFO | ORDER_RECEIVED     | trace_id: b-2, ...
...
10:30:15.500Z | INFO | PAYMENT_PROCESSED   | trace_id: a-1, ...
10:30:15.550Z | INFO | PAYMENT_PROCESSED   | trace_id: b-2, ...
```

When you hand this log to AI, its reasoning goes: identify the anomalous pattern (two independent order-submit operations started within 50 milliseconds) -> correlate the context (both operations ultimately led to a successful payment) -> form a hypothesis (the frontend did not "debounce" or one-time-lock the submit button, so a fast double-click was treated as two independent purchases) -> propose a fix (add a state lock to the submit button's click handler, disabling the button immediately after the first click until the API request returns).

trace_id is like a magic thread that strings the "pearls" (log events) scattered across various systems and time points into a clearly discernible "necklace" (operation flow). Those "devils" hiding in the gaps of time (race conditions) will have nowhere to hide under distributed tracing.

[Injection guide] Log-instrumentation plans for three major scenarios:

- Scenario one: Business Transactions -- record the end-to-end lifecycle of core business (registration, order placement, content publishing) in full. The entry point logs the `_STARTED` event and generates the trace_id; the service layer logs events at key business decision points and around external-dependency interactions; the exit point logs `_SUCCEEDED`/`_FAILED` events (tiered by business-exception WARN and system-error ERROR).
- Scenario two: Async & Background Jobs -- trace background tasks with no direct user interaction. When a task is enqueued, generate a job_id and trace_id as metadata to pass along; when a Worker picks up the task, restore the IDs and inject them into the log context; long-running tasks periodically log a "heartbeat" of progress; tasks that support retries log a `_RETRYING` event.
- Scenario three: Frontend User Interactions -- reconstruct the user's complete path of actions in the browser. Generate a session_id at session start; log core interaction events (PAGE_VIEW, BUTTON_CLICK, FORM_SUBMIT); use an axios interceptor or a fetch wrapper to automatically generate a trace_id for API requests; globally capture unhandled exceptions (window.onerror) with session "breadcrumbs" attached.

### 20.4 Reverse Feeding: Throw the Crash Log Directly at AI Instead of Describing It Verbally

Imagine going to the doctor. Description A (verbal): "Doctor, I haven't been feeling well lately, a little dizzy, and my stomach is off." Description B (data-driven): "Doctor, here is my blood pressure, heart rate, and temperature log from the past week, and here is my blood-test report. My discomfort usually hits about two hours after lunch." The answer is obvious -- description B provides objective, quantified data with time and context, from which the doctor can immediately hunt for "anomalous patterns."

It is the same when we report a bug to AI.

The inefficient "verbal description" mode:

> You: "My users report that when they upload a large file, sometimes the progress bar gets stuck at 99% and then fails, but sometimes it works. The code looks fine -- can you help me analyze what might be causing it?"

This statement makes several fatal mistakes: it is full of "uncertain" words ("sometimes," "might," "seems" -- huge noise for a probabilistic model); it embeds a "subjective judgment" ("the code looks fine" -- this unverified assumption may directly trigger AI's confirmation bias); it lacks "reproducible" context (how big a file? what format? what is the network like? -- all key environmental variables are missing). Faced with such a description, AI can only act like a search engine, listing a pile of "common causes of file-upload failure," light-years away from your specific bug.

The efficient "reverse feeding" mode:

Because the system already has well-embedded structured logs, when a problem occurs we do not "guess" -- we "extract." We pull from the log system the complete log stream for that failed operation, all carrying the same trace_id, and throw this raw, untampered data directly at AI:

> You (to AI):
> Context: We are investigating a file upload failure. I have extracted the complete, structured log stream for a failed operation, identified by trace_id: "trace-xyz-789".
>
> Your Role: Act as a Senior Site Reliability Engineer (SRE). Your task is to analyze these logs, identify the root cause of the failure, and propose a specific solution.
>
> Log Data: [paste the complete structured-log JSON array]
>
> Analysis Request: Based solely on the provided log data, please answer:
> 1. What is the precise point of failure?
> 2. What is the most likely root cause?
> 3. Propose a code-level fix for the identified service.

Why this approach is so efficient: fact-driven, not opinion-driven (you hand over cold, objective "evidence," forcing AI's analysis to rest entirely on data rather than "general knowledge"); self-contained context (the logs already carry every piece of key information needed for diagnosis: operation start, file size, chunking logic, external-dependency interactions, retry logic, final failure reason); precisely located problem (AI does not need to guess whether the problem is in the frontend, the backend, or the network).

Based on this data, AI's diagnosis will be surgically precise -- pinpointing the exact failure location (the 50th chunk upload failed), the most likely root cause (S3-side throttling or a too-short client-timeout config), and a concrete code-level fix (increase the S3 client timeout, implement exponential-backoff retry).

Stop "chatting" with AI about bugs -- start "feeding" it data.

### 20.5 Spotting False Trails in Logs: Unnecessary Fallbacks That Mask the True Problem

Feeding logs directly is powerful, but not foolproof. A poorly designed system's logs can themselves "lie." The most common "lies" come from overly broad, indiscriminate "error handling" and "fallback" logic in the code.

A robust system should keep running when errors occur, typically through try...catch blocks and degradation mechanisms (such as "falling back" to a generic hot-list when the recommendation system cannot reach the personalization engine). This is good for "user experience" but can be disastrous for "problem diagnosis -- it masks the true root cause, covering up what should be a serious ERROR ("Recommendation engine connection refused") with an INFO log that looks "normal" ("Fallback to generic recommendations").

Typical patterns of log false trails:

- Pattern one: the catch-all block. It "flattens" completely different errors -- "database timeout," "third-party API auth failure," "null pointer exception" -- into a single vague `OPERATION_FAILED` log. When AI sees this log, it has lost all the key information for judging the error's root.
- Pattern two: silent failures. After a cache update fails, it only logs a WARN and moves on without re-throwing. The main-flow logs look like everything is fine and the user operation "succeeded," but an important part of the system is already in an inconsistent state. AI analyzing the main-flow logs will be completely misled.

How to train AI to be a "log detective":

- Technique one: ask AI to find "pattern breaks." After feeding the logs, append the instruction: "In addition to finding errors, analyze the sequence of events. Are there any expected INFO logs that are missing? Is there a point where the log pattern naturally breaks from the typical success-case pattern?" This sends AI looking for "what should have happened but didn't" -- for example, no `CACHE_UPDATE_SUCCEEDED` after `CACHE_UPDATE_STARTED`, letting it infer the cache update "silently" failed.
- Technique two: cross-validate code and logs. When you suspect the logs contain a "false trail," feed the relevant code snippet along with the logs and have AI read the code inside the try block, list every specific error that might be caught, and compare them against the current problem. This is like handing the detective a suspect list to match against the crime-scene evidence.
- Technique three: proactively ask about "fallback paths." "Does this log stream suggest that any system fallback or graceful degradation logic was triggered? If so, what was the original failure that triggered this fallback?" This question directly guides AI to look for WARN-level logs or logs containing "fallback," "generic," or "default" keywords and link them to the earlier ERROR event.

A competent AI collaborator cannot blindly trust logs. You must always maintain a healthy skepticism, guiding AI to pierce the log's surface and dig out that original "first crime scene" masked by clumsy error-handling logic.

### 20.6 Closing the Loop: Run on Real Hardware -> Pull Logs -> Feed AI -> Root-Cause Reasoning -> Verify

We have now mastered the key skills of "evidence gathering" (telemetry) and "analysis" (feeding). Now let us chain them into a complete, repeatable, efficient "human-AI troubleshooting loop":

<!-- code-example:chapter-20-E4 mode:reference -->
> **Example type: reference.** Reference fragment; it is not guaranteed to run alone. Adapt it to the lesson context, project versions and real interfaces, then validate with observed output.
```text
Step 1: Reproduce & trigger    Step 2: Extract & isolate    Step 3: Feed & guide
(run on real hardware,         (pull logs with trace_id)     (structured feed +
 reproduce stably)                                            precise questions)
     |                        |                        |
     v                        v                        v
Step 5: Verify & fix      <----  Step 4: Reason & hypothesize
(verify before fixing,           (human evaluates AI's root-cause
 regression test,                hypotheses, applies critical thinking,
 update logs)                    makes the final decision)
```

- Step 1 Reproduce & trigger (your role: test engineer) -- stably reproduce the problem on "real hardware" or a Staging environment as close to production as possible. If the problem is intermittent, try to find the specific condition that triggers it; before triggering, make sure telemetry is ready and temporarily raise the log level to DEBUG if needed.
- Step 2 Extract & isolate (your role: data analyst) -- from the sea of logs, precisely extract the complete, clean log stream for that failed operation. Find the trace_id by timestamp and user_id, export all related logs from `_STARTED` to `_FAILED`, save them to a standalone file, and do not manually modify or trim -- keep them raw.
- Step 3 Feed & guide (your role: AI-interaction specialist) -- use the "reverse feeding" template, explicitly assign AI an expert role (SRE, DBA, etc.), paste the log data directly, and ask specific, closed questions. If you suspect the logs have false trails, use the "detective" techniques to follow up.
- Step 4 Reason & hypothesize (your role: senior engineer/architect) -- receive and evaluate the "root-cause reasoning" and "fix hypotheses" AI gives. AI is not a god; its hypotheses can also be wrong -- critical thinking is essential. The final decision of "which hypothesis to adopt" must be yours.
- Step 5 Verify & fix (your role: developer) -- verify first, fix second! Design a minimal experiment for the hypothesis (if the hypothesis is "database connection pool exhausted," write a script that furiously creates connections in the test environment to see if it reproduces). After the hypothesis is verified, adopt the fix, run regression tests to ensure the bug is truly fixed and no new problems are introduced. While fixing, improve the log instrumentation -- this is a valuable practice for continuously raising the system's observability.

This loop perfectly combines human and AI strengths: AI's strengths are processing massive structured data, fast pattern matching, and deep cross-domain knowledge; human strengths are deep understanding of the specific business domain, critical thinking and intuition, and the courage and responsibility to make the final decision under uncertainty.

### [Case Study] A Complete "Black-Box" Crash Fix, Reproduced End to End (a Pedagogically Recomposed Case)

Let us apply all the theory to a complete case (below is a pedagogically recomposed case: the narrative is assembled from typical fragments of multiple real debugging sessions; the numbers are demonstration magnitudes, not a record of any specific incident).

Background: you are a backend engineer at an e-commerce platform. The ops team reports: in production, the Order Service shows a few minutes of CPU 100% spikes every day around midnight, accompanied by a flood of API timeout errors, and then the service auto-restarts and recovers. Nobody knows why.

Step 1: Reproduce & trigger. The problem is scheduled, so "reproducing" is relatively easy. Your job is to prepare for "evidence gathering": confirm the log level is at INFO; temporarily raise the log level of "scheduled-task / batch-processing" modules to DEBUG; set an alert to notify you when CPU usage exceeds 95%.

Step 2: Extract & isolate. At 12:05 AM the alert fires; after a few minutes of "pseudo-death" the service is auto-restarted by the container orchestrator. You lock the time window 23:55-00:05. First you filter `level: ERROR` and find a flood of "Request timeout after 30s," but picking any trace_id at random shows only the request coming in and then nothing -- this path is a dead end. You change strategy: instead of focusing on "failed requests," look for "what the system was actively doing during that window" -- search for logs whose event name contains JOB/TASK/SCHEDULE. You find a `NIGHTLY_REPORT_GENERATION_JOB_STARTED`, followed by a flood of DEBUG logs showing this Job furiously looping over data. You isolate that Job's complete log stream.

Step 3: Feed & guide. You feed AI all the related logs from JOB_STARTED up to the crash, set the role to "Senior Go Performance Engineer," and ask: 1. based on the timestamps, find the most time-consuming operation in the loop; 2. the likely cause of the high CPU; 3. propose a Go code-level optimization.

Step 4: Reason & hypothesize. AI gives its analysis: `CALCULATING_USER_REPORT` is the obvious bottleneck -- 82 seconds for vip-user-1 (the database query FETCHING_USER_ORDERS took only 6 seconds); the system fetches 15,000 orders for a single user and then spends a long time doing in-memory aggregation in the Go application -- pulling large amounts of data into the application layer for CPU-intensive computation is a classic performance anti-pattern. It suggests pushing the computation down to the database layer with a single SQL query using aggregate functions (SUM/AVG/COUNT) and GROUP BY. You review it and find the hypothesis very reasonable, and decide to adopt it.

Step 5: Verify & fix. Verify: no need to wait until midnight -- you point locally at a test database with lots of order data, manually trigger the Job, and open the profiler at the same time. Sure enough, CPU time is 100% consumed in one large for-loop iterating over a sea of order objects. The hypothesis is verified! Fix: you tell AI "your hypothesis is correct; here is the problematic Go function, please refactor it into the SQL-aggregation approach you suggested." AI generates the new, efficient code. Regression test: the original 82-second processing drops to under 500 milliseconds, and CPU usage barely blips. Update logs: add a DEBUG log to the new aggregate SQL recording execution time, so future slowdowns are caught immediately.

The next midnight, you sleep soundly. When you wake up, the monitoring charts are calm. A "ghost" bug that haunted the team for weeks is thoroughly and elegantly resolved through a clear, human-AI, real-data-driven process.

> Part Six complete. You have now mastered the two great weapons of process constraints: stripping execution rights (the "three-step" standard flow: research -> strategize -> implement, guiding AI's "speed" toward "high-quality deep thinking") and telemetry-driven development (structured logs, distributed tracing, reverse feeding, spotting log false trails, forming the troubleshooting loop).

> Now, let us move on to Part Eight: Quality and Risk Governance -- Building a Permanent Automated Defense Line. There, you will learn to turn quality from "I think it's fine" into "the system proves it's fine."

<!-- cognitive-load-checkpoint -->
## Pause and organise: complete the minimum loop

First write one relationship between runtime signal and causal diagnosis, then place it in the telemetry diagnosis loop. Confirm that this step has an input, decision and evidence before adding evidence write-back; do not start the independent exercise until the three are connected.

## Independent exercise

Use fictional or authorized deidentified material. Answer independently before revealing the reference. Save `chapter-20.md` with versions, decisions, evidence and gaps.

For 'UI says submitted but no record exists,' propose three causes, redacted trace logs, a minimal reproduction and a verification sequence.

<!-- chapter-artifact-requirement -->
### Required chapter artifact: telemetry diagnosis loop

The submission for this independent exercise must contain the evidence below. The existing prompts supply content but do not replace these acceptance items.

- O1: The telemetry diagnosis loop draws both a forward path and a return path across runtime signal, causal diagnosis and evidence write-back
- O2: Every stage in the telemetry diagnosis loop states its input, decision, output and owner
- O3: The artifact records one signal, competing hypotheses, validation action and write-back location

<details>
<summary>Reveal reference feedback (answer first)</summary>

## Reference feedback

Check misleading UI state, API rejection and storage failure. Correlate trace IDs and record status and stages, not secrets, emails or raw tickets. Reproduce before narrowing the cause; code, configuration and platform are hypotheses. Share only minimal authorized logs and keep failures visible.

### Self-review and next steps

Check whether your decision is explicit, evidence reproducible and unknowns honestly recorded. The reference illustrates one defensible approach, not a unique answer. Seek peer review for alternatives with equivalent evidence. Mark unsupported parts unfinished and revisit the corresponding step.

</details>

## Sources and boundaries

Registered sources support only the external claims used here. The telemetry diagnosis loop, example numbers and exercise scenario are internal instructional design and require project-specific validation.

- [Google SRE: Monitoring Distributed Systems](https://sre.google/sre-book/monitoring-distributed-systems/) — Google, 2026-09-17. Monitoring should focus on actionable signals and distinguish symptoms, causes, white-box and black-box information.
- [The Twelve-Factor App: Logs](https://12factor.net/logs) — Adam Wiggins, 2026-09-17. Application logs can be treated as time-ordered event streams captured and routed by the execution environment.
