TechNova is a fictional company used as a running example throughout this series.
Part 6 ended with a question. The agent cancelled an order and issued a refund, the run reported success, and the closing line asked: when the verification read comes back wrong, what kind of failure is it, and how do you tell?
Here is what makes the question real. The team shipped the unsafe cancel-then-refund variant from Part 6: no verification gate, and a refund backend with no precondition check of its own. In review, both looked like optional hardening. The cancel request is accepted, the loop records the order as cancelled, and a refund goes out on an order that was never actually cancelled. The customer keeps both the item and the money. Nothing crashes, the run reports success, and nobody notices until the numbers do not reconcile at the end of the week.
Reaching for a better model, or wrapping the run in a retry, is guessing. The agent already wrote down what it did. This article is about reading that record: inspect the trace, name the kind of failure, decide whether a retry will actually help, and add a check so the same failure cannot come back quietly. Trace, classify, respond, eval.
The trace is the evidence
A demo trace and a production trace are not the same artifact. A demo trace, if it exists at all, is usually the conversation: what the user said, what the model said back, which tools got called. It is enough to see that the agent did something. It is not enough to see whether what it did was right.
A useful production trace should record the loop step by step, in the terms Part 6 built. For each step it holds what the agent observed, what it decided, which tool it called, what the tool returned, what the verification read came back with, the resulting state, the cost and latency, and, when the loop ends, why it stopped. Most of those fields are routine application observability: production services already record requests, responses, latency, errors, and state changes. What an agent trace adds is the model’s decision at each step and the state the loop carried forward, so you can see not only what the software called, but what the agent believed happened. The most important comparison is between the tool response and the verification read.
Part 6 singled out that gap. A tool response describes the request, not the world: accepted means the request was taken, not that the order is cancelled. That is ordinary API-contract reasoning, not something special to agents. A response guarantees only what its contract promises, whether that is a 202 or a 200 with a status field. When a later step depends on the finished state, the loop needs either a response whose contract guarantees completion or an authoritative read that confirms it.
The verification read asks the authoritative source, the system that owns the current order state, what business state actually holds. The trace should also record which source was read, because a cache or a copy of the data that updates later is not automatically ground truth. When the verification read is missing, or when it disagrees with the tool response, you have found the point where the run first diverged from the world. A trace that records that gap lets you see it after the fact instead of reconstructing it from a reconciliation report.

So the discipline is simple to state and easy to skip: when something goes wrong, read the trace before you change anything. The Part 6 trace is the right place to start because it already records the one comparison that matters. The companion lab’s naive trace records this exact failure, if you want to read one before your pager makes you.
Two places failures show up
Agent failures tend to come from one of two places: a step went wrong, or the loop around the steps went wrong. Treat these as a way to point your attention, not a list to memorize. The model is often blamed first because it is the most visible part of the system, but the trace may point somewhere else entirely.
Execution failures
Execution failures start when a step itself goes wrong. Some execution failures are familiar from ordinary software. A tool failure can be a wrong argument, a timeout, or a malformed response. An agent adds another possibility: a model decision failure, where the model chooses the wrong action or attempts an unsupported one.
The failure that matters most for this agent loop is a control-state failure: the state the loop recorded no longer matches the world. This is the case Part 6 was built around. The tool returned accepted, the loop wrote down cancelled, and the order stayed open. Nothing errored. The loop believed a claim about a request and recorded it as a fact about the world.
A control-state failure can also propagate: its root is one execution step, but its symptoms appear in later steps. Once cancellation_status: cancelled is in the loop’s state, the next step reads it as established truth. The refund step does not re-examine the cancellation; it trusts the state it inherited and fires. One wrong value becomes the false assumption later steps are built on, so authoritative verification matters before every consequential action, not only the last one. (In the Part 6 design, the backend’s own precondition check refuses this attempt; the trace still has to expose why the loop tried.)

Structural loop failures
Here the individual steps may look reasonable, but the loop that runs them goes wrong. In context degradation, the context the agent works from becomes stale or too large, so later decisions rely on outdated or degraded information. A run that drifts from its original objective usually traces back here, or to a decision step that lost the thread. In a loop runaway, cost and steps keep rising with no progress because the stopping condition was not tight enough. In a wrong escalation, the loop exits or hands off to a human when it should not, or, worse, keeps going when it should have stopped or escalated.
One structural failure hides better than the rest: the silent stall. A step can hang without erroring or returning, such as a stream that goes idle without closing, so a stopping rule that watches only for errors and iteration caps can wait forever. Production loops need a watchdog that treats silence as a failure too. Silence has to resolve into a state the loop can act on: a timeout, or the blocked condition Part 3 named. From there, policy decides whether to retry, escalate, or stop.
Telling them apart from the trace
Each kind of failure leaves a signature: a recognizable pattern across the recorded fields. For each step, five questions do most of the work: what did the agent see, what did it decide, what did the tool return, what did the verification read confirm, and what state did the loop write next.
Start with the category. If the failure is in what a step did, decided, returned, or recorded, it is an execution failure, even when the mistake happened earlier and later steps only inherited it. If the failure is in how the loop manages context, progress, stopping, or escalation across steps, it is structural. Then use the signature to find the specific failure and the first thing to check.
Execution failures: one step is wrong
| Trace signature | Likely failure | First thing to check |
|---|---|---|
| Bad arguments, tool error, timeout, or a malformed, empty, or incomplete response | Tool failure | Tool contract, required fields, and inputs; after a timeout, whether the call took effect |
| Decision or tool choice does not follow from the observed state | Model decision failure | Context and tool descriptions visible at that step |
| Tool response is treated as completion, but verification is missing or the world disagrees | Control-state failure | Verification read and source of truth |
| Inherited state contradicts the current world | Propagated control-state failure | Where state and world first diverged |
| Backend rejects a consequential action because an authoritative precondition is false | Propagated control-state or decision failure | Whether the loop marked the precondition satisfied without verifying, or the world changed after a correct check |
Structural loop failures: the loop around the steps goes wrong
| Trace signature | Likely failure | First thing to check |
|---|---|---|
| Context is stale or oversized | Context degradation | What the loop kept or dropped |
| Cost or steps rise with no progress | Loop runaway | Stopping condition and budget |
| Loop exits or escalates at the wrong time | Wrong escalation | Exit and escalation condition |
| A step starts but never completes: no response and no error | Silent stall | Timeout and watchdog coverage |
The tables are not the lesson. The lesson is the habit: a failure leaves a signature in the recorded fields, and reading the signature is faster and more honest than guessing at a cause. Classification matters because each class points to a different kind of fix. And the signature you name here is the same thing you will turn into a test at the end.
Does retry help?
Once you have named the failure, the most common reflex is to retry. But retry is a response strategy, not a diagnosis: whether it helps depends on the kind of failure you named, and retrying every failure the same way is how a system manages to be expensive and unreliable at once.
Start with the ordinary rules. Most retry policy comes from standard software engineering, and the same basic rules still apply to agents:
- Transient faults, such as a rate limit, a dropped connection, or a timeout on a read, may clear on their own. Retry with backoff, up to a limit.
- Repeatable failures, such as a bad argument or a validation error, fail the same way every time. Fix the cause; another attempt only spends budget reaching the same wall.
- Conditions the run cannot resolve safely, such as a bad credential, a missing permission, or another blocker, should not be retried blindly. Stop the loop and surface the problem with enough of the trace for someone to act. Stopping here is a control, not giving up; the real failure would be retrying into the wall, or continuing past the blocker as if it were cleared.
Then use the class you named. Several failure classes from the tables do not improve with a plain retry. A control-state failure needs a fresh authoritative read, not a second attempt at the action. A model decision failure may simply repeat, or pass by chance on the next run and hide the problem. A loop runaway needs a tighter stopping condition, not more steps. A silent stall has to become a timeout before any retry decision is possible.
Finally, retry safely. Before you replay a step, look at what it already did. A step that completed part of its work, or that called a model and got a non-deterministic answer, is not safe to re-run blindly: replaying it can repeat an effect or produce a different result than the rest of the run assumed. For side-effecting calls, this is the backend discipline of idempotency: repeating the same request should not create a second side effect. Reuse the same idempotency key for that logical operation, and do not create a new key for the retry; that turns one logical operation into two. When the outcome is uncertain, as after a timeout on cancel_order, re-read authoritative state before deciding whether another attempt is needed. Without that protection, retrying issue_refund can issue the refund twice.
The safer move overall: classify what failed, preserve the work that already succeeded, and retry only the part that genuinely needs it.
Add an eval so it does not come back
You have read the trace, named the failure, and chosen a response. The last step is making sure the same failure cannot return quietly on the next deploy. In ordinary software, that means writing a regression test after fixing a bug. For an agent the move is the same, turning the failed run into an eval, with two added considerations: the path and the outcome may need separate checks, and one passing run does not establish reliability because an agent is not deterministic.
The trace already contains the failed case, so start there. Turn the failure into a specified task (cancel an order, verify the cancellation, then issue one refund), define success as something a grader can check rather than a vague sense that it worked, run it across several trials, and keep it in the regression suite so it runs on every change.

For the cancel-then-refund failure, that means two graders. A trace grader checks the path: the loop never attempts issue_refund before an authoritative verification read confirms cancelled. An outcome grader checks the resulting world state: the backend never creates a refund unless the order is actually cancelled, and a successful run creates exactly one refund. The first catches bad sequencing even when the backend protects the money; the second verifies the final enforcement boundary. Together, they turn the failure you saw in production into something the suite catches before it ships again.
The two checks catch different problems. Checking only the outcome can miss unsafe or invalid behavior along the way: the agent reached the right answer but used the wrong tool, leaked data, retried a dozen times, or skipped an approval. Checking only the path can be too rigid, rejecting a valid run that reached the right result by a different route. So grade the outcome where you can, inspect the trace when the path is what matters, and do not force one exact sequence unless the sequence is part of the requirement. (Tools use terms like transcript, trace, and trajectory differently; here, trace means the recorded path of the run and outcome means the resulting world state.)
Structural failures can be turned into checks too, though the assertion looks different. For a step that stalls, the test is not about a final answer; it is that the run stops with a defined blocked state or surfaces the failure within its turn, time, or retry budget rather than hanging. The check is on the path and the stopping behavior, not the outcome.
You do not need a large suite to start. A useful first set is on the order of twenty to fifty real tasks, drawn from the checks you already run by hand, the failures you have actually seen in production, bug reports, and the edge cases you know are dangerous. The point is not coverage of everything. The point is that the failures you have already paid for become failures you never pay for twice.
Three Takeaways
-
A well-instrumented trace helps distinguish execution failures from structural ones. The recorded fields carry a signature; reading it beats guessing at a cause or reaching for a bigger model.
-
Retry is a separate decision from diagnosis. The failure class decides whether a retry helps, hurts, or just spends budget arriving at the same wall. Classify first, then choose the response.
-
Every production failure is a test case waiting to be written. Turn the trace into a task and a grader, run it across several trials, and add it to the suite before the next page fires.
Next: even a correctly running loop can fail at its boundaries, including what the agent may see, do, and remember, and a wrong action may be induced by untrusted input or inherited from poisoned memory. That is Part 8.
Source note: this article builds on the evaluation concepts in Anthropic’s Demystifying Evals for AI Agents, and on the human-oversight, simplicity, and tool-design principles in Building Effective Agents (Schluntz & Zhang). The failure framing, the trace-signature reading, the retry classification, and the diagnostic workflow are this series’ own synthesis.