TechNova is a fictional company used as a running example throughout this series.
The demo loop is not enough
The basic loop is still the same: observe, decide, act, check, repeat. That is enough to show an agent working in a demo.
In production, completing the task is only part of the job. The system also has to keep the right state between steps, call tools through clear contracts, verify that important actions actually happened, enforce business rules at the backend, handle retries safely, control cost and runtime, and leave a trace of what happened.
The loop itself did not change. What changed is the structure around it.
This article focuses on that production structure. We will use one cancel-then-refund example to see where a demo implementation needs stronger state, verification, boundaries, budgets, and stop rules before it can be trusted in production.
The loop still holds
The basic loop is the same one from Part 3: observe → decide → act → check → repeat.

The same control loop from Part 3.
In a demo, observe is “read the last tool result.” In production, observe reads from working state, not just the conversation. In a demo, act is “call the tool.” In production, act calls a tool with a contract and an idempotency key. In a demo, check is “did the tool return something.” In production, check sometimes has to ask a harder question: did the world actually change the way the tool said it did? A 200 OK can mean “request accepted,” not “done.” Strictly, 202 Accepted is the HTTP status defined for a request accepted but not yet completed. In practice, production APIs also return 200 OK with an application status such as accepted, and the agent has to honor the contract of the API it actually receives. For an action like a refund, that gap matters.
That last one is the heart of this article.
The loop did not change. The implementation got stricter.
A small build: cancel, then refund
Take one concrete task. A TechNova customer writes in about order #4482: they want to cancel it and get their money back. The order is still awaiting shipment, so the cancellation is legitimate. The agent’s job is to cancel the order and issue the refund.
The naive loop is short. The agent calls cancel_order, gets back 200 OK, and calls issue_refund. Two calls, one after the other, both appear successful.
The problem is what this API’s 200 OK actually guarantees. In TechNova’s contract, it means the cancellation request was accepted, not that the order is cancelled. Between accepted and cancelled, several things can happen: the cancellation might still be queued, it might not yet have propagated to the service that owns refunds, it might have partially applied during a retry, or it might be waiting on a fraud-review hold.
So the naive loop can issue a refund against an order that is not actually cancelled yet. It has moved money based on an action that had not landed. That is the failure this article is built around:
A tool response describes the request, not necessarily the world. Before the next consequential action depends on it, the system must confirm that the expected change actually happened.
The diagram below shows both paths. The naive path trusts the 200 OK. The safe path re-reads the order status and refunds only after the cancellation is confirmed.
The rest of this article builds the pieces that make that safe path hold up in production.

The production agent architecture map
Before fixing the loop, it helps to see what this agent needs around it. A demo may have little more than a model and a tool. A production agent needs more structure.

The pieces fall into three groups:
-
The model makes one decision each turn: which action to take next, given the working state and the tools it is allowed to call.
-
The agent runtime owns the agent-side controls around that decision: working state between steps, the tool registry and contracts, the skill for recurring procedures, approval and verification gates, stop rules and budgets, and the trace log.
-
The authoritative application owns the real order state. It is not part of the agent; it is what the agent’s tool calls reach. It revalidates preconditions, enforces the rules when a change is written, and is the final authority on what actually happened.
The main production difference is in this surrounding system, not in giving the model more capability. For the cancel-then-refund case, this article focuses on tool contracts, working state, and the verification gate, while describing the other controls around them.
This task does not need external reference documents, so there is no RAG component in the map.
Tool contracts
In a demo, a tool is often just a function with a one-line description in the prompt: “cancels an order.” That is enough for the model to call it. It is not enough to call it safely.
A production tool needs a contract. It does not need to be a heavyweight spec, but it should answer six questions:
- When to use: the name and description the model reads when choosing a tool: what the tool is for, and when not to reach for it. The remaining fields protect what happens after the model decides.
- Input shape: what arguments it accepts and in what form. For example:
order_id,idempotency_key. - Output shape: what comes back on success, structurally. For example: a status or identifier that code can check, not free text.
- Failure modes: what happens when the tool cannot complete the action. Does it raise an error, return an error status, or time out?
- Idempotency expectation: is calling it twice safe? For anything that moves money or changes state, the contract should require an idempotency key so a retry does not double-apply.
- Verification method: how you confirm the action actually happened. For
cancel_order, that means re-reading the order status and checking that it reached a terminal cancelled state.
That last field is the one demos skip and production cannot. A tool’s own response is a claim about the request; the verification method is how you check that claim against the world.
Two decisions shape the idempotency key. First, what it identifies. A bare order_id is too weak because the same order can carry more than one legitimate intent. Scope the key to the operation, not the resource: something like intent, action, and resource combined, or an identifier minted once for the run.
Second, who owns it. The model may propose the operation, the runtime preserves that operation’s identity across retries, and the backend enforces idempotency. When the backend sees the same operation identity again, it returns the earlier result instead of applying the mutation a second time. Retrying the same business intent reuses the identity; starting a genuinely new operation does not. If the model generates a new key for every attempt, the backend sees a new request rather than a retry of the same operation.
A contract is also something code can check. A typed output shape, validated at runtime, limits what the agent may treat as a valid result. That lets the next step check the result mechanically instead of trusting free text. The contract belongs to building the tool, not to documentation added afterward.
The skill: packaging the procedure
Tools tell the agent what it can call. Knowledge retrieval (RAG) tells the agent what it knows. A skill tells the agent how a recurring piece of work should be done: the procedure, in order, with the checks that matter.
For the cancel-then-refund task, that procedure is worth packaging because it is easy to get wrong and likely to be reused. Here is a compact version written as a SKILL.md file:
---
name: cancel-and-refund
description: >
Cancel an order and issue any refund owed, when cancellation is legitimate.
Use when a customer asks to cancel an order and get their money back.
---
1. Confirm the cancellation is allowed:
order not yet shipped; request is legitimate.
2. Call `cancel_order` with an idempotency key.
3. Verify the cancellation landed:
re-read order status; proceed only if terminal-cancelled.
4. Get human approval for the refund amount
if it crosses the approval threshold.
5. Issue the refund with an idempotency key.
6. Verify the refund landed:
re-read refund status.
7. If verification fails, do not continue.
Wait and re-read with backoff, or escalate.
Two things about this file are deliberate.
First, it is short. A skill is the name, the description, and the procedure, not a manual. The name and description can remain available to the agent, while the full procedure is loaded when the task needs it. The procedure carries the steps that are easy to get wrong, especially the verification steps.
Second, a useful way to create a skill is to write it after a successful manual run, not before. Run the task once, note where it fails or becomes unclear, and then capture the steps that worked. A skill written from imagination records what you think the work is. A skill written from a successful run records what the work actually turned out to be.
A skill is guidance the model follows, not enforcement. The steps that must never be skipped are also enforced in code. The runtime prevents the loop from proceeding to the refund until the cancellation is confirmed, and the backend rejects a refund that is no longer valid. If the model skipped a step, the system would still refuse the unsafe action.
State
State is what the loop knows between steps that is not in the conversation text. The conversation is where the customer’s words live; state is where the facts the loop has established live. Keeping them separate is what stops the agent from re-deriving the situation from chat history on every turn, and from believing something just because it was said.
For the cancel-then-refund task, the working state holds these fields. They are not all the same kind of thing, so it helps to group them by where they come from:
- Tracked by the loop itself:
step_count,budget_remaining,approval_status,verification_status. These describe the loop’s own progress: how far it has gone, what it has left to spend, and whether approval and verification have happened. - Taken from the request:
order_id,customer_intent. - The latest raw tool output:
last_tool_result. A claim about the request, not yet checked. - Observations of state owned elsewhere:
cancellation_status,refund_status.
The last group is the one to be careful with. cancellation_status is not the tool’s response, and it is not the order’s real status. It is the most recent verified observation of that status, copied from the system that owns the order, and it can become stale the moment it is written. Copying a fact into working state does not make the loop the owner of that fact.
Working state is what this series calls short-term memory, as Part 4 defined it: context the application manages for the current case and discards when the case closes. Long-term memory is different: facts kept across cases. This build uses none. What an agent should remember between cases, and what it is allowed to keep about a person, is a governance question we take up in Part 8.
Approval and verification are two different things
It is easy to assume that a human approval step makes a consequential action safe. It does not, on its own, because approval and verification answer different questions.
Approval is about intent. A human looks at “refund $740 on order #4482” and decides whether that should happen.
Verification is about outcome. It asks whether the action actually happened in the world. Approval cannot answer that, because at approval time the action has not run yet.
So a system can have a human approve the refund, issue it, and still be wrong if the refund was based on an unverified cancellation. Approval signs off on the plan; verification confirms the result. A consequential action that requires approval needs both, in that order: approve the intent, take the action, verify the outcome, then continue.
Not every consequential action needs approval. The trigger is the risk policy, not simply the fact that an action has consequences. A system can reasonably auto-refund $5 and require sign-off at $5,000. Risk policies usually consider how much is at stake, how reversible the action is, and how ambiguous the case is.
In the running example, cancelling order #4482 before it ships is cheap to undo and may move no money. The $740 refund is where money leaves, so if there is a human gate, it belongs on the refund, and possibly only above a threshold.
Check becomes verify-before-commit
You have already seen the safe path: re-read the order before issuing the refund. This section covers what makes that re-read trustworthy, and what it still cannot guarantee on its own.
Not every step needs this. For reads and reversible, low-stakes actions, a quick check that the result looks sane may be enough. For actions such as cancel, refund, delete, or publish, check becomes two questions:
- verify: did the world actually change the way the tool’s response implied?
- commit: given that it did, is it now safe to take the next consequential step?

Verification can happen at different levels, with different degrees of independence:
-
Schema / postcondition check. Does the response structurally match what success should look like? This is cheap and worth doing, but it only checks the response. It cannot catch a tool that reports success when the world did not actually change.
-
Ground-truth re-read. After the action, query the state again. For the cancellation, re-read the order status from the system that owns it. This can catch the gap between “request accepted” and “state actually changed.”
The read still has to be trustworthy. A separate endpoint is not enough if it reads from a lagging replica or stale cache. For consequential state changes, verification should use the authoritative source or a sufficiently consistent read path. If the system provides version numbers or concurrency tokens, which help detect whether the state has changed since it was read, those can strengthen the check.
-
Independent verifier. When there is no clean state to re-read, a separate judgment can decide whether the action achieved its purpose. Sometimes that may be another model. This is more expensive, and a second model can still share the first model’s blind spots.
For cancel-then-refund, mechanism two is the right one: re-read the order status from the system that owns it.
Verification is still not a transaction. The re-read and the refund are two separate calls, and the world can change between them. A warehouse worker, a support override, or another workflow could change the order after the re-read confirms cancellation but before the refund runs.
That is why the refund must check its precondition again when it executes. The backend makes the refund conditional on the state it finds at write time and rejects it if the order is no longer cancelled.
The agent’s verification gate prevents bad sequencing. The authoritative application enforces the rule where state changes. A guarantee that exists only in the loop can be bypassed by another caller, a retry, or a race.
If the re-read comes back pending, unknown, or blocked, the loop does not commit. It waits and re-reads with backoff, keeping the same operation identity rather than sending the mutation again, or escalates to a human.
A reader’s comment on Part 3 helped sharpen this framing: confirm the world changed, not just what the tool reported.
The code shape
Here is the safe path as a structural sketch. This is the shape, not the implementation. The working version, with real error handling and backoff, lives in the repo, not the article.
# observe → decide → act → check(verify → commit) → repeat
# Consequential path: cancel an order, then refund, only after verifying.
cancel = call_tool("cancel_order", order_id=order_id,
idempotency_key=key("cancel", order_id))
# check: verify the world, not the tool's claim
status = verify_action_landed(
read=lambda: call_tool("get_order_status", order_id=order_id),
expected="cancelled", # terminal state, not "accepted"
retries=3, backoff="exponential",
)
if status == "verified":
# commit: now it is safe to take the next consequential step
call_tool("issue_refund", order_id=order_id,
idempotency_key=key("refund", order_id))
else:
escalate(order_id, reason="cancellation not verified")
Three things in this sketch are the whole point: the idempotency keys (so a retry does not double-apply), the verify_action_landed re-read between the two consequential actions, and the else branch that refuses to commit when verification fails. Everything else is detail.
In this simplified sketch, the action plus the order ID identifies the one cancel and the one refund this run performs. If the same action can legitimately happen more than once on one order, such as two partial refunds, the key also needs an operation or run ID.
One scoping note: the companion lab uses a deterministic controller so its traces are stable and it runs without an API key. Under Part 5’s definition that controller is workflow-shaped: code owns the next-action choice. It is the production structure around the loop; swap the decision seam for a model and the choice becomes agentic while the state, contracts, verification gate, budgets, and traces stay the same.
The lab also does not implement the approval gate. Approval is part of the production design described here, not something the lab exercises.
You can run the complete deterministic example, compare the safe and naive traces, and inspect the tests in the Part 6 companion lab.
Budget and stop rules
A loop that can call tools can also run away. Budget and stop rules are how you bound the cost of that.
Stop rules cover four things, not one: step count (the loop has tried too many times), cost (it has spent its budget), time (it has run too long), and silence. The last is the one demos forget. A loop can hang on a step that never errors and never returns, a call that goes idle without closing, and a stop rule that only watches for errors and step caps will wait forever. Production loops need a watchdog that treats silence as a failure too.
Budget is worth thinking of as a blast-radius control, not just a bill. A per-run ceiling, checked before expensive calls rather than reconciled afterward, turns a worst case from a runaway into a clean stop. It is a boundary on what the agent can do, written in cost instead of permissions, the same family as scoping tools and gating consequential actions.
Observability
Everything the loop does should leave a trace. For each step, log: what the agent observed, what it decided, which tool it called, what the tool returned, what the verification read came back with, the resulting state, the cost and latency, and, when the loop ends, why it stopped.
The one line worth singling out is the gap between the tool response and the verification read. That gap is where the cancel-then-refund failure lives. A trace that records both, “tool said accepted; re-read said still pending,” lets you see it later instead of guessing. This trace is also what Part 7 will build its diagnostics on; for now, the point is to capture it.
Bounded autonomy
It is tempting to read a working build and decide the lesson is “give the agent more freedom” or “add more agents.” It is not. Everything that made this agent safe was a constraint, not a capability. The loop followed a fixed control structure, but the next action was chosen at runtime from the current state. Each tool was scoped to one job and carried a contract. Consequential actions were bounded by idempotency, approval where needed, and verification. The verification step refused to commit on an unconfirmed result. The stop rules were deliberate. In a production build the model chooses actions; the structure decides which actions are possible at all. In production, reliability depends less on giving the model more freedom and more on placing clear boundaries around it.
What this article does not build
- This is not a framework. It is the shape of a loop; the working pieces live in a repo.
- This is not a vendor-specific platform. Native tool calling is the production baseline here; a protocol layer like MCP standardizes how tools are exposed; a hand-rolled ReAct loop is a teaching artifact, not a production recommendation.
- This is not a multi-agent system. One bounded agent is the whole scope here. When a second agent is justified is outside the scope of this article.
- This is not a durable execution runtime. The sketch here assumes the loop stays alive while it waits. Real settlement often does not cooperate: a cancellation or a refund can take minutes or hours to reach a terminal state, and a loop holding an open process through that window will lose its state to a timeout or a restart. If a wait can outlive the process, the execution state has to survive the process, which means an orchestration layer that persists workflow state across restarts. The structure in this article does not change when you put it there. Where the waiting happens does.
- This is not the diagnostics story (Part 7) or the governance story (Part 8). When the verification read comes back wrong, why it went wrong and how you classify it is Part 7. Who is allowed to touch what, and the audit trail, is Part 8.
What this article builds is the production loop shape: the five words, unchanged, with the implementation made strict enough to trust.
Three Takeaways
-
A production agent is the loop plus its scaffolding. A loop alone is a demo. A production agent is the loop plus working state, tool contracts, a packaged procedure, an approval boundary, a verification gate, a budget with stop rules, and a trace. The model is only one part of the system.
-
Tool success is not ground truth. A tool response describes the request, not the world. For irreversible actions,
checkhas internal structure: verify that the world changed, then commit to the next step. Verification only counts when it is independent of the action it checks. For consequential state changes, a ground-truth re-read is a strong default. -
The safest agent is the most bounded one. Reliability does not come from more freedom or more agents. It comes from constraints: scoped tools, contracts, idempotency, approval, verification, budgets, and stop rules. The model chooses actions; the structure decides which actions are possible at all.
Next: when the verification read comes back wrong, what kind of failure is it, and how do you tell? That is Part 7.
Source note: this article builds on the “keep it simple” and bounded-autonomy principles from Anthropic’s Building Effective Agents (Schluntz & Zhang). The verify-before-commit gate, the production architecture map, and the “a tool response describes the request, not the world” framing are this series’ own synthesis.