TechNova is a fictional company used as a running example throughout this series.
The mistake: starting with agents instead of task shape
Imagine a fictional e-commerce company, TechNova, had started its support system with one assumption: “Let’s build an agent.”
The team gives a single agent access to everything it might need: order lookup, shipping status, cancellation, refund rules, warranty checks, customer messaging, and human approval. The demo works. The agent reads the customer’s message, checks the order, reasons through the policy, decides what to do next, and drafts a response.
Six months later, the same system is in production. It is slow, expensive, hard to debug, and brittle in ways nobody can quite explain. Some requests take two seconds. Others take forty. The on-call runbook has a page called “agent stuck in a loop.”
The uncomfortable part is that the model did not fail. The prompts are fine. The tools work. The architecture was wrong before the first prompt was written.
That is the mistake this article is about: not using an LLM, but choosing the most flexible shape before checking how much flexibility the task requires. Flexibility you do not need is not free. You pay for it in tokens, latency, debugging time, and on-call hours, every request, forever.
The architecture choice is the first decision in any project, and the most expensive one to reverse later. This article walks through five shapes a system can take, the one question that organizes the choice among them, the factors that sharpen it, and the warning signs that you reached too high on the ladder.
Five shapes for the same work
There are five practical architectures available to most production teams. They are not equally attractive options. They are a ladder. Most systems should live on the lower three rungs.
Single LLM call. One model call, one response. No agent loop, no dynamic tool choice. The model takes input, returns output, and the system either uses the output or doesn’t. The surrounding code may add validation, retries, or formatting, but the model itself is doing one task in one turn. Once the task itself becomes a designed sequence of distinct steps, rather than helper code around one model task, you have moved into workflow territory. This is the simplest possible shape and it solves more production problems than most engineers think: summarize this case, classify this ticket, draft a first reply.
Predefined workflow. A sequence of steps the developer designed. Steps may include LLM calls, code, tool calls, API requests, database lookups, retries, validation gates, parallel branches, and conditional routing. The graph of possible paths is fixed at design time. The model may make decisions inside steps, but the structure of the graph is the developer’s.
Hybrid workflow with one agentic step. A predefined workflow with one bounded decision point where the model is allowed to choose dynamically among predefined options. The workflow handles the predictable parts: authentication, data fetching, validation, the steps that have to happen in order regardless of input. The agent handles the one decision in the middle that doesn’t have a deterministic rule. Then the workflow takes over again.
Single agent. A loop where the model decides the next step at runtime based on what it has seen so far. The developer defines the available tools, boundaries, budget, and stopping condition, but the model decides the sequence. The loop is simple: observe → decide → act → check → repeat. The path emerges as the system works.
Multi-agent system. Multiple agents, each with its own scope, coordinating to solve a task that no single agent could solve cleanly alone. Specialization is the cost-justifying property: different domains, different tools, different memory, different review responsibilities. The coordination layer is itself a design problem and is rarely free.
Patterns such as chaining, routing, parallel work, and model-based evaluation can appear at several rungs. The pattern describes how work is arranged; the architecture describes who controls what happens next.
The Architecture Ladder

More runtime freedom usually means more cost, latency, complexity, and more ways for the system to fail.
The ladder is not about the technologies you use. RAG, tools, databases, queues, and APIs can appear at several rungs. What changes as you move up the ladder is who controls what happens next.
The real question: who decides the next step?
The deciding factor isn’t complexity. It’s who decides the next step.
In short:
- Workflow: the developer defines the possible paths in advance.
- Agent: the model decides the next action at runtime.
- Hybrid: code handles the predictable steps; the model handles one bounded decision.
Suppose TechNova has a rule: if the refund amount is over $500, route the request to human approval. The model might summarize the case, classify the reason, or draft the reply, but the next step was already decided by code. That is a workflow.
Now suppose the situation is less clear. The order data, the customer’s message, the warranty terms, and the shipping status point in different directions. The system cannot know in advance whether the next step should be to ask for photos, check inventory, start a warranty claim, or offer a replacement. If the model looks at the current situation and chooses what to do next at runtime, that step is agentic.
The same distinction holds even when the system becomes more complex. A workflow can contain several LLM calls, branches, retries, and routers if the developer defined the possible paths in advance. A system with one model and one tool can still be agentic if the model decides whether to use the tool and what to do next. The number of tools or model calls does not determine the architecture.
One clarification matters here: a model can influence the next step without controlling it. For example, the model might return a label such as REFUND, and code uses that label to select a predefined path. That is still a workflow. It becomes agentic when the model itself is allowed to choose the next action.
The diagram below turns this into a decision path. For cases that don’t fit one branch cleanly, the five decision factors later in the article help.
Who Decides the Next Step?

Workflows are graphs, not just pipelines
A common confusion makes the workflow option look weaker than it is. Many engineers picture a workflow as a linear pipeline: step one, then step two, then step three. Real production workflows are not linear. They are graphs.
A workflow can branch, run steps in parallel, route, retry, validate, handle errors, and involve human review. It can also call LLMs inside individual steps and use their outputs to select predefined branches.
What a workflow cannot do well is decide a new path that was not designed into the system.
That last sentence is the boundary. A workflow operates within a graph the developer drew. The graph can be dense and branching, but every path in it existed before the system ran. When a workflow hits an input it doesn’t know how to handle, it can take a default path, escalate to a human, fail with an error, or pattern-match imperfectly. It cannot choose a path that isn’t already there.
The implementation framework does not determine the architecture. A workflow engine can host agentic behavior, and a custom loop can implement a predefined workflow.
Consider a customer-support system that classifies incoming messages into refunds, order status, technical issues, and complaints, then dispatches each to a handler. Each handler can itself be a branching graph. The system is still a workflow, because every category, handler, and step was designed at build time.
The same applies to RAG: routing a query among predefined retrieval paths is still workflow behavior when those paths were designed ahead of time.
Now consider a message that mixes a refund question, a technical complaint, an emotional concern, and a deadline. The workflow can fall back to a default handler, send the case to manual review, or pattern-match on the loudest signal. What it does not have is a clean designed path for every combination of those signals. That is where runtime judgment starts becoming useful.
Structured business processes usually fit a predefined graph. Open-ended work such as research, investigation, or debugging an unfamiliar codebase is a stronger candidate for an agent because the path emerges as the system works.
Five decision factors that sharpen the choice
The diagram gives you the main decision path, but real tasks do not always fit neatly into one branch. You may know that most of the work is predictable but still be unsure whether one difficult step calls for a workflow, a hybrid design, or a full agent.
Five factors help sharpen that choice.
Path predictability. Can you draw the decision tree before runtime? If yes, a workflow can encode it. If no, the model has to choose paths at runtime.
Input variability. Is the input shape known and bounded? A bounded input space (orders, tickets, structured forms) favors workflows. An open-ended input space (natural-language conversations, exploratory research questions) favors agents.
Action range. How many distinct actions does the task need to choose among? A small fixed set fits a workflow. A large or open-ended set, especially when the choice depends on intermediate results, favors an agent.
Reliability and auditability. How easily must you be able to reconstruct what the system did and why? Regulated domains, financial transactions, anything with compliance or audit requirements: workflows give you traceability that agents don’t, by default. If you need to prove what the system did and why, the workflow’s predetermined graph is the answer.
Cost and latency tolerance. Agents typically run more LLM calls, more tool calls, and longer loops than workflows. A single LLM call is one round trip; an agent loop can easily become five to fifteen model/tool round trips before the user sees an answer. If the task budget is tight (chat-facing latency under two seconds, cost per request under a fraction of a cent), agents may be priced out before they are evaluated on capability.
The five factors don’t combine into a formula. They combine into a sense of which shape fits the task. A useful heuristic: if four of the five factors point toward “workflow,” it’s almost certainly a workflow. If four point toward “agent,” it’s probably an agent. If they split, you are likely in hybrid territory: most of the system is predictable, but one decision point isn’t.
The table below shows how the five shapes compare on each factor. The table is directional, not absolute. Hybrid systems retain strong auditability only when the bounded decision, its inputs, and its output are explicitly logged.
| Factor | Single LLM call | Predefined workflow | Hybrid | Single agent | Multi-agent |
|---|---|---|---|---|---|
| Path predictability | High | High | Mostly | Low | Low |
| Input variability | Low | Low–medium | Medium | High | High |
| Action range | None | Fixed, small | Fixed + one decision | Dynamic | Dynamic + delegated |
| Reliability / auditability | Medium | High | High if bounded/logged | Lower | Hardest |
| Cost / latency | Lowest | Low | Medium | Higher | Highest |
Single LLM calls are auditable at the input/output level, but they do not give you the same step-by-step path trace that a predefined workflow does. Agents can approach workflow-like auditability only when you invest in richer traces, explicit constraints, and explicit decision logs.
These tradeoffs are easier to design for upfront than to retrofit after the system is already in production.
Hybrid: the shape most production systems actually want
A customer writes in to TechNova:
“I bought the TechNova SmartHub two weeks ago. After the firmware update it stopped connecting. I threw away the box, but I need this working before Monday. Can you help or send a replacement?”
This is a real-shape support request. It is not a clean refund question, not a clean technical question, and not a clean replacement question. It is partly all three. It has a deadline. It has a customer with a thrown-away box. It has a firmware update as the suspected cause.
A pure workflow handles part of this case well. The system needs to authenticate the customer, fetch the order, check the purchase date, check the return window, check warranty status, and check for known issues with the firmware update. All of these steps are predictable. Every support request needs them. The graph is the same regardless of what the customer wrote.
An agentic decision step handles a different part well. Given the gathered facts, what should the system actually do? Process a return? Send a replacement? Offer troubleshooting? File a warranty claim? Ask for clarification? Escalate to a human?
The first part is rule-based. The second requires judgment because the branches overlap and the right choice depends on the conversation context, the customer’s tone, the deadline, the firmware history, and how the previous steps resolved. You could try to enumerate the rules. You would build a decision matrix with thirty rows and find it still doesn’t cover real cases. The branching logic isn’t simple enough for code and isn’t open-ended enough to need a full agent.
The hybrid shape splits the difference cleanly. The predictable steps run as a workflow. The messy decision runs as one bounded agentic step. Then the workflow resumes.
A Hybrid System in Practice

Two things make this work in production. First, the agentic step is bounded: the model chooses among a known set of next paths, not from an open space. The choice is “which of these six branches,” not “what should we do.” Second, the output is structured: the model returns a path identifier the workflow can route on, not free text.
{"next_path": "REPLACEMENT"}
That schema limits the agentic step to approved path IDs and lets the workflow treat the result as a deterministic input.
This is the shape many production support, customer service, claims processing, and routing systems actually need. The vast majority of the work is predictable. One decision point in the middle is genuinely ambiguous. A pure workflow forces you to enumerate every rule, and you will get it wrong on edge cases. Giving the model the whole process, including the parts that don’t need its judgment, means paying for that judgment on every request.
Hybrid can be a steady-state design when one part of the problem is messy and the rest is predictable.
The cost of hybrid is operational. You now have two runtimes inside one system, and the handoff between them needs to be solid: what state the workflow passes in, what the agent is allowed to return, what happens if the agent fails or exceeds its budget. For example: the workflow may pass order status, warranty status, firmware version, known-issue flag, and customer deadline; the agent may return only a structured path such as RETURN, REPLACEMENT, TROUBLESHOOT, WARRANTY, ASK_CLARIFICATION, or MANUAL_REVIEW. These aren’t glamorous engineering problems, but they’re the difference between a hybrid that ships and one that gets quietly replaced six months later.
When a full agent is justified
Not every uncertain decision needs a full agent. If only one part of the process requires runtime judgment, a hybrid design may be enough.
A full agent becomes more appropriate when several of these conditions appear together:
- The next step depends on what just happened. You cannot fully predict step four until you see what happened in step three.
- The model may need to choose among many possible actions. The choice depends on intermediate results and is too large to encode cleanly as branches.
- The task requires repeated interaction. The agent may need several rounds of acting, checking the result, and changing its approach before it can finish.
- The environment can change the plan. Errors, missing data, or unexpected results may require the model to try a different approach rather than simply retry the same step.
When all four are present, a full agent is a strong fit. When only one or two are present, a workflow or hybrid design may still be simpler.
Coding is a good example because it gets used both ways. Coding is the domain. Control flow is the architecture. The same coding task can be solved by a single LLM call (“explain this function”), by a workflow (“read issue → fetch likely files → generate patch → run tests → report”), or by an agent (“read issue → choose which files to inspect → search → open files → edit → run tests → inspect failures → choose next action → repeat”). The architecture isn’t determined by the fact that the task involves code; it’s determined by whether the next step can be designed upfront.
Agents are powerful where they fit, and more expensive across most dimensions that matter once they’re running. The boundaries an agent operates within (tools available, budget allowed, stopping condition, escalation path) aren’t optional. They are the work. Building an agent is mostly the work of constraining it.
Multi-agent: a different question
A common path through the architecture decision goes: workflow feels too rigid, so the team builds an agent; the agent feels too messy, so they build several agents to coordinate. That second step is usually wrong.
Multi-agent is not the next step after a single agent feels hard. It is a separate design decision that must earn its coordination cost through specialization, separation, or measurably better results.
Coordination is not free. Each agent has its own context, memory, and scope; the protocol between them is its own design problem. Token cost grows with every agent and every coordination turn. Sequential agents add latency one after another; parallel agents still require time to launch multiple tasks and combine their results. Debuggability is significantly worse than a single agent: failures can come from any one agent, from the coordination layer, or from the interaction between them.
Multi-agent earns its cost in a small number of cases. When the work genuinely splits across specialized domains where one model can’t hold all the context: say, a system that needs a security expert and a performance expert each reasoning about the same change with their own knowledge bases. When the work needs independent review: one agent generates, another checks, kept separate so the generator can’t coach the reviewer. When the work needs separation of authority: one agent has write access to one system, another to a different one, with the boundary enforced by design. In those cases, coordination cost is the price of admission. In most other cases, a single agent with the right tools and the right context does the same work for less.
The most common failure mode is coordination overhead on a problem that didn’t need it. Three agents pass messages back and forth to do work one agent could have done directly. The system looks architecturally impressive in design reviews; it costs three times as much, takes three times longer, and fails in ways that take three times as long to diagnose.
There is a useful parallel with how teams approached microservices a decade ago: a legitimate pattern that often got applied to problems that didn’t need it. Multi-agent has a similar risk profile. Earn the second agent. Then earn the third.
Warning signs you chose too much architecture
Over-engineering is often harder to spot than under-engineering because the system may work while quietly accumulating cost, latency, and operational complexity.
Four warning signs are worth recognizing early.
Your agent keeps running past the point where the answer was already correct. The system finds the right answer in step three but doesn’t stop. It keeps reasoning, keeps calling tools, keeps revising. By step eight the answer is the same as step three, but the user has waited twenty seconds and the system has spent ten times the cost. This usually means the stopping condition is underspecified or the agent has been given too open a goal. Sometimes it means a workflow would have been better. If the answer is reliably correct by step three, perhaps step three didn’t need an agent in the first place.
Your multi-agent system is just routing in a costume. Three agents pass messages, but the messages always flow the same direction. One classifies. One handles. One responds. There is no genuine coordination: no negotiation, no specialization that couldn’t have been a tool call, no review loop that adds value. The system would be cheaper and more reliable as a routing workflow with one or two specialist agents at the leaves.
Your agent never escalates to a human, even when it clearly should. The agent is allowed to take any action within its tool set, but the tool set doesn’t include “stop and ask.” The agent improvises through situations it doesn’t understand, produces confident but wrong outputs, and the team notices only when a customer reports it. Human escalation should be an explicit part of the system design, not an accidental fallback. If your agent doesn’t have one, you have built something more dangerous than what you needed.
Your “agent” is really doing one fixed sequence of steps the developer wrote in the prompt. The system prompt contains instructions like “first do X, then Y, then Z, then return the result.” The model follows the prompt because it’s a competent model. The system functions. But the architecture is a workflow being executed by a model that has no idea it’s a workflow. The team is paying agent prices for workflow behavior, and getting workflow rigidity wrapped in agent unpredictability. The right move is to take the steps out of the prompt and put them in code, where they belong.
These warning signs usually point to a mismatch between the task and the architecture. Before tuning the prompt or changing models, ask whether the system can move down a rung: from multi-agent to one agent, from an agent to a workflow, or from a workflow to a single call.
Default to the simplest shape that works
Start at the bottom of the ladder. Climb only when the rung below provably cannot carry the load.
Add architectural freedom only when the task forces you to. Move beyond a single call when the work requires explicit orchestration, add bounded agency when one decision cannot be encoded cleanly, and use open-ended agent loops only when the path itself must emerge at runtime.
The most expensive production agents are the ones that should never have been agents in the first place.
Three Takeaways
- Start lower on the ladder than your instincts suggest. A single LLM call or predefined workflow often solves the problem with less cost, latency, and debugging pain.
- The key question is who decides the next step. If the developer can draw the path ahead of time, use a workflow. If the model must choose the next action at runtime, you are moving into agent territory.
- Use agency only where uncertainty earns it. Hybrid is often the practical middle ground: keep predictable steps in the workflow, let the agent handle one bounded decision, then return control to the workflow.
Next: once you have chosen the agent shape, what does it take to build the loop so it can survive production? That is Part 6: Building the Production Agent Loop.
Source note: this article builds on the workflow-versus-agent distinction and the “start simple” principle from Anthropic’s Building Effective Agents (Schluntz & Zhang). The architecture ladder, the “who decides the next step” framing, and the treatment of hybrid as its own rung are this series’ own synthesis.