Multi-agent systems: when they help and when they don't
The pitch for multi-agent systems is appealing. Instead of one model juggling everything, a team of specialists collaborates: a planner, a researcher, a coder, a reviewer. It sounds like how human organizations work, so it feels natural.
The reality is more mixed. Multiple agents solve a few problems very well, and they create a fresh set of problems that a single agent never has. This post sorts out which is which, so you can decide on evidence instead of on how impressive the architecture diagram looks.
What "multi-agent" actually means
A multi-agent system is several agent loops, each with its own context window and possibly its own tools and instructions, that pass work and information between them. If you're new to the loop itself, how the agent loop works is the place to start.
Two things differ from a single agent. First, each agent sees only its own context. What one learns isn't visible to the others unless it's passed along. Second, something has to coordinate: decide who does what, combine results, and handle disagreement. That coordination is where most of the complexity, and most of the bugs, live.
Why you'd split the work
There are four solid reasons to use more than one agent.
-
Context isolation. A task with lots of noisy exploration, like searching a big codebase or reading dozens of web pages, can fill a window with material the main task doesn't need. A sub-agent does the digging in its own window and returns a short summary. The parent's context stays clean. This is the most common and most defensible reason.
-
Parallelism. Independent sub-tasks can run at the same time. Researching five companies, checking twelve files, or trying several approaches to a problem finishes faster with several agents working at once.
-
Specialization. Different steps may need different instructions, tools, or models. A cheap, fast model can handle bulk extraction while a stronger one does the final reasoning.
-
Separation of duties. Sometimes you want an agent that can't do certain things. A reviewer with read-only access can check work from an agent that writes, and a component that handles untrusted content can be walled off from one that holds credentials. This is a security benefit as much as a modeling one, and it connects to least privilege for AI agents.
Why you might not
The costs are real and often underestimated.
Tokens multiply. Anthropic's engineering write-up on its multi-agent research system reported that agents use several times more tokens than ordinary chat, and multi-agent setups use several times more again. Check their post for the figures. The intuition is straightforward: every agent has its own window, its own instructions, and its own tool definitions, and they all get billed.
Coordination is hard. Agents can do the same work twice, leave gaps between them, or make decisions that contradict each other. A vague delegation ("look into pricing") produces a vague result.
Errors compound. If a worker returns a confident, wrong summary, the orchestrator has no way to know. Mistakes propagate through the system looking authoritative.
Debugging gets much harder. Instead of one transcript you have many, and the bug may live in the handoff between them.
Shared context is a real problem for tightly coupled work. Teams building coding agents have written about this: when several agents edit the same codebase and each makes reasonable but incompatible assumptions, the result is a mess. Tasks where every step depends on the details of the previous one don't split cleanly.
Latency can go up. Parallel workers help, but sequential handoffs add model calls and waiting.
The common patterns
Anthropic's guide to building effective agents describes a handful of patterns that show up repeatedly. Most are workflows with agentic pieces, which is a good sign, since predictable structure is easier to trust.
Prompt chaining. A fixed sequence: step A's output feeds step B, with checks in between. Not really "agents" in the strict sense, but often the right answer.
Routing. A first step classifies the request and sends it to a specialized handler, such as billing questions to one prompt and technical questions to another.
Parallelization. Run independent sub-tasks simultaneously, or run the same task several times and compare (voting). Useful for breadth and for confidence.
Orchestrator-workers. A lead agent breaks the task into sub-tasks dynamically, delegates each to a worker, and synthesizes the results. This is the pattern behind most research agents, and the one people usually mean by "multi-agent."
Evaluator-optimizer. One agent produces work, another critiques it against clear criteria, and the loop repeats. It helps when there's a checkable standard, such as "does this pass the tests" or "does this follow the style guide."
Handoffs. One agent transfers the whole conversation to another better suited to the situation, common in customer support with escalation paths.
Notice how few of these need agents to talk to each other freely. The best designs use structured handoffs, not open-ended chatter.
When it pays off
Multi-agent designs tend to earn their cost in a few situations.
Breadth-first research. "Find everything about these twenty vendors" splits naturally: one worker per vendor, each doing its own searching and reading, with a lead agent assembling the report. The work is parallel, the sub-tasks are independent, and noisy exploration stays isolated.
Large search spaces. Exploring many candidate solutions, or scanning a big repository for every use of a deprecated function.
Heavy exploration with small outputs. When the process generates enormous context but the result is a paragraph.
Tasks that need different permissions. A component that reads untrusted web content with no access to secrets, feeding a component that takes actions.
Independent verification. A separate reviewer with a fresh context catches things the author is blind to.
When to stay with one agent
Stay with a single agent, or a simple workflow, when:
- The task is small and fits comfortably in one window.
- Steps depend closely on each other, as in most coding and editing tasks.
- You need consistency of style or decisions across the whole output.
- Cost or latency is tight.
- You haven't yet measured what one agent can do.
That last point matters most. It's common to reach for multiple agents to fix a problem that a better prompt, better tools, or better context handling would solve at a fraction of the price. Context engineering and tool design usually pay off first.
How to design the handoffs
Most multi-agent bugs are communication bugs, so treat every delegation like an API contract.
Give workers a real brief. A good delegation states the objective, the specific question to answer, the output format, which tools to use, and the boundaries: what's out of scope, and how much effort to spend. "Find the current pricing tiers for Vendor X. Return a table with plan name, monthly price, and limits. Use only the vendor's own site. Stop after six searches" works. "Research Vendor X pricing" invites wandering.
Return summaries, not transcripts. The worker's raw log defeats the purpose of isolation. Ask for a compact result with the key findings and the sources.
Ask for evidence. Have workers include where each claim came from, so the lead can spot-check.
Scale effort to the task. Tell the orchestrator how many workers a simple question deserves versus a complex one. Without guidance, systems tend to over-spawn.
Use shared files for state. Instead of passing large artifacts through messages, write them to a location that others can read. It keeps windows small and gives you a trail to inspect.
Set budgets. Cap workers per task, steps per worker, and total tokens. A runaway swarm is an expensive way to find out that a limit was missing.
Where multi-agent shows up in practice
You may already be using this without building it. Coding tools such as Claude Code support subagents: specialized helpers that run in their own context, take a focused task, and hand back a result. A "search the codebase" helper or a "review this diff" helper is the isolation benefit in action. There's more on this in Claude Code skills, hooks, and subagents.
Vendors are also offering hosted infrastructure for running agents, with sandboxed code execution, checkpoints, and scoped credentials. Anthropic announced managed agents along those lines at its 2026 developer event. Whether you build your own orchestration or rent it, the design questions are the same.
A walkthrough: a small research orchestrator
Concrete beats abstract, so here's how an orchestrator-worker setup for a research question might run. The question: "Which of these five observability vendors supports on-premises deployment, and what does it cost at 50 million events a month?"
The lead agent reads the question and makes a plan: one worker per vendor, each with the same brief. It writes that brief once, carefully: find the deployment options and published pricing for the vendor, use only the vendor's own documentation and pricing pages, return three fields (deployment options, price at the stated volume or "not published," source URLs), and stop after a handful of searches.
Five workers start at the same time, each with a clean context and two tools, web_search and fetch_page. They wander through documentation, and each burns tens of thousands of tokens reading pages. None of that noise reaches the lead. Each worker returns about a hundred words plus links.
The lead agent then compares the five summaries, notices that two vendors list pricing per gigabyte instead of per event, and does a small follow-up calculation. It spot-checks one source that looks suspicious, then writes the final comparison with citations.
Total cost is higher than a single agent would have spent on one vendor, and probably lower than one agent reading all five in sequence would have spent, since that agent's window would have filled with pages from earlier vendors. Wall-clock time is roughly that of the slowest worker. That's the profile where the pattern shines: independent, parallel, noisy, with small outputs.
Now imagine the same architecture applied to "refactor our authentication module." The workers would need shared understanding of every file, their edits would collide, and the lead would spend most of its effort reconciling conflicts. Same design, wrong problem.
Failure stories to expect
Some failures come up so often that you should assume you'll meet them.
The overeager spawner. Asked a simple question, the lead launches ten workers. Cap the fan-out, and tell the lead what proportion of effort each kind of task deserves.
The telephone game. A detail gets lost or distorted as it passes from worker to lead to final answer. Have workers quote sources, and have the lead verify anything the conclusion depends on.
Duplicated effort. Two workers research overlapping topics because their briefs weren't distinct. Assign clearly separate slices.
Endless refinement. An evaluator and an optimizer keep trading edits without converging. Set a maximum number of rounds and a definition of good enough.
Silent failure. A worker hits an error and returns "no results," and the lead treats it as fact. Make workers report failures explicitly and distinguish "found nothing" from "couldn't look."
None of these are exotic. They're what happens when you hand work between components without a contract, so write the contract.
Testing a multi-agent system
Evaluate the parts and the whole.
- Test workers alone. Given a good brief, does each produce a correct, well-formatted result?
- Test the orchestrator's decomposition. Does it split tasks sensibly, without overlap or gaps?
- Test end to end on realistic tasks, measuring quality, cost, and time.
- Compare against a single-agent baseline. If the simpler version scores close at a third of the cost, you have your answer.
- Read the transcripts, especially the handoffs, where misunderstandings hide.
The approach in how to evaluate an AI agent works directly here.
A rule of thumb
Start with a single agent. Add sub-agents only when you can name the specific problem they solve, such as "the search results are flooding the context," "these five lookups are independent and slow," or "the reviewer must not have write access." Then measure whether the added complexity bought what you expected.
Multi-agent systems are a tool for particular jobs, and the teams that use them well tend to be the ones who tried the simple version first and know exactly why it wasn't enough.