AI-Driven Developmentadvanced9 min

Subagents & Multi-Agent Orchestration

When one agent isn't enough, an orchestrator fans the work out to parallel subagents and synthesizes what they find.

A single AI agent is remarkably capable, but it has one body and one desk. Everything it's working on has to fit inside its single context window, and it does its tasks one after another. Hand it a job that's genuinely wide — "audit every file in this repo for hardcoded secrets" — and you feel the limits: the relevant material won't all fit at once, and reading it serially is slow. The fix is to stop thinking of one agent and start thinking of a team.

From one worker to a crew

Multi-agent orchestration introduces a lead agent — the orchestrator — whose job is not to do the work itself but to divide and delegate. It breaks the task into independent pieces and spins up a subagent for each one. The subagents run in parallel, each tackling its slice at the same time. When they finish, the orchestrator collects their findings and synthesizes them into a single result for you. It's the manager-and-team pattern, applied to AI.

The quietly important detail is that each subagent gets its own fresh context window. A subagent searching the billing module isn't carrying around everything the auth-module subagent read. That isolation is what makes the approach scale: instead of one window straining to hold the whole repo, you have many windows each holding just one focused chunk. The orchestrator only needs the summaries the subagents send back, not their raw working memory.

How it works: fan out, merge, verify

The shape is always the same. The orchestrator fans out, launching one subagent per piece with a clear, self-contained brief: what to look at, what to produce, and what not to touch. Then it fans in, reading every subagent's report and merging them. Finally it verifies the merged result — running the tests, or handing it to a separate reviewer subagent that has no shared context to bias it.

The timing follows from that shape. The subagents' work overlaps, so it costs only as long as the slowest piece. The split, the merge and the check don't overlap with anything, so they add up. Four 10-minute pieces don't finish in 10 minutes; they finish in the split, plus 10, plus the merge and the check.

Step through one job below. You'll predict when it finishes before you see it. Then flip the pieces to overlapping — the same four modules, but every change also edits one shared file — and watch what happens at the merge.

Check yourself

An orchestrator spends 3 minutes splitting a job into 5 independent pieces. The subagents take 6, 8, 8, 9 and 12 minutes, running in parallel. Merging and checking take 5 minutes. When is the job done?

Note

In our stack — Claude Code can act as an orchestrator and launch subagents — separate Claude instances, each with its own fresh context window — to work in parallel. A common use is breadth: spin up several subagents to comb different parts of a large codebase at once, then have the lead agent synthesize their reports. Another is an independent reviewer subagent that checks the main agent's output without inheriting its context. Each subagent runs on one of Anthropic's Claude models, and you can mix model sizes — a lighter model for wide search, a stronger one for synthesis.

Split along seams that don't overlap

Parallel agents can't see each other's work. Each one starts from the same snapshot of the code and edits its own copy. That's harmless when the pieces live in different places — four modules, four folders, four findings lists. It breaks down the moment two pieces touch the same file: a shared routes table, a config file, a types file everyone imports. The first edit to land changes the file; every other edit was made against the old version and no longer fits. Something has to be redone, and the redo can't run in parallel either, or it would clash again.

So before fanning out, look for the seams. Good splits follow boundaries the code already has: one module, one service, one directory or one question each. If every piece needs a small change to a shared file, pull that change out — have the orchestrator make it first, or last, by itself — so the parallel pieces really are independent.

Watch out

Overlapping pieces cost you twice. If two subagents edit the same file, you pay for both edits, then pay again to redo one of them on top of the other — serially. A fan-out with overlaps can end up slower than one agent and more expensive. Split by boundaries that don't share files, and keep shared edits with the orchestrator.

The cost: coordination and tokens

Orchestration isn't free, and treating it as a default is a mistake. Every subagent burns its own tokens: each one re-reads the brief, the conventions and whatever shared code it needs before it does anything useful. So four subagents don't cost a quarter each — together they usually cost more than one agent doing the same job, and in the example above (illustrative numbers) the clean split used about 1.35× the tokens to finish more than twice as fast. There's coordination overhead too: the orchestrator has to write clear sub-tasks, wait on the slowest subagent, and stitch together results that might disagree. For a small, linear task — fix one function, rename one variable — all of that is pure waste, and a single agent is faster and cheaper.

So the rule of thumb is simple: reach for multiple agents when the task is wide and splits cleanly, or when it needs an independent check, and stick with one agent when it's narrow, sequential, or tangled through shared files. Used well, orchestration buys you breadth and a second, unbiased opinion — which is exactly what you want when you move on to verifying AI output.

Check yourself

Which of these jobs is the best fit for fanning out to parallel subagents?

Key takeaways

  • A single agent shares one context window; big or broad tasks can overflow it or get tangled.
  • An orchestrator splits the task into pieces and hands each to a subagent that runs in parallel.
  • Each subagent gets its own fresh context window, so they don't crowd each other out.
  • The orchestrator then merges the subagents' results and verifies them before calling the job done.
  • Total time is the split, plus the slowest piece, plus merging and checking — not the work divided by the number of agents.
  • Parallel only pays when the pieces are independent: pieces that edit the same files conflict and have to be redone one at a time.
  • The cost is coordination overhead and more tokens — don't reach for it on simple tasks.

Keep going