← Back to blog
    September 29, 20269 min read

    Building a Multi-Agent OS with Claude: Architecture, Tools, and Workflows for Production AI Agents

    Claude as the orchestrator of a multi-agent stack is worth it for exactly one kind of work: parallel, breadth-first tasks. Anthropic measured a 90.2% improvement at 15x the token cost - here is the architecture, the failure modes, and the real dollars, from someone who ships it.

    AI agentsClaudemulti-agentAI automationAI consulting

    Building a Multi-Agent OS with Claude: Architecture, Tools, and Workflows for Production AI Agents

    A multi-agent system with Claude as the orchestrator is worth building for exactly one kind of work: tasks where the information you need does not fit in one context window and the subtasks genuinely don't depend on each other. Anthropic's own engineering team measured a 90.2% performance improvement on internal research evaluations when they moved from a single agent to a lead agent (Claude Opus 4) coordinating specialized subagents (Claude Sonnet 4) in parallel. That number is real and I have built variations of this stack for clients. But the same team published the second number less loudly: the multi-agent setup burns roughly 15x the tokens of a normal chat interaction. If your task doesn't have parallel breadth, you are paying a 15x premium to move information around for no reason, and a single agent with a good skill file will beat your agent fleet. I've watched it happen.

    This article is the architecture I actually ship: hub-and-spoke orchestration with Claude as the coordinator, what goes in each layer, where the failure modes live (they are structural, not prompt bugs), and what it costs to run in production. Numbers below come from real client invoices, not vendor math.

    What is a multi-agent OS, actually?

    A multi-agent OS is a control layer that breaks a goal into subtasks, assigns each subtask to a specialized agent with its own context window, tools, and prompt, and merges the results. The "OS" framing is not a metaphor I love, but it's accurate in one way: like a kernel, the orchestrator doesn't do the work itself. It schedules, delegates, verifies, and cleans up.

    The shape that survives production is hub-and-spoke. One coordinator agent receives the task, decomposes it, spawns workers, and checks what comes back. Workers never talk to each other directly; everything routes through the coordinator. Claude Code's subagent system works exactly this way, and so does Anthropic's production research system behind Claude Research, described in their engineering post "How we built our multi-agent research system". The free-form alternative, a mesh of peer agents chatting with each other, is where the 2025-2026 failure studies concentrated their ugliest findings - the MAST authors logged inter-agent misalignment and duplicated, contradictory context as core failure categories, and those are mesh-native diseases. Open mesh reads as cool in a demo and falls apart at week two.

    Why does the orchestrator need to be the strongest model?

    Because decomposition quality caps the whole system's output. A 2025 UC Berkeley study (Cemri et al., the MAST taxonomy - Multi-Agent System Failure Taxonomy, arXiv 2503.17957) found that when a multi-agent system produces a report that misses entire categories of what you asked for, the root cause is almost never a lazy subagent. It's the coordinator's task decomposition. A worker can only answer the question it was handed; if the coordinator hands out three subtasks where the job needed five, no amount of worker competence recovers the missing two.

    That's why the standard assignment is the strongest model as lead and cheaper models as workers. Claude Opus 4 (or whatever your frontier Claude is at read time) plans and verifies; Claude Sonnet or Haiku does the fetching, extracting, drafting. The expensive tokens concentrate on the two moments that decide output quality: planning and verification. I ran a competitor-analysis workflow last spring for a client, and the lead's decomposition pass was 6% of total tokens but, in my judgment, 80% of the value. Spend frontier budget where it earns it.

    There's a second reason specific to Claude Code. Subagents cannot ask you a clarifying question mid-task, and (in versions before roughly v2.1.186, mid-2026; check your changelog, the behavior was being reworked when I last tested) a subagent running in the background auto-denies any tool call that would normally prompt for permission. So if the orchestrator's instructions are ambiguous, the worker fails silently and reports as if it succeeded. The strongest model belongs in the seat that writes the instructions.

    What are the core layers of a production stack?

    From watching my own deployments and fixing other people's, four layers matter:

    • Orchestration. The coordinator agent, its decomposition prompt, and the rule that all results return through it. In Claude Code this is the main conversation spawning subagents; with the Anthropic API it's an agent whose tool calls recursively invoke other agent sessions. Define input and output schemas for every worker. Structured JSON out of workers makes the verification pass reliable; free-form text makes it vibes.
    • Tools and integrations. Workers get scoped tool access: web search, file operations, an MCP (Model Context Protocol, the open standard for connecting LLM agents to external systems) connector to your CRM or database, a code interpreter. Scope matters. A research worker needs search and nothing else; giving it write access to your production systems because it was convenient to copy a config is how you end up with a very fast, very confident agent editing the wrong table.
    • Memory and state. The coordinator owns shared state; workers get only the context they need, passed as explicit scoped payloads rather than a growing shared transcript. This is the single biggest correctness lever. The MAST failure data shows most coordination failures are information-flow failures: the wrong context passed to the wrong agent, duplicated contradictory context, or lost results between steps. Less context, deliberately routed, beats more context sprayed everywhere.
    • Observability and gates. Every message logged, every worker run scored, human approval gates on anything that touches money, customers, or production data. Anthropic added a citation-checking pass separate from synthesis in their research system for exactly this reason: the agent that writes the answer should not be the only thing checking the answer. Langfuse, an open-source LLM observability platform, is what I use for tracing; anything that lets you replay a run is fine.

    Which workflows justify multi-agent in practice?

    Be honest here, because the 15x token multiplier has no sense of humor. In my client work, these earn their cost:

    • Deep research and competitive analysis: one question, many independent sources, synthesis at the end. The textbook case. This is the exact workload Anthropic's 90.2% number came from, and it's the workflow I trust least to a single agent for anything a client will actually read.
    • Document-heavy due diligence or compliance review: a folder of contracts, each summarized by a separate worker in parallel, discrepancies flagged by the coordinator.
    • Content pipelines with a verification stage: research, draft, then an independent critic subagent - sequenced after the draft worker, reporting to the coordinator like everything else, never talking peer-to-peer - that checks claims against sources before anything publishes. I run a version of this on my own blog. The critic rejects maybe a third of drafts and the published third is better for it.
    • Code investigation across a large repo: fan out readers to map different subsystems, then one coordinator assembles the architecture picture.

    What does not justify it: anything sequential with tight dependencies between steps, anything that fits in one context window, and anything where a human can verify the output in two minutes. For those, one good agent with a well-written skill file wins on cost, latency, and debuggability. I need to be honest about this: most "multi-agent" proposals I get from excited clients collapse into a single agent with extra steps once we map the actual dependencies. That's fine. Discovering you don't need the fleet is the cheapest outcome of the whole exercise.

    Where do multi-agent systems fail in production?

    Structurally, which is the uncomfortable finding. The Berkeley MAST study built a taxonomy of 14 failure modes, grouped under specification issues (bad system design), inter-agent misalignment, and missing task verification, and the headline lesson for operators is that you don't debug a multi-agent system by rewriting prompts. You debug it by fixing who talks to whom, what context they receive, and where verification happens.

    The three failures I hit most in the field:

    • Silent permission denials. Background workers auto-deny gated tool calls and report success (version caveats above apply - verify the behavior in your Claude Code release before you trust it). Symptom: a task "completed" but a file was never written. Fix: run permission-needing steps in the foreground or pre-grant explicitly scoped permissions.
    • Context contradiction. Two workers return conflicting facts and the coordinator averages them into mush. Fix: a verification step that compares worker outputs and escalates discrepancies instead of smoothing them.
    • Decomposition gaps. The report is thinner than the brief. Don't blame the workers; the coordinator never asked. Fix the decomposition prompt and add a completeness check against the original task before synthesis.

    The pattern across all three: every surviving multi-agent system I know has phase gates, shared artifacts, or a final supervisor. Nobody in 2026 is shipping an ungoverned peer mesh to production, and if they tell you they are, ask what happens when two agents disagree.

    What does it cost to run?

    Real numbers from a client research pipeline I ran in June 2026: roughly 40 deep-research reports a month, multi-agent, averaged about $3.80 per report in model spend versus $0.35 for the single-agent equivalent on the same tasks. It stings until you notice the multi-agent reports passed the client's accuracy review 9 times out of 10 in my (small, one-client, not your-mileage) sample and the single-agent versions needed a human rewrite more often than not. At consultant rates, $3.45 of extra tokens per report is noise. At ten thousand reports a day, it is very much not noise, and that's the actual decision boundary: multi-agent economics work when the value of the output is high and the volume is bounded. If you're at high volume and low unit value, you need caching, cheaper workers, or a different architecture, in that order.

    Token math, since Anthropic was kind enough to publish it: research-style orchestration runs about 15x chat token usage, and performance gains correlate strongly with token spend and parallel tool calls. Multi-agent is not a efficiency trick. It is a spend-more-tokens-get-more-coverage trade, and it should be sold internally as exactly that.

    Should you build one?

    If your workload matches the profile above, yes, and Claude plus either Claude Code subagents or the Anthropic Agent SDK is the fastest credible path. Start with one coordinator and two workers, log everything, gate the outputs, and add agents only when a real bottleneck appears. If your workload doesn't match the profile, the honest answer your consultant should give you is a single agent, a good skill file, and a cron job. I do this evaluation maybe once a month for someone, and I save them the fleet more often than I sell it.

    If you want a second pair of eyes on whether your task graph actually needs multiple agents, that's the kind of thing I do - reach out via ishchuk.eu.

    Frequently asked questions

    What is a multi-agent OS in AI?
    A multi-agent OS is a control layer that breaks a goal into subtasks, assigns each subtask to a specialized AI agent with its own context window, tools, and prompt, and merges the results through a central coordinator. The coordinator schedules, delegates, and verifies work but does not do the granular tasks itself, similar to how an operating system kernel manages processes. Production systems almost always use a hub-and-spoke design where worker agents never communicate directly with each other.
    How much better is a multi-agent system than a single AI agent?
    Anthropic reported a 90.2% performance improvement on internal research evaluations when using a multi-agent setup (Claude Opus 4 as lead, Claude Sonnet 4 subagents) compared to a single agent. The gain comes from parallel exploration across independent context windows. The trade-off is cost: the multi-agent system used roughly 15 times the tokens of a normal chat interaction, so the pattern only pays off for breadth-first tasks like deep research.
    Why do multi-agent AI systems use more tokens?
    Multi-agent systems use about 15x more tokens than a chat interaction because each subagent runs its own full context window, repeats task instructions, makes its own tool calls, and returns results that the coordinator then processes and synthesizes. Anthropic found that performance gains correlate strongly with token usage and parallel tool calls, meaning the extra spend is largely the mechanism behind the performance improvement, not overhead you can strip away.
    What is the MAST taxonomy for multi-agent AI failures?
    MAST (Multi-Agent System Failure Taxonomy) is a 2025 UC Berkeley study by Cemri et al., published as arXiv 2503.17957, that categorizes 14 failure modes of multi-agent AI systems under three groups: specification issues (flawed system design), inter-agent misalignment, and missing task verification. Its key finding is that multi-agent failures are structural, caused by how information flows between agents, rather than by prompt wording, so debugging means changing the architecture and context routing, not rewriting prompts.
    Should the orchestrator agent use a stronger model than its workers?
    Yes. The standard assignment is the strongest model as the lead orchestrator and cheaper, faster models as workers, because decomposition quality caps the whole system's output. If the coordinator's task breakdown misses subtasks, no worker competence can recover them, so the expensive tokens concentrate on planning and verification while cheaper models do the fetching and drafting.
    When should you not use a multi-agent AI architecture?
    Skip multi-agent for tasks with tight sequential dependencies, tasks that fit in a single context window, and tasks a human can verify in a couple of minutes. In those cases a single agent with a well-written skill file wins on cost, latency, and debuggability, because multi-agent orchestration burns roughly 15x the tokens and adds structural failure modes without adding coverage.