Office Hours — How do you structure multi-agent workflows to scale beyond a handful of agents in production?
A daily developer question about AI/LLMs, answered with a direct, opinionated take.
How do you structure multi-agent workflows to scale beyond a handful of agents in production?
The Real Scaling Problem
Most teams don’t fail because one agent breaks. They fail because five agents working together produce exponential token bloat, silent coordination failures, and costs that become unmanageable before anyone notices. A single agent calling Claude Opus 5 costs money. Ten agents each maintaining their own context while messaging each other? That’s where you find yourself with a $50K monthly bill for something that should cost $5K.
The scaling problem isn’t orchestration. It’s token economics and failure visibility.
Token Multiplication Is Stealth
When you have one agent solving a task, you pay once. When you have five agents working in parallel, each maintaining its own context about the task, the same information gets encoded five times. Worse, when agents message each other, they often duplicate context: Agent A sends 12K tokens to Agent B, Agent B processes them and sends 8K tokens to Agent C, who sends 10K to Agent D. You’ve paid for 40K tokens to move a single piece of information through the system.
Cursor’s architecture separation—frontier models for planning, cheaper models for execution—works because it front-loads the expensive reasoning and delegates repetitive work. But this only works if you build it intentionally. Default multi-agent systems don’t do this.
Start by modeling token cost as part of your agent design, not after. If you’re using Claude Opus 5.5 (matching Fable performance at 40% lower cost) for every agent, you’re overpaying on the execution layer. Reserve Opus for tasks that actually need that reasoning capability: planning, arbitration, complex trade-off decisions. Use Claude Sonnet 5 ($2/$10 per M, permanently) for most agent work. Use cheaper models or deterministic code for pure execution.
Structured Handoff Reduces Silent Failures
Agents fail silently when they don’t know what went wrong. If Agent A sends a message to Agent B and B doesn’t understand it, B might hallucinate a response instead of asking for clarification. Now you’ve corrupted the workflow and the error propagates downstream.
Use explicit schemas for inter-agent communication. Don’t rely on free-form text. Define what Agent B expects to receive, what success looks like, and what counts as a retry.
Here’s a concrete pattern:
from dataclasses import dataclass
from enum import Enum
class AgentRole(Enum):
PLANNER = "planner"
EXECUTOR = "executor"
VALIDATOR = "validator"
@dataclass
class TaskHandoff:
task_id: str
source_agent: AgentRole
target_agent: AgentRole
payload: dict
required_fields: list[str]
success_criteria: str
retry_budget: int
def validate(self) -> bool:
return all(f in self.payload for f in self.required_fields)
# In your orchestration loop:
handoff = TaskHandoff(
task_id="build-api-integration",
source_agent=AgentRole.PLANNER,
target_agent=AgentRole.EXECUTOR,
payload={
"api_endpoint": "https://api.example.com",
"auth_method": "oauth2",
"required_scopes": ["read:data", "write:data"],
},
required_fields=["api_endpoint", "auth_method"],
success_criteria="successful_auth_token_obtained",
retry_budget=3,
)
if not handoff.validate():
raise ValueError(f"Handoff missing required fields")
# Now the executor knows exactly what it's receiving,
# and failure modes are explicit.
This removes the ambiguity that causes hallucination. Agent B knows what fields it’s receiving, what it’s supposed to do, and when to fail loudly instead of improvising.
Observability Needs to Scale Before Agents Do
With three agents, you can read logs. With fifty agents, you can’t. Build observability into your architecture from the start.
Log every agent invocation with structured metadata: which agent, what model, input token count, output token count, latency, success/failure, and the decision it made. Use time-series metrics to track:
- Total tokens per agent per day
- Cost per agent per day
- Failure rate (tasks that required retry)
- Handoff failures (messages that didn’t parse or were rejected)
- Average latency per agent
When an agent starts drifting—producing worse results or consuming 3x tokens for the same task—you’ll see it in the metrics before it becomes a production incident. Without this, you’re flying blind.
TypeSafe’s “System One” decision models (structured outputs that return probabilities instead of generated text) are gaining traction because they’re also easier to monitor. You get clear pass/fail signals instead of trying to judge whether free-form text “looks good.”
Implement a Resource Budget, Not Just a Token Limit
A global token limit (“stop if we hit 1M tokens today”) is useless because it doesn’t tell you which agent to cut off or when. A resource budget assigns each agent a quota and tracks it in real time.
class AgentBudget:
def __init__(self, agent_id: str, daily_token_limit: int, daily_cost_limit: float):
self.agent_id = agent_id
self.daily_token_limit = daily_token_limit
self.daily_cost_limit = daily_cost_limit
self.tokens_used = 0
self.cost_incurred = 0.0
def can_invoke(self, estimated_tokens: int) -> bool:
if self.tokens_used + estimated_tokens > self.daily_token_limit:
return False
if self.cost_incurred + self.estimate_cost(estimated_tokens) > self.daily_cost_limit:
return False
return True
def record_invocation(self, input_tokens: int, output_tokens: int, cost: float):
self.tokens_used += input_tokens + output_tokens
self.cost_incurred += cost
Each agent gets a daily allowance. When an agent hits its limit, it degrades gracefully: fall back to cached results, or delegate to a cheaper model. The system doesn’t crash; it just becomes cheaper.
Use Broadcast for Shared State, Not P2P Messaging
Ten agents each sending messages to ten other agents creates O(n²) communication complexity. Broadcast to a shared state layer instead. Use a lightweight message bus or a database as your single source of truth.
When Agent A completes a task, it publishes the result to a topic. All agents subscribed to that topic see it simultaneously. No duplication, no lost messages, and the state is consistent across the system.
This also makes recovery easier: if an agent crashes, it can replay the message log to get back to the correct state instead of having to ask every other agent “what did I miss?”
Validation Must Happen Before Handoff
Before Agent B trusts Agent A’s output, Agent B should verify it against the success criteria. This prevents hallucinated work from propagating.
def validate_task_completion(result: dict, criteria: str, validator_model: str) -> bool:
# Use a cheap, fast model for validation
# Opus is overkill here
prompt = f"""
Task completion criteria: {criteria}
Result: {result}
Did the result meet the criteria? Answer only YES or NO.
"""
response = invoke_model(validator_model, prompt, temperature=0)
return "YES" in response
Keep the validator lightweight. Claude Sonnet 5 is fine. The goal is to catch obviously wrong results before they waste time downstream.
Example: Document Processing Pipeline at Scale
Say you have a pipeline that ingests PDFs, extracts data, validates it, and stores it. With five agents:
- Splitter (breaks PDFs into chunks)
- Extractor (pulls structured data from chunks)
- Reconciler (merges overlapping extractions)
- Validator (checks data quality)
- Storer (writes to database)
Each agent gets a Sonnet 5 for execution logic, with Opus 5.5 only for Reconciler (which needs judgment about conflicting extractions). Cost drops from $0.50 per document (all Opus) to $0.12 per document (mostly Sonnet).
Each agent publishes to a topic: `pdf
Question via Hacker News