Office Hours — How do you handle hitting AI coding agent usage limits in production?
A daily developer question about AI/LLMs, answered with a direct, opinionated take.
How do you handle hitting AI coding agent usage limits in production?
You’re running agents at scale, they’re hitting rate limits or token budgets, and now you’ve got a choice: queue the work, fail gracefully, route to a cheaper model, or actually think about whether agents are the right tool here. Most teams don’t plan for this until it breaks.
The Real Problem Isn’t the Limit, It’s Visibility
Rate limits and usage caps exist on every frontier API. OpenAI, Anthropic, Google all enforce them. The problem is that coding agents can burn through your monthly allocation in hours if they’re looping on a hard problem or retrying failed tasks. A single agent debugging a complex codebase can generate 10–50x more tokens than a human writing the equivalent code.
The first thing you need is honest telemetry. Track tokens per task, cost per task, and failure rates separately. If you’re not measuring this, you’re flying blind. Asana cut costs 76x on browser automation by switching to GPT-6 Astra, but only because they measured it. Most teams don’t.
# Basic cost tracking (pseudo-code)
import time
def track_agent_task(model, task_name):
start_tokens = get_token_count()
start_cost = 0
result = run_agent(model, task_name)
end_tokens = get_token_count()
task_cost = (end_tokens - start_tokens) * get_model_price(model)
log_metric({
"task": task_name,
"model": model,
"tokens_used": end_tokens - start_tokens,
"cost": task_cost,
"success": result.success,
"retries": result.retry_count
})
return result
Practical Strategies When You Hit the Wall
Tiered fallback routing. Start with your preferred model (Claude Opus 5, GPT-6 Astra), fall back to a cheaper tier (Claude Sonnet 5, GPT-5.6 Luna) when you hit rate limits. The tradeoff is latency and quality, but for many agentic tasks it’s acceptable. Databricks benchmarked their codebase and found GLM-5.2 matched Claude Opus 4.8 performance at lower cost, so cheaper models can actually be the right default for your specific workload.
Asynchronous queueing with backpressure. Instead of blocking on every agent task, queue work to a job system (Celery, Bull, Temporal). Set per-queue rate limits and process tasks during off-peak windows. This trades latency for cost: a code review agent that finishes in 2 hours instead of 30 seconds is cheaper per token and often acceptable for async workflows.
Model-specific budget caps. Set hard limits on spend per day or per task. When you hit the cap, degrade gracefully: fail the task with a clear error, suggest human review, or route to a synchronous alternative. Don’t silently queue infinite work hoping budgets magically expand.
# Pseudo-code: hard budget enforcement
MAX_DAILY_SPEND = 1000 # USD
TOKENS_PER_DOLLAR = 2000000 # depends on model
class CostGate:
def __init__(self):
self.daily_spend = 0
self.last_reset = time.time()
def can_proceed(self, estimated_tokens, model_price):
if time.time() - self.last_reset > 86400:
self.daily_spend = 0
self.last_reset = time.time()
task_cost = (estimated_tokens / 1e6) * model_price
if self.daily_spend + task_cost > MAX_DAILY_SPEND:
return False, f"Daily budget exceeded: {self.daily_spend:.2f}/${MAX_DAILY_SPEND}"
self.daily_spend += task_cost
return True, None
Aggressive caching and deduplication. Coding agents often repeat the same queries. Prompt caching (GPT-6 Astra, Claude Fable 5.1) cuts re-input costs dramatically. Long context cached at $2.50/M (Fable 5.1) versus $10/M for uncached input is a 4x saving on reads. Track which tasks agents run repeatedly and cache aggressively.
The Harder Question: Is This Task Worth Automating?
Before you optimize around limits, ask whether the task should be automated at all. Autonomous agents work reliably when success is verifiable (tests pass/fail, linter output, CI checks). They drift badly on subjective tasks (code review quality, architectural judgment). If your agent is looping because it can’t reliably measure success, no budget will fix it.
The Remote Labor Index (Center for AI Safety) shows agents complete only about 16% of real freelance jobs at professional quality. That means 84% of tasks still need human judgment. If your agent is burning budget failing on ambiguous tasks, the solution isn’t better queuing, it’s narrower scope.
Long-running coding agents (35-hour autonomous chip optimization by Qwen3.7-Max) are real but rare. Most production agents work best on narrow, well-defined tasks with fast feedback loops. Be honest about which category your task falls into.
Cost Optimization You Actually Control
Prompt construction matters more than most teams realize. A “context compiler” approach (deciding what to keep and discard from the context window) can unlock agent performance without longer context windows or newer models. If you’re hitting limits because agents are loading entire repositories into every call, that’s a prompt engineering problem, not a budget problem.
Model routing as a design choice (tools like Jev from TypeSafe now exist for this) lets you use cheaper models for planning and structure-enforcement, then hit expensive models only for complex reasoning. Cursor’s agent architecture separates planning (frontier models) from execution (cheaper models) and rebuilt SQLite in Rust with 100% test coverage at reasonable cost. That pattern works.
Bottom line: Track tokens per task obsessively, set hard budget caps with graceful degradation, and route to cheaper models when they’re sufficient for your workload. But first, narrow your agent’s scope to tasks with verifiable success criteria—throwing budget at unsolvable problems wastes both money and time.
Question via Hacker News