Office Hours — How should you test AI agents in production to ensure reliability and catch failures?
A daily developer question about AI/LLMs, answered with a direct, opinionated take.
How should you test AI agents in production to ensure reliability and catch failures?
Testing agents in production is different from testing traditional software because the failure modes are subtler and the stakes are higher. An agent can look like it worked, generate plausible-sounding logs, and still have made a critical mistake. You need signal before you have visibility.
Start with Outcome Monitoring, Not Log Inspection
Don’t trust what the agent claims it did. Watch what actually happened in the systems it touched. If your agent submits a pull request, check whether the tests pass. If it executes a database query, validate the result set. If it calls an API, verify the response matches expectations.
This sounds obvious, but most teams still rely on parsing LLM-generated explanations of what the agent did. That’s insufficient. An agent might confidently report “I successfully updated the user record” when it actually updated the wrong user or partially failed mid-transaction.
Set up outcome validators for each critical action the agent takes. These are simple assertions: “Did the PR get created?” “Does the file exist?” “Can we query the modified data back?” Keep them lightweight because they’ll run frequently.
Implement Token and Cost Budgets as Safety Valves
Runaway agents burn money faster than they cause damage. Set hard limits on:
- Total tokens per task (not just per request)
- Total API spend per day
- Consecutive failed retries before human escalation
When an agent hits a budget ceiling, it should fail loudly and stop. Don’t let it keep trying.
Here’s a simple budget pattern:
class AgentBudget:
def __init__(self, max_tokens=100_000, max_cost=5.00):
self.tokens_used = 0
self.cost_used = 0.0
self.max_tokens = max_tokens
self.max_cost = max_cost
def deduct(self, tokens, cost):
self.tokens_used += tokens
self.cost_used += cost
if self.tokens_used > self.max_tokens or self.cost_used > self.max_cost:
raise BudgetExceededError(f"Tokens: {self.tokens_used}/{self.max_tokens}, Cost: ${self.cost_used:.2f}/${self.max_cost:.2f}")
This isn’t sophisticated, but it prevents the agent from burning through your month’s budget in one bad loop.
Test on Real Tasks With Synthetic Baselines
Run your agent against the exact tasks it’ll face in production, but compare its outcomes to a known-good baseline. If the agent is supposed to categorize support tickets, run it against a test set where you already know the correct category. If it’s meant to write code, compare the generated code against passing test suites.
The baseline doesn’t have to be perfect. It just needs to be better than random. You’re looking for: does the agent do better than a deterministic fallback?
Databricks benchmarked its coding agents by running them on tasks from their production codebase and measuring whether the generated code passed tests. They found that GLM-5.2 matched Claude Opus 4.8 on their specific workload at lower cost. They didn’t trust vendor benchmarks. They built their own.
Do the same. Don’t run SWE-Bench or other public benchmarks. Run your agent on your actual tasks. The gap between benchmark performance and real-world performance is usually massive.
Watch for Silent Failures and Hallucinated Tool Calls
The worst agent failure is the one it doesn’t report. An agent might hallucinate a tool call (“I successfully called the API to fetch user data”), but the tool actually failed or never ran. The agent then bases downstream decisions on information that doesn’t exist.
Instrument every tool call the agent makes:
def call_tool(name, **kwargs):
try:
result = tools[name](**kwargs)
log_tool_call(name=name, args=kwargs, result=result, status="success")
return result
except Exception as e:
log_tool_call(name=name, args=kwargs, error=str(e), status="failed")
# Don't let the agent keep going with a hallucinated result
raise
Require the agent to acknowledge tool failures. If a tool call fails, the agent should note it explicitly and adjust its plan. If it pretends the tool succeeded when it didn’t, escalate to a human.
Implement Circuit Breakers for Suspicious Patterns
Watch for agent behaviors that smell wrong, even if the outcome looks correct:
- Repeated retries on the same task without variation
- Tool calls in illogical order (calling delete before reading)
- Requests for excessive permissions or data access
- Confidence scores that don’t match outcome volatility
When you see these patterns, fail safely. Log a detailed trace and route to human review rather than continuing.
Claude Code with Auto Mode catches 89% of dangerous commands versus 13.6% for humans, but it’s still catching and rejecting, not blindly executing. Copy that pattern: classify risky actions and require human approval before proceeding.
Test Failure Recovery, Not Just Happy Paths
Most teams test agents assuming everything works. Test what happens when tools are slow, timeouts fire, APIs return errors, or databases are locked. Does the agent handle graceful degradation? Can it retry intelligently without looping?
Run failure injection tests. Simulate tool outages. Force latency spikes. Return malformed responses. Real production will do all of these anyway.
Trace Every Decision, Not Just Final Outputs
You need to know why the agent did what it did, especially when the outcome is wrong. Log:
- Every tool call and its result
- Model reasoning at each step
- Branching decisions (if X, then Y; else Z)
- Context used to make decisions
Tools like Weights & Biases Weave let you trace thousands of agent executions and spot patterns. A practitioner used Weave to identify redundant model calls in an apartment-search agent across 2,500 traces and systematically removed waste. You can apply the same approach to catch systematic failures.
Set Clear Success Criteria Before Deployment
Define what “working” means for your agent before it goes live. Is it:
- Completing the task without human intervention?
- Completing it faster than a human would?
- Completing it with zero errors, or just fewer errors than the baseline?
If the success criterion is vague (“the agent seems to work well”), you won’t know when it’s failing. Make it concrete and measurable.
Expect Long Debugging Cycles
Finding the root cause of agent failures is slow. The agent might fail due to a bad prompt, a missing tool, incorrect context, an edge case in the data, or a model limitation. Narrow down systematically. Disable tools one at a time. Change the prompt in isolation. Test on simpler versions of the task first.
Reserve more time for debugging agents than for tuning traditional systems. The feedback loop is longer and the failure modes are less obvious.
Bottom line: Test agents by watching what they actually changed in real systems, not what they claim they did. Set hard budgets, validate outcomes against baselines, instrument tool calls, and fail safely when patterns look suspicious. You’ll catch failures faster than waiting for the perfect safety classifier.
Question via Hacker News