Office Hours — How do you evaluate whether your LLM application is actually working better than baseline, and what evaluation frameworks hold up in practice? A daily developer question about AI/LLMs, answered with a direct, opinionated take. 2026-09-14T12:00:00.000Z Office Hours Office Hours office-hoursq-and-apractical-ai

Office Hours — How do you evaluate whether your LLM application is actually working better than baseline, and what evaluation frameworks hold up in practice?

A daily developer question about AI/LLMs, answered with a direct, opinionated take.

Daily One question from the trenches, one opinionated answer.

How do you evaluate whether your LLM application is actually working better than baseline, and what evaluation frameworks hold up in practice?

This question hits at the gap between “the model produces plausible output” and “the system is materially better than what we had before.” Most teams skip this part and ship anyway. Then six months in, you can’t tell if GPT-6 Astra is actually solving the problem or just confidently wrong in new ways.

The baseline problem

The first trap is not having a baseline at all. You need to define what you’re comparing against: the human doing it manually, a simpler heuristic, the previous model version, or just random guessing. Without it, you’re measuring capability in a vacuum, which is useless for business decisions. Anthropic’s work benchmarking Claude on their own production codebase revealed that a closed-weight open-source model (GLM-5.2) matched Claude Opus 4.8 on their specific tasks while costing 34% less per job. They only discovered this because they built evals on their actual workload, not some generic benchmark. Your domain is not the same as MMLU or SWE-Bench Pro.

What’s worse: SWE-Bench Pro itself has ~30% broken tasks, and OpenAI withdrew its endorsement after discovering this. Generic benchmarks are gamed. Your benchmarks won’t be.

Three evaluation buckets that work

Automated metrics on historical data. If you have ground truth (customer satisfaction ratings, conversion events, bug reports resolved, etc.), you can replay your LLM output against that dataset and measure precision, recall, F1, accuracy—whatever’s appropriate. This is cheap and repeatable. The catch: it only works if success is binary or easily measurable. If your task is “improve code readability” or “make the UI less confusing,” you’ll struggle. Also, your historical distribution might not match the live distribution, especially if the LLM is already changing user behavior.

Human annotation on a held-out sample. Pick 100–200 random outputs from your LLM and from the baseline. Have humans (ideally not you, and definitely not the person who built the prompt) score them on whatever dimension matters: correctness, clarity, safety, speed, user satisfaction. Use a clear rubric. This is more expensive than automated metrics but gives you a realistic picture of real-world quality. The downside is it doesn’t scale for continuous improvement—you can’t annotate every output. Use it sparingly to validate your automated metrics.

Production A/B testing when possible. Split traffic. Half gets the LLM, half gets the baseline. Measure actual business outcomes: engagement, task completion time, error rate, cost-per-transaction. This is the gold standard because it tells you what users actually experience. But you need volume and a fast feedback loop. If your feature touch only 10 users a month, you’ll need months to reach statistical significance. Also, if the LLM is genuinely worse, you’re shipping bad experiences to real people—so run it on a smaller cohort first.

What breaks in practice

User experience metrics lie. Your LLM application might reduce perceived latency because streaming makes it feel faster, but actually increase total inference cost. Databricks’ benchmarking discovered that “proxy wins” (picking cheaper models) only work if you measure on your codebase, not generic benchmarks. They found GLM-5.2 matched Opus 4.8 on their million-line production system while others might see the opposite.

Selection bias corrupts observational comparisons. If you only A/B test your LLM on users who opt into a beta, those users are systematically different from everyone else—they’re more tech-forward, more forgiving of bugs, more likely to give feedback. Your improvement signal is confounded with user type. This is why Towards Data Science published “Your AI Adoption Lift Is a Selection Effect”—teams were measuring genuine feature value when they were actually measuring who chose to try the feature.

Benchmark inflation is real. The “Benchmark Paradox” showed GPT-6 Astra scoring anywhere from 39% to 100% depending on which evaluation you ran. Same model, wildly different stories. This matters because you might pick the wrong model based on which benchmark you trust.

Token budget masquerades as capability. The UK AI Security Institute found that agentic system success on tasks like TerminalBench 2.0 and SWE-Bench Pro rises ~25% when test-time compute grows 10x (1M→10M tokens). What looked like an inferior model might just need more reasoning tokens. If you don’t measure inference cost and time together with quality, you’ll make wrong decisions.

A concrete framework

Here’s what works:

  1. Define success clearly before you start. Is it speed? Cost? Accuracy? User satisfaction? Pick one primary metric and 1–2 supporting metrics. If you can’t measure it, don’t ship it.

  2. Build a small eval set on your domain. 50–100 examples max. Annotate ground truth yourself or with a co-founder. This becomes your canary—if the LLM fails here, it will fail in production.

  3. Run automated metrics on that eval set continuously. Compare your LLM output to baseline output using your chosen metrics. This is your dashboard. Make it a blocker for pushes if the score drops.

  4. Pick one high-risk production cohort. Route 5–10% of traffic to your LLM initially. Measure actual business outcomes: task completion, time-on-task, error rate, cost-per-result. Watch for silent failures (the model produced output but it was wrong in a way your automated metrics didn’t catch).

  5. After two weeks of stability, expand to 50%, then 100%. If you hit a regression, roll back immediately and debug.

  6. Track infrastructure costs alongside quality. An LLM that’s 5% more accurate but costs 3x as much is not an upgrade. Use Cursor’s architecture pattern: route cheap models for execution once a frontier model has planned the work. Costs drop while quality stays high.

Example from production: a team using Claude Sonnet 5 for customer support classification was getting 87% accuracy on their eval set but noticed their live accuracy was 73%. The gap? Sonnet’s tokenizer now emits ~30% more tokens for the same text, so each inference was more expensive but not proportionally more accurate. They switched to Gemini 3.5 Flash for this task (simpler problem, lower compute needed) and recovered both accuracy and cost efficiency.

Frameworks that hold up

Synthetic evals on your domain are better than generic benchmarks. Build three versions of your eval: happy path, edge cases, and adversarial examples (things designed to break the LLM). If it passes all three, it probably won’t crater in production.

Cost-normalized metrics matter. Don’t just measure accuracy. Measure accuracy-per-dollar and accuracy-per-millisecond. A 10% cheaper model with 2% lower accuracy is often the right call.

Continuous monitoring beats one-time evals. Your baseline drifts. User behavior changes. New edge cases emerge. Sample 20–50 outputs per week from production, annotate them, and watch your metrics drift. This is where you catch the “one capital letter breaks the bot” class of failures.

Comparison requires the same compute budget. If you give GPT-6 Astra 10M tokens to reason and give Claude Opus 5 1M tokens, GPT-6 wins by default. Compare fairly: same input, same time limit, same cost. Then measure.

Bottom line: Build a small, domain-specific eval set on your actual data before shipping anything. Automate metrics on it. A/B test on a small production cohort. Measure business outcomes, not just benchmark scores. Cost per unit of quality is your real north star, and most teams never calculate it.

Question via Hacker News