The Benchmark — MT-Bench
A plain-English explainer of one AI evaluation benchmark: what it measures, how it works, and when to trust it.
MT-Bench
Evaluates conversational AI quality through multi-turn dialogues judged by GPT-4, measuring a model’s ability to follow instructions, maintain context, and produce coherent responses across extended exchanges.
What it measures
MT-Bench tests how well language models handle realistic, multi-turn conversations where context and instruction-following matter. Instead of single isolated questions, it presents 8 dialogue scenarios spanning writing, roleplay, reasoning, math, coding, and creative tasks. Each model generates responses to follow-up prompts in sequence, simulating how humans actually interact with chatbots.
Why it was created
Released by LMSYS (UC Berkeley) in March 2023, MT-Bench was designed to address a key limitation of single-turn benchmarks: they don’t capture how models perform in sustained conversations where earlier responses influence later ones. As chat interfaces became the primary way users engaged with LLMs, existing benchmarks felt detached from real-world use. The authors wanted a lightweight alternative to human evaluation that could quickly rank models.
How it works
MT-Bench contains 80 questions across 8 categories (writing, roleplay, extraction, reasoning, math, coding, creativity, and knowledge). For each question, a model receives an initial prompt, generates a response, then receives a follow-up question requiring it to build on that response. A second turn follows similarly. Crucially, responses aren’t scored by rubric — instead, GPT-4 (using Claude in recent versions) compares two model responses side-by-side and declares a winner, or calls it a tie. Scores range from 1–10 based on aggregate win rates. The benchmark runs quickly (minutes per model) and costs roughly $100–150 per full evaluation.
What scores mean in practice
GPT-4 scores around 8.6–9.0 on MT-Bench, Claude 3 Opus reaches ~9.0, and Claude 3.5 Sonnet scores ~9.2. A year ago (late 2023), GPT-4 at 8.6 represented the frontier; now models regularly exceed that. Open models like Llama 2 70B scored ~6.6 initially, while Llama 3 70B reached ~8.0. A score of 7.0+ is generally considered “quite capable” at multi-turn tasks; 6.0–7.0 suggests competence with limitations; below 6.0 indicates notable struggles with context retention or complex instructions.
Known limitations
- Judging bias: Relying on GPT-4/Claude as a judge introduces circular dependency — benchmarking LLMs with LLMs, which may favor certain model architectures or writing styles. The judge itself has preferences and blindspots.
- Limited sample size and narrow scope: Only 80 questions across 8 categories means high variance and vulnerability to contamination (models may have seen similar prompts during training). Math and coding tasks especially feel underrepresented for evaluating modern capabilities.
- Lacks measurement of harmful behavior: MT-Bench doesn’t explicitly test refusals, robustness to adversarial inputs, or long-context performance — critical for production systems. A high MT-Bench score doesn’t guarantee safety or reliability.
When to trust it (and when not to)
- Trust it for: Quick comparative ranking of conversational models. If you’re choosing between two chat interfaces and want a signal faster than human evaluation, MT-Bench scores correlate reasonably with user preference in real conversations. It’s useful for detecting major regressions during model development.
- Don’t trust it alone for: Mission-critical decisions on which model to deploy. Use it alongside human evals, domain-specific benchmarks (coding, math, knowledge), and safety evaluations. A high MT-Bench score tells you a model is conversationally fluent — not that it’s accurate, safe, or suitable for your use case.