The Benchmark — MMMU (Massive Multi-discipline Multimodal Understanding)
A plain-English explainer of one AI evaluation benchmark: what it measures, how it works, and when to trust it.
MMMU (Massive Multi-discipline Multimodal Understanding)
A benchmark that tests whether AI models can understand and reason about images across 30 academic disciplines, combining visual and textual comprehension.
What it measures
MMMU evaluates how well models handle questions that require both reading text and interpreting images—charts, diagrams, photographs, screenshots, equations rendered as images—to arrive at correct answers. The benchmark covers subjects from medicine and engineering to art history and law, testing whether models can extract information from visual content rather than just process text tokens.
Why it was created
Most multimodal benchmarks at the time (2023) focused on simple image captioning or visual object recognition. MMMU was created by researchers at CMU, Salesforce, and other institutions to fill a gap: existing vision-language benchmarks didn’t test the kind of rigorous, domain-specific reasoning that actually matters in professional and academic contexts. The benchmark was released in November 2023 as multimodal AI became a competitive frontier.
How it works
MMMU contains 12,000 multiple-choice questions across 30 disciplines (medicine, chemistry, physics, art history, etc.). Each question is paired with one or more images—these can be photographs, plots, architectural diagrams, microscopy images, or other visual content. Models receive both the image and question text, and must select the correct answer from typically 4 options. Scoring is straightforward: accuracy percentage. The dataset is split into validation (500 questions) and test (11,500 questions) sets, with answers held out from public access.
What scores mean in practice
As of late 2024, Claude 3.5 Sonnet and GPT-4V score around 59-61% on MMMU, while older models like Claude 3 Opus scored ~50-52%. For reference, human expert performance varies by domain but averages around 78-82% depending on expertise level in each field. Two years ago (2022), no multimodal models were being evaluated on this benchmark; the first published results showed even the best models at ~30-40%, making the current 60% a meaningful improvement but still a large gap from expert performance.
Known limitations
-
Domain imbalance and expert annotation concerns: Some fields are underrepresented (only ~400 questions on law or philosophy vs. ~600 on medicine). Answers were validated by experts but inter-annotator agreement metrics aren’t always published, creating ambiguity around edge cases.
-
Contamination risk in training data: Popular images from textbooks, academic papers, and Wikipedia likely appear in the training data of frontier models. While MMMU tries to use less-common images, models trained on the full internet have probable exposure to similar visual content and associated captions.
-
Still primarily visual QA, not deep reasoning: While harder than typical benchmarks, MMMU is fundamentally multiple-choice selection. Models can succeed through pattern matching rather than the kind of step-by-step visual reasoning required in real professional work (e.g., a radiologist interpreting an X-ray and writing a detailed report).
When to trust it (and when not to)
-
Trust it for: Comparing multimodal models on domain-aware visual understanding, especially when you care about whether a model can handle academic or professional visual content. It’s genuinely harder than image captioning and correlates reasonably with real-world multimodal capability.
-
Don’t trust it alone for: Concluding a model is “production-ready” for specialized work (medical, legal, engineering). MMMU doesn’t measure the kind of accountability, error detection, or open-ended reasoning that matters in high-stakes domains. Use it as one signal among benchmarks testing longer-form output and real-world task performance.