LLM benchmarks are standardized test sets, such as MMLU, GPQA Diamond or SWE-bench, that score language models on the same questions so their abilities can be compared. A benchmark score tells you how a model did on one specific task format under one specific setup; it does not tell you how the model will perform on your data. In 2026 the most useful benchmarks are hard, recently updated and resistant to memorization, while older ones like MMLU are saturated and mostly of historical interest.
This guide explains what the common benchmarks actually measure, why scores stop meaning much over time, and how to read a model announcement without being misled. It deliberately does not rank models; for how ranking sites work, see our guide to LLM leaderboards.
What LLM benchmarks measure
A benchmark has three parts: a fixed set of tasks, a way of prompting the model, and a scoring rule. Change any of the three and the number changes. Most benchmarks fall into a handful of families:
- Knowledge and reasoning: multiple-choice or short-answer exams across subjects.
- Math: problems with a single checkable final answer.
- Coding: programs that must pass unit tests, from single functions up to real repository issues.
- Agentic tasks: multi-step work in a terminal, browser or tool environment, scored by whether the end state is correct.
- Tool use: whether the model calls functions with the right names and arguments.
- Human preference: people compare two answers and pick the better one, which is how arena-style leaderboards work.
The first five are scored automatically against a reference. Human preference is not a benchmark in the strict sense, because there is no fixed answer key.
The benchmarks you will see in 2026
| Benchmark | Measures | Size | Status in 2026 |
|---|---|---|---|
| MMLU | Knowledge across 57 subjects, multiple choice | About 14,000 test questions | Saturated |
| MMLU-Pro | Harder MMLU with 10 answer options | About 12,000 questions | Still used, nearing saturation |
| GPQA Diamond | Graduate-level science questions written to resist web search | 198 questions | Widely reported, small sample |
| Humanity's Last Exam | Expert-written questions at the frontier of knowledge | 2,500 questions | Hard, actively used |
| SWE-bench Verified | Fixing real GitHub issues in Python repositories | 500 tasks | Contamination and test problems documented |
| Terminal-Bench | Agentic tasks in a command-line environment | Version 4.0 | Actively maintained |
| ARC-AGI-3 | Learning novel interactive game environments | Game set | Launched March 2026 |
| BFCL | Function and tool calling accuracy | Version 4 | Actively maintained |
| LiveBench | Mixed tasks with monthly fresh questions | Rotating | Built to limit contamination |
Put simply: the older multiple-choice exams are saturated or close to it, the newer expert exams and agentic benchmarks are where differences still show up, and benchmarks that refresh their questions are the best defense against memorization.
MMLU and MMLU-Pro
What is MMLU? Massive Multitask Language Understanding is a multiple-choice exam covering 57 subjects, from elementary math to law and medicine, introduced in 2020 (paper). For years it was the headline number in model launches. Frontier models now score so close to the ceiling that the remaining gap is partly label noise, and several trackers, including the Vellum leaderboard, exclude it as saturated. MMLU-Pro raised the difficulty with about 12,000 questions and ten answer options instead of four, which also reduces lucky guessing (paper).
GPQA Diamond
GPQA is a set of 448 graduate-level biology, physics and chemistry questions written by domain experts and designed so that skilled non-experts with web access still struggle (paper). The Diamond subset is the 198 highest-quality questions. Because it is small, a difference of a couple of points between two models can be within the noise of a single run, so look for averaged runs or confidence intervals.
Humanity's Last Exam
Humanity's Last Exam, from the Center for AI Safety and Scale AI, contains 2,500 expert-written questions across many fields, including some with images, and was published in Nature (article). It was built because the older exams had stopped separating top models. Some aggregators use only the text questions, so check which version a reported score refers to.
SWE-bench Verified and its successors
SWE-bench asks a model to resolve real GitHub issues; a fix counts only if the repository's tests pass. The Verified subset of 500 tasks was human-screened to remove broken ones and became the standard coding number. In February 2026 OpenAI said it would stop reporting it: an audit of tasks models frequently failed found that at least 59.4 percent of them had flawed tests or descriptions, and frontier models could reproduce the original fixes from memory (OpenAI). OpenAI later reported major problems in SWE-bench Pro too. The lesson is general: even curated benchmarks contain errors, and popular public tasks leak into training data.
Agentic and tool-use benchmarks
Terminal-Bench, now at version 4.0, measures whether an agent can complete realistic tasks in a terminal, and reports confidence intervals alongside resolution rates. ARC-AGI-3, the newest member of the ARC-AGI series, launched by the ARC Prize Foundation in March 2026, drops static puzzles in favor of interactive game environments that agents must learn without instructions; a 100 percent score means beating every game as efficiently as humans (ARC Prize). The Berkeley Function Calling Leaderboard measures whether models call tools with correct names and arguments.
Older classics: HumanEval, GSM8K and BIG-bench
HumanEval, 164 hand-written Python problems from 2021, and GSM8K, about 8,500 grade-school math word problems, still appear in papers (HumanEval, GSM8K). Both are saturated for frontier models and widely present in training data, so treat high scores as a minimum bar rather than a differentiator. BIG-bench, a collaborative suite of more than 200 tasks contributed by 450 authors at 132 institutions, is another research benchmark you will see cited, especially in papers about how capabilities change with scale (paper).
Why benchmark scores stop meaning much
Saturation. When the best models all score near the ceiling, the benchmark can no longer tell them apart. Researchers respond by building harder sets, which is why MMLU gave way to MMLU-Pro, GPQA and Humanity's Last Exam.
Benchmark contamination. Public test questions end up in web crawls and then in training data, so a model can score well by recall rather than reasoning. Benchmark creators fight back with hidden test sets, canary strings that ask crawlers to exclude the data, and fresh questions. LiveBench, for example, releases new questions regularly and scores them against objective answers instead of an LLM judge (paper).
Setup differences. The same model can score differently depending on the number of examples in the prompt, chain-of-thought settings, answer extraction, tool access and how many attempts are allowed. Vendor numbers often use the most favorable setup. Independent harnesses such as lm-evaluation-harness and HELM exist to make settings explicit and reproducible.
Errors in the answer key. As SWE-bench showed, some tasks are simply wrong. When a benchmark is near saturation, label errors can be a large share of what remains.
Funding and access. Some benchmarks are funded by labs whose models they test. Epoch AI's FrontierMath, for example, disclosed in December 2024 that OpenAI funded it; Epoch keeps a holdout set to limit the advantage. Disclosure is good practice, and it is worth checking.
How to read a benchmark claim
- Check the exact benchmark and version. "SWE-bench" could mean the full set, Lite, Verified or Pro; Humanity's Last Exam could be full or text-only.
- Check the setup. Look for the number of attempts, tool use, reasoning effort and whether it was run by the vendor or a third party.
- Look for independent replication. Trackers like the Epoch AI benchmarking hub and Artificial Analysis run evaluations themselves with published methods.
- Mind the error bars. On a 198-question test, one question is half a percentage point.
- Prefer fresh and hidden sets. Recently created or private tasks are harder to game.
- Test on your own task. Benchmarks shortlist models; your own evaluation set decides. Our playbook on how to evaluate LLM apps shows how.
Pros and cons of relying on benchmarks
Pros: cheap, repeatable comparisons; a shared vocabulary across labs; useful for spotting large capability jumps and regressions.
Cons: saturation, contamination and setup tricks; narrow task formats that rarely match real products; incentives for labs to optimize for the test.
Who should care about which benchmarks
Product teams should look at benchmarks close to their use case, such as tool calling for agents, coding benchmarks for developer tools or long-context tests for document workflows, and then build their own test set. Researchers need the full methodology and raw outputs, which open harnesses provide. Buyers comparing vendors should rely on independent trackers rather than launch slides. If your application uses retrieval, generic benchmarks say little about it; see our RAG evaluation metrics guide and our comparison of LLM evaluation frameworks.
FAQ
What is the best benchmark for LLMs?
There is no single best benchmark. For reasoning, harder recent sets such as GPQA Diamond and Humanity's Last Exam still separate models; for coding and agents, task-based benchmarks such as Terminal-Bench are more informative; for your product, your own test set matters most.
What is MMLU?
MMLU, or Massive Multitask Language Understanding, is a multiple-choice benchmark covering 57 subjects. It was the standard knowledge test for years but is now saturated, so frontier models score too close to the ceiling for it to separate them.
What is GPQA Diamond?
GPQA Diamond is the 198-question, highest-quality subset of GPQA, a benchmark of graduate-level science questions written to be hard to answer even with web search. Its small size means small score differences may not be meaningful.
What is benchmark contamination?
Contamination happens when test questions or answers appear in a model's training data, so it can score well by memorization. Fresh questions, hidden test sets and canary strings are the main defenses.
Why did OpenAI stop using SWE-bench Verified?
In February 2026 OpenAI said SWE-bench Verified no longer measured frontier coding ability, because many hard tasks had flawed tests and models showed signs of having seen the original fixes during training.
Are LLM benchmarks reliable?
They are reliable for what they narrowly measure under a stated setup, and useful for spotting big differences. They are unreliable as a guide to real-world performance on your task, especially when scores are close or reported only by the vendor.