No single LLM leaderboard is best, because each one measures something different. Arena ranks models by millions of human votes on which answer people prefer; Artificial Analysis and the Epoch AI hub run standardized benchmark suites themselves; LiveBench refreshes its questions to limit memorization; and Scale's SEAL leaderboards use private, expert-built test sets. Use a preference board to judge style and helpfulness, a benchmark aggregator to judge raw capability, and a specialist board for tasks like coding, tool calling or embeddings.

This guide explains how each leaderboard produces its numbers and where each can mislead you. It does not tell you which model is currently on top, since that changes weekly and every board below shows it live.

How an LLM leaderboard produces its rankings

There are three basic methods, and knowing which one a board uses tells you most of what you need to know about it.

  • Human preference: real users submit prompts, see two anonymous answers and vote for the better one. Votes are turned into ratings with a statistical model. This captures what people like, including tone and formatting, but not whether a long technical answer is actually correct.
  • Benchmark aggregation: the operator runs a fixed suite of benchmarks on every model under the same settings and combines the results into an index. This is reproducible and objective but inherits every weakness of the underlying benchmarks, which we cover in our guide to LLM benchmarks.
  • Private or refreshed test sets: questions are kept secret or replaced regularly so that models cannot have seen them in training. This limits contamination but makes results harder to audit from outside.

The major leaderboards compared

LeaderboardMethodWhat it coversMain caveat
Arena (formerly LMArena)Anonymous pairwise human votesText, coding, vision, image, video, search, agentsPreference, not correctness; can be gamed by private testing
Artificial AnalysisRuns its own benchmark suite; also speed and priceIntelligence Index of 10 evaluationsMostly English and text; index weights are a choice
LiveBenchFresh questions with objective answersMath, coding, reasoning, language, data analysisSmaller question pools per category
SEAL by ScalePrivate expert-built datasets20+ benchmarks, including agentic and frontier testsScale also sells training data to labs
Epoch AI Benchmarking HubRuns and collects benchmark resultsEpoch Capabilities Index, FrontierMath and moreSome benchmarks have funding ties, which Epoch discloses
VellumCurated benchmark resultsRecent, non-saturated benchmarks onlyCompiles reported numbers rather than testing everything itself
Specialist boardsTask-specific testsAider for code editing, BFCL for tool calls, MTEB for embeddingsNarrow by design

In prose: Arena is the only major board built on human votes; Artificial Analysis, Epoch AI and LiveBench all run evaluations themselves with published methods; SEAL keeps its tests private; Vellum filters out saturated benchmarks; and the specialist boards are the right place to look when you have one specific task.

Arena (formerly LMArena)

How does LMArena work? You type a prompt, two anonymous models answer, and you vote for the better one, call it a tie, or mark both as bad. Only votes cast while the model names are hidden count. The votes are fitted with a Bradley-Terry model, which estimates each model's probability of beating each other model, and a style control adjustment reduces the advantage of longer or more heavily formatted answers. LMArena renamed itself Arena in January 2026 and moved to arena.ai (announcement). The Arena leaderboard now has separate boards for text, web development, vision, document, search, image generation, image editing, video and agents.

Arena's strength is scale and realism: it reflects what many people actually ask. Its weaknesses are well documented. The Leaderboard Illusion paper found that some providers privately tested many model variants before release and published only the best result, and that proprietary models received far more battles than open-weight ones (paper). A 2026 MIT study found that removing just two of more than 57,000 votes in one dataset was enough to change the top-ranked model (MIT News). Treat small rating gaps as ties and look at the confidence intervals Arena publishes.

Artificial Analysis

Artificial Analysis runs every model through the same benchmark suite under identical prompts and settings and combines the results into the Artificial Analysis Intelligence Index. Version 4.3.2 combines ten evaluations in four groups: agents, coding, scientific reasoning and general knowledge, weighted 30, 20, 20 and 30 percent. It includes tests such as Humanity's Last Exam, Terminal-Bench 4.0, SciCode, AA-Omniscience and GDPval-AA, and the operator estimates the index's 95 percent confidence interval at under one point (methodology). It also measures output speed, latency and price per token, which makes it the most practical board for cost and performance trade-offs.

The caveats: the weights are an editorial choice, the suite is mostly text and English, and the index changes version as benchmarks are swapped, so compare scores only within the same version.

LiveBench

LiveBench, launched in 2024, was designed against contamination. It adds new questions regularly, draws on recent sources, and scores every answer against an objective ground truth rather than an LLM judge; its authors note that LLM judges can make large errors on hard questions (paper). It is a good cross-check when a model looks suspiciously strong on older public benchmarks.

SEAL leaderboards

Scale's SEAL leaderboards use private datasets built by domain experts, covering more than 20 benchmarks including agentic tasks and Humanity's Last Exam, which Scale co-created. Keeping tests private makes them hard to train on. The trade-off is that outsiders cannot inspect the questions, and Scale also sells data services to model developers, so read its results alongside independent boards.

Epoch AI Benchmarking Hub and Vellum

The Epoch AI Benchmarking Hub runs and compiles benchmark results and publishes the Epoch Capabilities Index, plus research on trends in compute and capability. Epoch disclosed that OpenAI funded its FrontierMath benchmark and keeps a holdout set to limit any advantage. The Vellum leaderboard is a clean summary view that deliberately leaves out saturated benchmarks such as MMLU, which makes it useful for a quick look but less useful for methodology.

Specialist leaderboards

When you have a specific job, a narrow board beats a general one. The Aider leaderboards test how well models edit code in real files. The Berkeley Function Calling Leaderboard, now in its fourth version, tests tool and function calls. MTEB ranks embedding models across retrieval, clustering and classification tasks, which matters if you are building search or retrieval; pair it with our RAG evaluation metrics guide. Some benchmarks also run their own official boards: the SWE-bench site ranks results on its coding subsets, and the ARC Prize Foundation publishes leaderboards for ARC-AGI.

The Hugging Face Open LLM Leaderboard, once the default for open models, was retired in March 2025, so older articles that cite it are out of date.

How to use leaderboards without being misled

  1. Match the board to the question. Human preference for chat quality, benchmark indexes for reasoning, specialist boards for specific tasks.
  2. Look at confidence intervals. Models a few points apart are often statistically tied.
  3. Cross-check two different methods. A model that leads on both a preference board and an independent benchmark index is a safer bet than one that leads on only one.
  4. Check dates and versions. Rankings move weekly, and indexes change composition.
  5. Shortlist, then test. Take the top few candidates and run them on your own test set, as described in our playbook on how to evaluate LLM apps.

Pros and cons of leaderboards

Pros: free, continuously updated, and a fast way to narrow a long list of models; independent boards are more trustworthy than vendor launch charts.

Cons: each measures a narrow slice; preference boards can be gamed and reward style; benchmark boards inherit contamination and saturation; none know your data, latency budget or failure costs.

Who should use which leaderboard

Developers choosing a model for a product should start with Artificial Analysis for capability, speed and price, then confirm on a specialist board for their task. Teams building chat experiences should weigh Arena's preference data. Researchers and analysts will get the most from Epoch AI and LiveBench, whose methods are open. Anyone building retrieval should check MTEB. Everyone should finish with their own evaluation, using the tools in our guide to LLM evaluation frameworks.

FAQ

Which LLM leaderboard is best?

It depends on what you need. Arena is best for human preference, Artificial Analysis for standardized capability plus speed and price, LiveBench for contamination-resistant results, and specialist boards such as Aider, BFCL and MTEB for coding, tool calling and embeddings.

How does LMArena work?

Users compare two anonymous model answers to their own prompt and vote for the better one. Only anonymous votes count, and Arena fits them with a Bradley-Terry model, with a style control adjustment, to produce ratings. LMArena was renamed Arena in January 2026.

Is LMArena reliable?

It is a useful signal of what users prefer, but research has shown that private pre-release testing can inflate scores and that rankings near the top can be fragile. Treat close ratings as ties and confirm with benchmark-based boards.

What is the Artificial Analysis Intelligence Index?

It is a composite score from Artificial Analysis that combines ten evaluations across agents, coding, scientific reasoning and general knowledge, all run under identical settings. Version 4.3.2 weights those groups at 30, 20, 20 and 30 percent.

What happened to the Hugging Face Open LLM Leaderboard?

Hugging Face retired the Open LLM Leaderboard in March 2025, so articles that still cite it are out of date. Independent boards like Artificial Analysis, Epoch AI and LiveBench now fill much of that role.