MMLU
github.com · Benchmarks
Massive Multitask Language Understanding: a multiple-choice benchmark spanning 57 subjects from math to law, introduced in 2020 and now largely saturated by frontier models.
About the Benchmarks category
The standard tests used to compare models on knowledge, reasoning, coding and abstraction. Check what each one measures, how it is scored and whether results may be contaminated.
MMLU is one of 5 benchmarks tools indexed on TrueIQ. Facts on this page come from the tool's official site and public APIs; prices and features change, so confirm details on the official website before you commit.