BIG-bench
github.com · Benchmarks
Beyond the Imitation Game: a collaborative benchmark of 200+ diverse tasks from 450 authors at 132 institutions, built to probe and extrapolate language model capabilities.
About the Benchmarks category
The standard tests used to compare models on knowledge, reasoning, coding and abstraction. Check what each one measures, how it is scored and whether results may be contaminated.
BIG-bench is one of 5 benchmarks tools indexed on TrueIQ. Facts on this page come from the tool's official site and public APIs; prices and features change, so confirm details on the official website before you commit.