TrueIQ.si
Submit a tool

// category · 9 tools · 2 open source

Leaderboards

Public rankings that compare models on preference, accuracy, speed or price. Read the methodology: each leaderboard measures something different.

Artificial Analysisartificialanalysis.aiIndependent benchmarking of model intelligence, speed and price across providers, updated as new models launch.LeaderboardsFreeHN 87LiveBenchlivebench.aiBenchmark refreshed with new questions regularly and scored against objective answers to limit test-set contamination.LeaderboardsOpen source★ 1.3kAider Leaderboardsaider.chatAider's code-editing benchmarks that test how well models edit real source files across several programming languages.LeaderboardsFree★ 49.4kMTEBhuggingface.coMassive Text Embedding Benchmark and leaderboard for comparing embedding models on retrieval, clustering and more.LeaderboardsOpen source★ 3.4kBerkeley Function Calling Leaderboardgorilla.cs.berkeley.eduLeaderboard measuring how accurately models call functions and tools across languages and multi-turn settings.LeaderboardsFree★ 13kArena (LMArena)arena.aiCrowdsourced arena, formerly LMArena, where people compare anonymous model responses side by side, producing preference-based leaderboards.LeaderboardsFreeHN 3Epoch AI Benchmarking Hubepoch.aiEpoch AI's independent runs of key benchmarks with open data showing how frontier capabilities change over time.LeaderboardsFreeSEAL Leaderboardslabs.scale.comScale AI's expert-driven leaderboards using private datasets to reduce contamination and gaming.LeaderboardsFreeVellum Leaderboardvellum.aiComparison table of recent models across reasoning, coding and cost metrics, maintained by the Vellum platform.LeaderboardsFree

More categories