TrueIQ.si
Submit a tool

// category · 15 tools · 12 open source

LLM Evals

Frameworks and platforms for testing language models and LLM applications — from academic benchmark harnesses to CI-friendly unit tests and LLM-as-judge scoring.

DeepEvaldeepeval.comPytest-style framework for unit-testing LLM apps with metrics for hallucination, relevance, RAG quality and more.LLM EvalsOpen source★ 18.6kGiskardgiskard.aiTesting library and platform that scans AI agents and models for hallucinations, bias and security issues.LLM EvalsOpen source★ 5.9kPromptfoopromptfoo.devDeveloper CLI for testing prompts and models with assertions, side-by-side comparisons and automated red-teaming in CI.LLM EvalsOpen source★ 25.7kEvidentlyevidentlyai.comOpen-source library for evaluating, testing and monitoring ML and LLM systems, from data drift to response quality.LLM EvalsOpen source★ 8kOpenCompassopencompass.org.cnComprehensive evaluation platform from Shanghai AI Lab supporting many models, datasets and a public leaderboard.LLM EvalsOpen source★ 7.5kOpenAI Evalsgithub.comFramework and open registry of evaluations for testing language models and the systems built on them.LLM EvalsOpen source★ 19.6kRagasragas.ioLibrary of metrics and synthetic test-set generation for evaluating retrieval-augmented generation pipelines.LLM EvalsOpen source★ 15.9kTruLenstrulens.orgInstrumentation and feedback functions for evaluating and tracking LLM apps, known for the RAG triad metrics.LLM EvalsOpen source★ 3.6klm-evaluation-harnessgithub.comEleutherAI's unified harness for running hundreds of academic benchmarks against any model with consistent prompts.LLM EvalsOpen source★ 14.1klightevalgithub.comHugging Face toolkit for evaluating LLMs across many backends with reusable tasks and detailed sample-level results.LLM EvalsOpen source★ 2.6ksimple-evalsgithub.comLightweight reference implementations OpenAI uses to report zero-shot benchmark results transparently.LLM EvalsOpen source★ 4.7kHELMcrfm.stanford.eduStanford's Holistic Evaluation of Language Models, measuring accuracy, robustness, bias and efficiency across many scenarios.LLM EvalsOpen source★ 2.9kBraintrustbraintrust.devEnd-to-end platform for evaluating AI products, with datasets, scorers, prompt playgrounds and production logging.LLM EvalsFreemiumHN 1Patronus AIpatronus.aiAutomated evaluation platform with research-backed judge models for catching hallucinations and unsafe outputs.LLM EvalsSee siteHN 1Galileogalileo.aiAI reliability platform that evaluates, monitors and guards agent and LLM applications with low-latency metrics.LLM EvalsFreemium

More categories