lm-evaluation-harness
github.com · LLM Evals
EleutherAI's unified harness for running hundreds of academic benchmarks against any model with consistent prompts.
About the LLM Evals category
Frameworks and platforms for testing language models and LLM applications — from academic benchmark harnesses to CI-friendly unit tests and LLM-as-judge scoring.
lm-evaluation-harness is one of 15 llm evals tools indexed on TrueIQ. Facts on this page come from the tool's official site and public APIs; prices and features change, so confirm details on the official website before you commit.