HELM
crfm.stanford.edu · LLM Evals
Stanford's Holistic Evaluation of Language Models, measuring accuracy, robustness, bias and efficiency across many scenarios.
About the LLM Evals category
Frameworks and platforms for testing language models and LLM applications — from academic benchmark harnesses to CI-friendly unit tests and LLM-as-judge scoring.
HELM is one of 15 llm evals tools indexed on TrueIQ. Facts on this page come from the tool's official site and public APIs; prices and features change, so confirm details on the official website before you commit.