simple-evals
github.com · LLM Evals
Lightweight reference implementations OpenAI uses to report zero-shot benchmark results transparently.
About the LLM Evals category
Frameworks and platforms for testing language models and LLM applications — from academic benchmark harnesses to CI-friendly unit tests and LLM-as-judge scoring.
simple-evals is one of 15 llm evals tools indexed on TrueIQ. Facts on this page come from the tool's official site and public APIs; prices and features change, so confirm details on the official website before you commit.