LLM evaluation metrics fall into four groups: exact and statistical metrics such as exact match, F1, BLEU and ROUGE, which compare output with a reference; embedding metrics such as BERTScore, which compare meaning; execution metrics such as pass@k, which run code against tests; and LLM-judge metrics such as G-Eval rubrics, faithfulness and answer relevancy, which use a model to grade open-ended output. For most modern applications, the useful set is a few deterministic checks plus one or two calibrated judge metrics tied to failures you have actually seen, alongside latency and cost.
This guide explains what each metric measures, where it breaks down, and how to choose a small set that tracks real quality.
The main families of LLM evaluation metrics
| Family | Examples | Needs a reference answer? | Best for |
|---|---|---|---|
| Exact and rule-based | Exact match, regex, JSON schema validity | Sometimes | Classification, extraction, structured output |
| Overlap (n-gram) | BLEU, ROUGE, token F1 | Yes | Translation, summarization baselines, short-answer QA |
| Embedding similarity | BERTScore, cosine similarity | Yes | Paraphrase-tolerant comparison |
| Execution-based | pass@k, unit tests | Tests instead of answers | Code generation |
| LLM-as-a-judge | G-Eval rubrics, pairwise preference, correctness | Optional | Open-ended answers, tone, helpfulness |
| RAG-specific | Faithfulness, context precision and recall, answer relevancy | Often | Retrieval-augmented apps |
| Safety and operational | Toxicity, refusal rate, latency, cost per request | No | Production readiness |
In plain terms: use deterministic checks wherever the right answer can be checked by code, overlap and embedding metrics only as rough baselines, execution for code, and LLM judges for everything open-ended, while always tracking latency and cost.
Exact match, F1 and rule-based checks
Exact match scores 1 if the output equals the reference and 0 otherwise. It is ideal for classification labels, multiple-choice answers and short factual answers, which is why many academic benchmarks use it. Token-level F1 gives partial credit for overlapping words and is common in extractive question answering.
Rule-based checks are the unsung heroes of application evaluation: does the output parse as valid JSON, match a schema, include a required disclaimer, avoid a banned phrase, or stay under a length limit? They are fast, free and never disagree with themselves. Tools such as Promptfoo and DeepEval make them one-line assertions.
BLEU vs ROUGE
BLEU was introduced in 2002 for machine translation and measures how many n-grams, meaning short word sequences, in the output also appear in one or more reference translations, with a penalty for outputs that are too short (paper). It is precision-oriented: it rewards saying things the reference says.
ROUGE, introduced in 2004 for summarization, counts overlapping units such as n-grams and longest common subsequences between a generated summary and reference summaries (paper). It is recall-oriented: it rewards covering what the reference covers.
Both share the same weakness for LLM output: a correct answer phrased differently from the reference scores poorly, and a fluent but wrong answer that reuses the reference's words can score well. Use them as cheap regression signals for translation or summarization, not as measures of quality.
BERTScore and embedding similarity
BERTScore compares outputs with references using contextual embeddings from a pretrained language model, matching each token to its most similar token in the other text, so paraphrases get credit (paper). Simpler cosine similarity between sentence embeddings does the same at the whole-text level. These are better than n-gram overlap for open phrasing, but they still measure similarity to a reference, not correctness: two sentences that differ only by "not" can look very similar.
pass@k for code
For code generation, the most meaningful metric is whether the code works. pass@k is the probability that at least one of k generated samples passes the unit tests for a problem. It was popularized by the paper that introduced the HumanEval benchmark, which also gave an unbiased way to estimate it from more than k samples (paper). pass@1 is the number that matters for most products, since users usually see one answer.
LLM-as-a-judge metrics
For open-ended output, a second model grades the first. G-Eval asks a judge model to follow a rubric, generate evaluation steps and score the output, and reported better alignment with human judgments than older metrics on summarization and dialogue tasks (paper). Other judge formats include pairwise comparison, where the judge picks the better of two answers, and reference-guided correctness, where it compares an answer with a gold answer. Research on MT-Bench and Chatbot Arena found strong judges can agree with human preferences at levels similar to agreement between humans, but also documented position, verbosity and self-preference biases (paper).
Judge metrics are only as good as their calibration. Write binary or narrowly scoped criteria, check agreement against human labels, and track judge drift. Our LLM-as-a-judge guide covers this in depth.
RAG, hallucination and safety metrics
Retrieval-augmented applications need metrics for both halves of the pipeline: context precision and recall for retrieval, and faithfulness and answer relevancy for generation. Ragas and DeepEval implement these; see our RAG evaluation metrics guide for definitions and the LLM hallucination detection guide for unsupported-claim checks.
Safety metrics include toxicity, bias and refusal rates, and attack success rate in red teaming, covered in our LLM red teaming tools guide. Operational metrics, meaning latency, time to first token, tokens per request and cost per request, belong in every scorecard, because a quality gain that doubles cost may not be worth shipping.
Model benchmarks vs application metrics
Benchmark harnesses such as lm-evaluation-harness and HELM mostly use exact match, multiple-choice accuracy and pass@k on public tasks, so results are comparable across models. Application evaluation uses your own data and usually leans on rule-based checks and judges. Both are valid; they answer different questions, as our LLM benchmarks explainer explains.
How to choose the right metrics
- Start from failures, not a metric menu. Review 50 to 100 real outputs, write down what goes wrong, and group the failures.
- Use code wherever possible. If a failure can be checked deterministically, such as format, length or a required field, write a rule-based check.
- Add a judge only for what code cannot check, such as tone, completeness or correctness of free text, with one narrow criterion per judge.
- Calibrate. Label a sample by hand and measure how often each automated metric agrees.
- Keep it small. Three to six metrics that map to real failures beat a dashboard of twenty generic scores.
- Track cost and latency next to quality on every run.
- Wire it into CI and production. Use a framework from our LLM evaluation frameworks guide, or a hosted platform such as Braintrust, and follow our playbook on how to evaluate LLM apps.
Pros and cons of each metric type
Deterministic and overlap metrics: cheap, fast and reproducible, but blind to meaning and paraphrase.
Embedding metrics: tolerate paraphrase, but still measure similarity rather than correctness.
Execution metrics: directly measure whether code works, but only as well as the tests.
Judge metrics: handle open-ended quality and scale far beyond human review, but cost money, carry biases and need calibration.
Who should use which metrics
Teams building extraction, classification or structured-output features should rely mainly on exact match and rule-based checks. Translation and summarization teams can keep BLEU or ROUGE as baselines but should add judge or faithfulness metrics. Code tools should measure pass@1 on realistic tests. Chat assistants and agents need rubric-based judges plus safety and cost metrics. Model researchers comparing checkpoints should use benchmark harnesses with standard metrics.
FAQ
What are the most common LLM evaluation metrics?
Exact match and F1 for short answers, BLEU and ROUGE for translation and summarization, BERTScore for semantic similarity, pass@k for code, and LLM-judge metrics such as G-Eval rubrics, faithfulness and answer relevancy for open-ended output, plus latency and cost.
What is the difference between BLEU and ROUGE?
BLEU measures how much of the generated text's n-grams appear in the reference, so it is precision-oriented and was designed for translation. ROUGE measures how much of the reference is covered by the output, so it is recall-oriented and was designed for summarization.
Is BLEU a good metric for LLMs?
Not on its own. BLEU penalizes correct answers that are phrased differently from the reference and can reward fluent wrong answers, so it is best used as a cheap regression signal alongside judge-based or task-specific metrics.
What is pass@k?
pass@k is the probability that at least one of k code samples generated for a problem passes its unit tests. pass@1, a single attempt, is the most relevant number for most real products.
What is G-Eval?
G-Eval is an LLM-as-a-judge method in which a model follows a rubric, writes evaluation steps and scores an output. It showed better agreement with human judgments than older automatic metrics on summarization and dialogue tasks, but like all judges it needs calibration.