If you want to know how to evaluate LLM applications, start by reading real outputs, not by picking metrics. Collect traces, write down what goes wrong, group those failures into categories, and only then build automated checks for the categories that matter. Use plain code checks wherever a rule can decide pass or fail, use an LLM judge only for fuzzy criteria, and validate that judge against human labels before you trust it.

This playbook walks through that loop in order, shows which kind of check fits which failure, and names open-source and hosted tools for each stage. It is written for engineers and product managers shipping a chatbot, a retrieval (RAG) assistant, or any feature that calls a language model.

Why generic LLM metrics are not enough

Public benchmarks tell you how a model does on someone else's test. They say little about whether your support bot cites the right refund policy or whether your summarizer drops the one sentence your users care about. For that background, see our explainer on what LLM benchmarks actually measure.

Ready-made metric libraries have the same problem at a smaller scale. A "helpfulness" or "coherence" score that runs on every output feels rigorous, but it rarely maps to a failure your users actually hit. The evals FAQ by Hamel Husain and Shreya Shankar, one of the most widely shared practitioner resources on this topic, calls error analysis "the most important activity in evals" for exactly this reason: it tells you which evals to write in the first place. You can read it at hamel.dev.

The three kinds of checks

Every LLM evaluation, whatever the tool, comes down to three kinds of checks. Knowing which one to reach for saves money and avoids noisy scores.

Check typeGood forCost per runMain risk
Code-based assertionFormat, JSON schema, required fields, banned phrases, exact answers, tool called or notNear zeroOnly catches what a rule can express
LLM judgeTone, faithfulness to sources, policy compliance, whether an answer is completeOne model call per checkJudge bias and drift; needs validation
Human reviewDiscovering new failure types, labeling ground truth, final sign-offHighestSlow, and reviewers disagree without clear guidelines

In plain terms: code checks are cheap and exact, so use them first. LLM judges handle the criteria a regular expression cannot, but they are classifiers that can be wrong, so treat them like any other model you ship. Human review is where your definition of "good" comes from, and it never fully goes away.

How to evaluate LLM outputs in six steps

1. Capture traces

A trace is the full record of one request: the user input, the system prompt, retrieved documents, tool calls, intermediate model responses and the final answer. You cannot do error analysis on the final answer alone, because the root cause is often upstream, such as a retrieval step that returned the wrong document.

An observability tool makes this painless. Langfuse and Arize Phoenix are open-source options you can self-host; hosted platforms such as LangSmith and Braintrust do the same job. Our comparison of LLM observability tools covers the trade-offs, including licensing and which vendors changed owners in 2026.

2. Do error analysis by hand

Pull a varied sample of traces and read them. Write a short free-text note on anything that looks wrong from the user's point of view. Focus on the first thing that went wrong in each trace, since early mistakes cause later ones.

How many? The Husain and Shankar FAQ suggests starting with about 100 diverse traces and annotating at least the first 30 yourself before letting an AI assistant suggest more. Have one domain expert own the final call on what counts as a failure, so labels stay consistent.

3. Build a failure taxonomy

Group your notes into categories, such as "cites the wrong policy," "ignores the user's language," "invents an order number," or "returns malformed JSON." Count how often each one appears. This turns a pile of anecdotes into a ranked list of problems, and it usually reveals that a few categories cause most of the pain.

Many items on the list will be plain bugs, like a broken prompt template or a missing document in the index. Fix those directly. Build automated evaluators only for failure types that are likely to come back.

4. Write the cheapest check that catches each failure

For each remaining category, ask whether code can decide it. Malformed JSON, a missing citation marker, a response over a length limit, or a wrong tool call can all be checked with a few lines of code. Those checks are deterministic, fast and free to run on every commit.

When the failure needs judgment, such as "the answer contradicts the retrieved document," write an LLM judge with a narrow, single-purpose prompt and a binary pass or fail output. Binary labels are easier to agree on and easier to validate than one-to-five ratings. Our guide to LLM-as-a-judge covers prompt design, known biases and how to measure judge accuracy.

For retrieval systems, split the evaluation into retrieval quality and answer quality, so you know which half broke. The RAG evaluation metrics guide maps the standard metrics, such as faithfulness and context recall, across Ragas, DeepEval and TruLens.

5. Validate your judges

An LLM judge is a classifier. Before you use its scores to make decisions, label a set of examples yourself and compare. Split your labeled examples into three groups: a few to put in the judge prompt as examples, a development set you use while editing the prompt, and a held-out test set you only check at the end.

Report two numbers on the test set: how often the judge catches real failures (the true positive rate) and how often it correctly passes good outputs (the true negative rate). A single "accuracy" figure hides a judge that passes everything. The FAQ cited above suggests labeling roughly 100 to 200 examples per failure mode for this step.

6. Run evals in CI and on live traffic

Offline and online evaluation answer different questions.

  • In CI (offline): run a curated dataset of core workflows, past bugs and edge cases every time you change a prompt, model or retrieval setting. Favor code checks here because they are cheap and stable. Fail the build if a known regression returns.
  • In production (online): sample live traces and run evaluators asynchronously, so users never wait on them. You usually have no reference answer for live traffic, so reference-free LLM judges do more of the work here. Track the trend and investigate when it moves.

The two loops feed each other. Every new failure you discover in production becomes a test case in the CI dataset, so it cannot quietly return.

Choosing tools for each stage

You do not need one platform for everything, and in 2026 several popular vendors changed owners, which is worth weighing. Here is how the main options line up.

StageOpen-source optionsHosted options
Tracing and reviewLangfuse (MIT core), Arize Phoenix (Elastic License 2.0), Opik (Apache 2.0)LangSmith, Braintrust, HoneyHive
Test runner and CI gatePromptfoo (MIT), DeepEval (Apache 2.0), Inspect AI (MIT)Braintrust, Confident AI
RAG metricsRagas, DeepEval, TruLensMost platforms above
Adversarial testingPromptfoo red teaming, garak, PyRIT, DeepTeamVendor platforms

To summarize the table: pair one tracing tool with one test runner. Promptfoo suits teams that like a YAML config and a command-line workflow, while DeepEval suits Python teams who want evals that look like pytest unit tests. Inspect AI, from the UK AI Security Institute (documentation), suits teams that need more rigorous, agent-style evaluations. Braintrust bundles tracing, datasets and experiments in one hosted product if you would rather not run infrastructure. Our roundup of LLM evaluation frameworks compares them in more depth.

Once quality checks are in place, add adversarial tests for prompt injection and data leakage. The LLM red teaming tools guide covers the main open-source scanners.

Pros and cons of an eval-first workflow

Pros

  • You can change models or prompts with evidence instead of gut feel.
  • Regressions are caught in CI rather than by users.
  • The failure taxonomy doubles as a product roadmap: it shows what to fix next.

Cons

  • It takes real human time. The FAQ authors report spending 60 to 80 percent of development time on error analysis and evaluation in the projects they have worked on.
  • LLM judges add cost and need periodic re-validation when you change the judge model.
  • A suite that passes 100 percent of the time is not giving you information. Keep adding harder cases.

Who this playbook is for

It fits small teams shipping their first LLM feature, who should start with the hand review in steps 2 and 3 and a handful of code checks. It also fits larger teams with an existing product, who will get the most value from steps 5 and 6: validated judges, a CI gate and sampled production monitoring. If you are comparing base models rather than evaluating your own application, you want public benchmarks and LLM leaderboards instead, plus a small custom test set built from your own tasks.

FAQ

How do you evaluate the performance of an LLM?

For your own application, measure task success on examples drawn from real usage, using code checks for objective rules and validated LLM judges for subjective criteria. For a base model, look at several independent benchmarks and leaderboards, then confirm on a small test set built from your own tasks.

What are the most common LLM evaluation metrics?

For open-ended apps, the most useful metrics are pass rates on application-specific checks, such as "answer cites a retrieved source" or "output is valid JSON." For RAG, faithfulness, answer relevancy, context precision and context recall are standard. Older text-overlap metrics like BLEU and ROUGE correlate poorly with human judgment on open-ended tasks.

What is the difference between offline and online LLM evaluation?

Offline evaluation runs a fixed dataset before deployment, usually in CI, to catch regressions. Online evaluation scores a sample of live production traffic after deployment to find new failures and track quality over time.

How many examples do I need to evaluate an LLM app?

Start by reading about 100 varied traces, annotating at least 30 yourself. A CI regression set often grows past 100 examples. To validate an LLM judge, label roughly 100 to 200 examples for each failure mode it checks.

Can I use an LLM to evaluate another LLM?

Yes, and it is standard practice for subjective criteria. Treat the judge as a classifier: give it one narrow question with a pass or fail answer, and measure its agreement with human labels on a held-out set before relying on it.

How often should I run LLM evals?

Run cheap code checks on every change. Run LLM-judge suites when prompts, models or retrieval settings change, and on a sample of production traffic continuously. Retire checks that always pass and replace them with harder cases.