The core RAG evaluation metrics split into two groups: retrieval metrics, which ask whether the right passages were fetched (context precision, context recall, context relevance), and generation metrics, which ask whether the answer sticks to those passages and addresses the question (faithfulness, also called groundedness, and answer relevancy). Faithfulness and answer relevancy can run without a reference answer, so they work on live traffic. Context recall, and in some libraries context precision, need a known-good answer, so they belong in your test set.

This guide defines each metric in plain language, shows how Ragas, DeepEval and TruLens name and compute them, and gives a symptom-to-metric table for debugging a retrieval-augmented generation pipeline.

How to evaluate RAG systems: retrieval versus generation

A RAG system fails in two different places. The retriever can return the wrong chunks, or miss the right one. The generator can ignore good chunks, add facts that are not in them, or answer a different question. If you only score the final answer, you cannot tell which half broke, so you will tune the wrong component.

Every serious RAG evaluation therefore records four things per example: the user question, the retrieved contexts, the generated response and, when you have one, a reference answer written or approved by a person. An observability tool such as Langfuse or Arize Phoenix captures the first three automatically from traces.

The core RAG evaluation metrics

Faithfulness (groundedness)

Faithfulness asks: is every claim in the answer supported by the retrieved context? Ragas computes it by breaking the response into individual claims, checking each against the context, and dividing supported claims by total claims, giving a score from zero to one (Ragas docs). TruLens calls the same idea groundedness. It needs no reference answer, which makes it the single most useful metric for catching hallucinations in production.

Note what it does not measure: a faithful answer can still be wrong if the retrieved context was wrong. That is why you pair it with retrieval metrics.

Answer relevancy

Answer relevancy asks: does the response actually address the question? In Ragas, an LLM generates a few questions that the response would answer, and the score is the average similarity between those and the original question. It penalizes incomplete or padded answers but, by design, does not check facts. DeepEval and TruLens offer equivalent metrics. No reference is needed.

Context precision

Context precision asks: are the relevant chunks ranked near the top of what was retrieved? Ragas computes a rank-weighted precision across the retrieved chunks; DeepEval's contextual precision does the same with an LLM judge. In DeepEval it requires an expected output; Ragas offers variants with and without a reference.

Context recall

Context recall asks: did retrieval fetch everything needed to answer? Because "everything needed" is only knowable if you know the right answer, it always needs a reference. Ragas splits the reference answer into claims and measures the share that can be attributed to the retrieved context (Ragas docs).

Context relevance

Context relevance, called contextual relevancy in DeepEval, asks: how much of the retrieved text is relevant to the question at all? It needs no reference, so it is a cheap way to spot a retriever that returns noisy chunks.

Same metrics, different names

Each library uses slightly different names for overlapping ideas. This table maps them.

What it checksRagasDeepEvalTruLensNeeds reference answer?
Answer supported by contextFaithfulnessFaithfulnessGroundednessNo
Answer addresses the questionResponse (answer) relevancyAnswer relevancyAnswer relevanceNo
Retrieved text is on topicContext relevance (NVIDIA metric)Contextual relevancyContext relevanceNo
Relevant chunks ranked firstContext precisionContextual precisionNot a core triad metricDeepEval yes; Ragas optional
Retrieval found everything neededContext recallContextual recallNot a core triad metricYes
Errors caused by noisy chunksNoise sensitivityNot built inNot built inYes

The short version: all three libraries cover faithfulness and answer relevance without a reference. TruLens packages the reference-free trio as its "RAG triad" of context relevance, groundedness and answer relevance (TruLens docs). Ragas and DeepEval go further on retrieval ranking and recall when you can supply reference answers.

Ragas vs DeepEval vs TruLens

All three are open source and use LLM judges under the hood for most metrics, so the guidance in our LLM-as-a-judge guide applies: validate scores against a small human-labeled set before trusting them.

RagasDeepEvalTruLens
LicenseApache 2.0Apache 2.0MIT
StyleMetrics library plus test-set generationPytest-style test cases and CIInstrumentation plus feedback functions and a dashboard
StrengthWidest RAG metric menu, synthetic test dataCI gates, thresholds, many non-RAG metricsTracing-first; RAG triad on live apps
Version checked (PyPI, Oct 2026)0.4.34.2.82.15.0

Ragas is the most RAG-specific of the three, and its repository now lives under the vibrantlabsai organization on GitHub, so update old bookmarks. DeepEval is the easiest way to turn RAG metrics into pass or fail unit tests with a threshold, which its docs default to 0.5 for every metric. TruLens, maintained under Snowflake, leans toward instrumenting a running app and scoring it continuously.

A five-step RAG evaluation workflow

  1. Build a small labeled set. Collect 50 to 100 real user questions, including hard and ambiguous ones, and write or approve a reference answer for each. Start with real traffic if you have it.
  2. Score retrieval first. Run context recall and context precision. If recall is low, the answer was never available to the model, and no prompt change will fix that.
  3. Score generation next. Run faithfulness and answer relevancy. Low faithfulness with good recall means the model is ignoring or embellishing the context.
  4. Read the failures. Open the lowest-scoring examples and read the traces. Metrics point to where to look; they do not explain why. Our playbook on how to evaluate LLM apps covers this error-analysis step.
  5. Gate and monitor. Put the labeled set in CI with thresholds, and run the reference-free metrics on a sample of production traces in your observability tool.

Diagnosing problems: symptom to metric

Symptom users reportMetric that confirms itLikely fix
"It made that up"Low faithfulnessTighter grounding instructions, citations, lower temperature, smaller context
"It didn't know something that's in our docs"Low context recallBetter chunking, hybrid search, query rewriting, more top-k
"The answer is buried in fluff"Low answer relevancyPrompt for concise answers; check the question is passed correctly
"It cites the wrong document"Low context precisionAdd a reranker; improve metadata filters
"Answers got worse after adding documents"Rising noise sensitivity, falling context relevanceFilter or deduplicate the index; rerank

Spoken plainly: hallucinations show up as low faithfulness, missing knowledge as low recall, rambling as low relevancy, and wrong citations as low precision. Each points to a different part of the pipeline. For other ways to catch unsupported claims, including detector models and sampling-based methods, see our guide to LLM hallucination detection.

Pros and cons of automated RAG metrics

Pros

  • They separate retrieval failures from generation failures.
  • Reference-free metrics run on live traffic without labeling.
  • They make regressions visible when you change chunking, embeddings or models.

Cons

  • Most are LLM judges, so they cost money and have their own error rates.
  • Scores from different libraries are not comparable, even for metrics with the same name.
  • A high score on a weak test set is meaningless; the labeled questions matter more than the metric choice.

Who should use which tool

Choose Ragas if you want the broadest set of RAG metrics and help generating synthetic test questions. Choose DeepEval if your priority is CI: failing a build when faithfulness drops. Choose TruLens if you want to instrument a running app and watch the RAG triad over time. Many teams combine one of these with a tracing platform. For a wider view of testing tools, see our roundup of LLM evaluation frameworks.

FAQ

What are the main metrics for evaluating RAG?

Faithfulness and answer relevancy for the generated answer, and context precision, context recall and context relevance for retrieval. Together they show whether a bad answer came from missing context or from the model misusing good context.

How do you evaluate RAG retrieval?

Use context recall to check that everything needed was retrieved, which requires reference answers, and context precision to check that relevant chunks are ranked first. Context relevance gives a reference-free view of how noisy the retrieved text is.

What is the difference between Ragas and DeepEval?

Both are open-source and Apache 2.0 licensed with overlapping RAG metrics. Ragas focuses on RAG metrics and synthetic test-set generation; DeepEval is built around pytest-style test cases with thresholds for CI and covers many non-RAG metrics, such as agent and conversation checks.

Can you evaluate RAG without ground truth?

Yes, partly. Faithfulness, answer relevancy and context relevance need only the question, retrieved contexts and response. Context recall, and context precision in some libraries, need a reference answer.

How do you measure RAG accuracy?

Compare responses with reference answers, for example with Ragas factual correctness or an LLM judge that checks the answer against the reference. Pair that with faithfulness, because an answer can match the reference by luck while ignoring the retrieved context.