Building a Retrieval-Augmented Generation (RAG) application is one of the most effective ways to supercharge Large Language Models (LLMs) with private, domain-specific knowledge. However, moving a RAG system from a basic prototype to a production-ready application presents a major hurdle: How do you know if your RAG pipeline is actually performing well?
When an LLM produces an inaccurate or irrelevant answer in a RAG system, identifying the root cause can be tricky. Did the retriever fail to fetch the correct context from your vector database? Or did the generator (LLM) hallucinate despite receiving the right information?
To solve this problem, developers turn to RAGAS (RAG Assessment), the open-source evaluation framework designed specifically to benchmark, score, and optimize RAG pipelines.
In this comprehensive guide, we will explore what RAGAS is, its core evaluation philosophy, the 5 key metrics every AI engineer should know, hands-on Python code implementation, real-world examples, and key interview concepts.
What Is the RAGAS Framework? #
RAGAS (short for RAG Assessment) is an evaluation framework tailored for measuring the performance of Retrieval-Augmented Generation systems.
Rather than treating a RAG pipeline as an unpredictable “black box,” RAGAS decouples the system into its two core architectural components:
- The Retriever: Responsible for fetching relevant context chunks from a vector database or knowledge base.
- The Generator (LLM): Responsible for reading the retrieved context and synthesizing a clear, accurate answer for the user.
By evaluating these components independently and together, RAGAS gives developers actionable numeric scores (ranging from 0.0 to 1.0) that directly point to pipeline bottlenecks.
The “LLM-as-a-Judge” Philosophy
Traditional Natural Language Processing (NLP) metrics—such as BLEU or ROUGE—rely on exact keyword and string matching. These metrics fail when evaluating modern RAG pipelines because LLMs express facts using varying phrasing, synonyms, and structures.
RAGAS solves this by pioneering the LLM-as-a-Judge approach. It leverages the Natural Language Understanding (NLU) capabilities of an advanced LLM (such as GPT-4 or other Company LLM ) to evaluate intent, factual consistency, semantic similarity, and contextual relevance. This brings human-level grading accuracy at scale.
Key Concepts and Terminology #
Before diving into specific metrics, it is helpful to understand the standard terminology used across RAGAS evaluation workflows:
- Sample: A single test entry in your evaluation dataset consisting of a user query, retrieved context, generated answer, and reference ground truth.
- User Input (Query): The raw question or prompt submitted by the user to the RAG system.
- Retrieved Context: The list of text chunks fetched from your vector database by the retriever in response to the query.
- Response: The final text answer generated by your LLM using the retrieved context.
- Reference (Grounded Response): The “gold standard” or human-curated answer key used as a benchmark for accuracy.
- Reference Context: The ideal context chunks that should have been retrieved from the knowledge base for a specific query.
- Evaluation Dataset: A curated collection of samples used to benchmark and score the RAG system.
- Experiment: A specific configuration run of your RAG pipeline where a single hyperparameter (e.g., chunk size, top-k retrieval count, or prompt template) is adjusted to test performance changes.
The 5 Core RAGAS Evaluation Metrics #
RAGAS provides targeted metrics to evaluate both the retrieval and generation phases. Here is a detailed breakdown of the five primary metrics.
1. Context Recall (Retriever Metric) #
Context Recall measures the completeness of the information retrieved by your retriever. It evaluates whether the retrieved context contains all the important facts needed to construct a complete answer.
Target Component: Retriever
Core Focus: Completeness & Coverage
Core Question: Did the retriever retrieve all the necessary information needed to answer the query?
How Context Recall Works #
The evaluation process generally works as follows:
- An evaluator LLM extracts individual factual claims from the human-curated reference answer.
- It checks whether those reference claims can be supported by the retrieved context.
- The score is calculated as the number of reference claims supported by the retrieved context divided by the total number of reference claims.
The formula is:
Ideal Score: 1.0 (Higher is better)
A low Context Recall score indicates that the retriever failed to retrieve some of the information necessary to answer the query completely.
Example :- Suppose a user asks: #
“What are the benefits of using RAG in an AI application?”
The reference answer contains 4 important claims:
- RAG provides the LLM with external knowledge.
- RAG can reduce hallucinations by grounding responses in retrieved information.
- RAG allows knowledge to be updated without retraining the LLM.
- RAG can provide source information for generated answers.
Now imagine your retriever returns context that supports only these 3 claims:
- External knowledge ✅
- Reducing hallucinations ✅
- Updating knowledge without retraining ✅
- Providing source information ❌
Therefore:
So the Context Recall score is 0.75 (75%).
This means the retriever successfully retrieved context covering 75% of the important information, but it missed the information about source attribution.
# Simple Example
# Reference answer contains 4 important factual claims
reference_claims = [
"RAG provides external knowledge to the LLM.",
"RAG can reduce hallucinations.",
"RAG allows knowledge to be updated without retraining the LLM.",
"RAG can provide source information for answers."
]
# Retrieved context contains information supporting only 3 claims
retrieved_context = """
RAG provides external knowledge to the LLM.
It can reduce hallucinations by grounding responses in retrieved documents.
Knowledge can be updated without retraining the LLM.
"""
# Simplified evaluation:
# In a real RAG system, an evaluator LLM would determine
# whether each reference claim is supported by the retrieved context.
supported_claims = [
True, # External knowledge → found
True, # Reduce hallucinations → found
True, # Update knowledge → found
False # Source information → missing
]
# Calculate Context Recall
context_recall = (
sum(supported_claims) / len(reference_claims)
)
print(f"Context Recall: {context_recall:.2f}")
print(f"Context Recall: {context_recall * 100:.0f}%")
Context Recall: 0.75
Context Recall: 75%
Implementing Context Recall with Ragas and Ollama
Install Library
- pip install ragas
- pip install python-dotenv
from google import genai
from google.genai import types
import json
import os
from dotenv import load_dotenv
load_dotenv()
# make .env file and GOOGLE_API_KEY=" "
# ============================================================
# 1. Create Gemini Client
# ============================================================
client = genai.Client()
# ============================================================
# 2. Reference Claims
# ============================================================
reference_claims = [
"RAG provides external knowledge to the LLM.",
"RAG can reduce hallucinations by grounding responses in retrieved documents.",
"RAG allows knowledge to be updated without retraining the LLM.",
"RAG can provide source citations such as document names or URLs."
]
# ============================================================
# 3. Retrieved Context
# ============================================================
retrieved_context = """
RAG provides external knowledge to the LLM.
RAG can reduce hallucinations by grounding responses
in retrieved documents.
Knowledge can be updated without retraining the LLM.
"""
# ============================================================
# 4. Evaluation Prompt
# ============================================================
prompt = f"""
You are a strict evaluator for Context Recall in a RAG system.
Your task is to evaluate EVERY reference claim.
REFERENCE CLAIMS:
{json.dumps(reference_claims, indent=2)}
RETRIEVED CONTEXT:
{retrieved_context}
RULES:
1. Evaluate every reference claim.
2. Do not skip any claim.
3. A claim is supported ONLY when the retrieved context
explicitly states the information or clearly entails it.
4. If the information is missing from the retrieved context,
mark it as false.
5. Do NOT use your own knowledge.
6. Do NOT make assumptions.
7. Only use the Retrieved Context to make the decision.
Return one evaluation for every reference claim.
"""
# ============================================================
# 5. JSON Schema
# ============================================================
response_schema = {
"type": "OBJECT",
"properties": {
"evaluations": {
"type": "ARRAY",
"items": {
"type": "OBJECT",
"properties": {
"claim": {
"type": "STRING"
},
"supported": {
"type": "BOOLEAN"
}
},
"required": [
"claim",
"supported"
]
}
}
},
"required": [
"evaluations"
]
}
# ============================================================
# 6. Call Gemini
# ============================================================
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=prompt,
config=types.GenerateContentConfig(
temperature=0,
response_mime_type="application/json",
response_schema=response_schema
)
)
# ============================================================
# 7. Check Gemini Response
# ============================================================
print("Raw Gemini Response:")
print(response.text)
# ============================================================
# 8. Parse JSON
# ============================================================
if not response.text:
raise ValueError(
"Gemini returned an empty response. "
"Check the model response and safety feedback."
)
result = json.loads(response.text)
# ============================================================
# 9. Get Evaluations
# ============================================================
evaluations = result["evaluations"]
# ============================================================
# 10. Calculate Context Recall
# ============================================================
total_claims = len(reference_claims)
supported_claims = sum(
1
for evaluation in evaluations
if evaluation["supported"] is True
)
context_recall = supported_claims / total_claims
# ============================================================
# 11. Display Evaluation
# ===========================================================
print("\n" + "=" * 60)
print("CONTEXT RECALL EVALUATION")
print("=" * 60)
for evaluation in evaluations:
symbol = "✓" if evaluation["supported"] else "✗"
print(
f"{symbol} {evaluation['claim']}"
)
# ============================================================
# 12. Display Final Score
# ============================================================
print("\n" + "=" * 60)
print(f"Supported Claims : {supported_claims}")
print(f"Total Claims : {total_claims}")
print(f"Context Recall : {context_recall:.2f}")
print(f"Context Recall : {context_recall * 100:.0f}%")
print("=" * 60)
Raw Gemini Response:
{
"evaluations": [
{
"claim": "RAG provides external knowledge to the LLM.",
"supported": true
},
{
"claim": "RAG can reduce hallucinations by grounding responses in retrieved documents.",
"supported": true
},
{
"claim": "RAG allows knowledge to be updated without retraining the LLM.",
"supported": true
},
{
"claim": "RAG can provide source citations such as document names or URLs.",
"supported": false
}
]
}
============================================================
CONTEXT RECALL EVALUATION
============================================================
✓ RAG provides external knowledge to the LLM.
✓ RAG can reduce hallucinations by grounding responses in retrieved documents.
✓ RAG allows knowledge to be updated without retraining the LLM.
✗ RAG can provide source citations such as document names or URLs.
============================================================
Supported Claims : 3
Total Claims : 4
Context Recall : 0.75
Context Recall : 75%
2. Context Precision #
Context Precision is a retrieval evaluation metric used in RAG systems to measure whether the relevant retrieved chunks are ranked higher than irrelevant chunks.
It focuses not only on whether relevant information was retrieved, but also on where that information appears in the ranked retrieval results.
In simple terms: Context Precision answers the question:
“Are the most useful chunks appearing near the top of the retrieved context?”
Key Information #
| Property | Description |
|---|---|
| Target Component | Retriever & Re-ranker |
| Primary Focus | Relevance and ranking order |
| Input | User query + retrieved chunks |
| Output | Score between 0 and 1 |
| Ideal Score | 1.0 |
| Higher Score | Better retrieval ranking |
| Lower Score | More irrelevant chunks are ranked before relevant ones |
2.1 Why Context Precision Matters #
A RAG system may retrieve the correct information but still perform poorly if the relevant chunks are ranked too low.
For example, suppose the query is:
“What are the benefits of Retrieval-Augmented Generation?”
The retriever returns the following top-5 chunks:
| Rank | Retrieved Chunk | Relevant |
|---|---|---|
| 1 | RAG combines retrieval with generation | ✅ |
| 2 | History of neural networks | ❌ |
| 3 | RAG can reduce hallucinations | ✅ |
| 4 | CNN architecture and image classification | ❌ |
| 5 | RAG allows knowledge to be updated without retraining | ✅ |
The retriever found 3 relevant chunks, but two irrelevant chunks were placed between them.Context Precision penalizes this type of ranking.
2.2 How Context Precision Works #
The evaluation can be understood in three steps:
Step 1 — Retrieve Top-K Chunks #
The retriever returns a ranked list of chunks:
Query
↓
Retriever
↓
Top-K Retrieved Chunks
For example:
Rank 1 → Relevant
Rank 2 → Irrelevant
Rank 3 → Relevant
Rank 4 → Irrelevant
Rank 5 → Relevant
Step 2 — Determine Relevance #
Each retrieved chunk is evaluated against the query.
A simple relevance representation is:
[1, 0, 1, 0, 1]
Where:
1 = Relevant
0 = Irrelevant
Step 3 — Calculate Precision at Relevant Positions #
For every relevant chunk, calculate the precision up to that rank. The resulting values are combined into the final Context Precision score.
2.3 Formula #
Let:
The precision at rank is:
Context Precision is then calculated as:
Where:
- = number of retrieved chunks
- = relevance indicator for rank
- = precision among the first retrieved chunks
- = total number of relevant retrieved chunks
Important intuition #
The metric gives more credit when relevant chunks occur early in the ranking.
2.4 Worked Example #
Consider the following retrieval result:
Rank 1 → Relevant ✅
Rank 2 → Irrelevant ❌
Rank 3 → Relevant ✅
Rank 4 → Irrelevant ❌
Rank 5 → Relevant ✅
Therefore:
relevance = [1, 0, 1, 0, 1]
Precision@1 #
There is 1 relevant chunk in the first 1 result:
Precision@3 #
There are 2 relevant chunks in the first 3 results:
Precision@5 #
There are 3 relevant chunks in the first 5 results:
Only ranks 1, 3, and 5 contain relevant chunks, so: Context Precision≈0.756
Sample code Example
def context_precision(relevance):
"""
Calculate Context Precision.
Parameters:
relevance: list[int]
1 = relevant chunk
0 = irrelevant chunk
Returns:
float: Context Precision score
"""
relevant_count = 0
precision_sum = 0.0
for rank, is_relevant in enumerate(relevance, start=1):
if is_relevant == 1:
relevant_count += 1
precision_at_k = relevant_count / rank
precision_sum += precision_at_k
if relevant_count == 0:
return 0.0
return precision_sum / relevant_count
relevance = [1, 0, 1, 0, 1]
score = context_precision(relevance)
print(f"Context Precision: {score:.3f}")
Context Precision: 0.756
Complete Gemini (LLM) Example
import json
import os
from dotenv import load_dotenv
from google import genai
from google.genai import types
load_dotenv()
api_key = os.getenv("GOOGLE_API_KEY") or os.getenv("GEMINI_API_KEY")
if not api_key:
raise ValueError(
"API key not found. Add GOOGLE_API_KEY or GEMINI_API_KEY to your .env file."
)
client = genai.Client(api_key=api_key)
MODEL_NAME = "gemini-2.5-flash"
query = "What are the benefits of RAG?"
retrieved_contexts = [
"RAG allows an LLM to use external knowledge from a retrieval system.",
"Convolutional neural networks are commonly used for image classification.",
"RAG can reduce hallucinations by grounding an LLM's response in retrieved documents.",
"Transformers use self-attention mechanisms to process relationships between tokens.",
"RAG allows knowledge to be updated by changing the external knowledge base "
"without retraining the LLM."
]
def evaluate_context(query, context):
prompt = f"""
You are an expert RAG evaluation system.
Determine whether the retrieved context is relevant and useful
for answering the user's query.
User Query:
{query}
Retrieved Context:
{context}
Return JSON with exactly one field:
{{
"relevant": "1"
}}
Use "1" if the context is relevant and useful.
Use "0" if the context is irrelevant or not useful.
Do not return explanations.
"""
response = client.models.generate_content(
model=MODEL_NAME,
contents=prompt,
config=types.GenerateContentConfig(
temperature=0,
response_mime_type="application/json",
response_schema={
"type": "OBJECT",
"properties": {
"relevant": {
"type": "STRING",
"enum": ["0", "1"]
}
},
"required": ["relevant"]
}
)
)
response_text = (response.text or "").strip()
if not response_text:
print("Gemini returned an empty response.")
return 0
try:
result = json.loads(response_text)
return int(result.get("relevant", "0"))
except (json.JSONDecodeError, TypeError, ValueError) as error:
print("Invalid Gemini response:")
print(response_text)
print("Error:", error)
return 0
relevance = []
for rank, context in enumerate(retrieved_contexts, start=1):
verdict = evaluate_context(query, context)
relevance.append(verdict)
print(
f"Rank {rank}: "
f"{'Relevant' if verdict == 1 else 'Irrelevant'}"
)
def context_precision(relevance):
relevant_count = 0
precision_sum = 0.0
for rank, is_relevant in enumerate(relevance, start=1):
if is_relevant == 1:
relevant_count += 1
precision_sum += relevant_count / rank
if relevant_count == 0:
return 0.0
return precision_sum / relevant_count
score = context_precision(relevance)
print("\nRelevance:")
print(relevance)
print(f"\nContext Precision: {score:.3f}")
Rank 1: Relevant
Rank 2: Irrelevant
Rank 3: Relevant
Rank 4: Irrelevant
Rank 5: Relevant
Relevance:
[1, 0, 1, 0, 1]
Context Precision: 0.756
3. Noise Sensitivity (Generator Robustness Metric) #
Noise Sensitivity evaluates how robust the generation LLM is when presented with irrelevant or distracting information (“noise”) in the retrieved context.
- Target Component: Generator (LLM)
- Core Focus: Noise Robustness
- Core Question: Is the LLM smart enough to discard irrelevant context chunks, or does it incorporate false/unrelated claims into its final answer?
- How It Works:
- RAGAS separates the retrieved context into useful facts (matching the ground truth) and noisy facts (irrelevant to the query).
- The generated response is broken down into factual claims.
- RAGAS counts how many claims in the final response were pulled from the noisy context chunks.
- Ideal Score: 0.0 (LOWER IS BETTER). A score of 0.0 means the LLM completely ignored the noise. A score of 0.5 means half of the generated response consists of irrelevant noise.
Formula #
For example:
Total claims = 4
Noise-induced incorrect claims = 1
Noise Sensitivity = 1 / 4 = 0.25
A score of 0.0 means the LLM completely ignored the noise.
Gemini-Based Implementation #
Gemini can act as the LLM-as-a-Judge to identify claims influenced by irrelevant context.
import json
import os
from dotenv import load_dotenv
from google import genai
from google.genai import types
load_dotenv()
api_key = os.getenv("GOOGLE_API_KEY") or os.getenv("GEMINI_API_KEY")
if not api_key:
raise ValueError("Gemini API key not found.")
client = genai.Client(api_key=api_key)
MODEL_NAME = "gemini-2.5-flash"
query = "What are the benefits of RAG?"
reference = """
RAG provides external knowledge to the LLM.
RAG can reduce hallucinations by grounding responses in retrieved documents.
RAG allows knowledge to be updated without retraining the LLM.
"""
retrieved_contexts = [
"RAG allows an LLM to use external knowledge from a retrieval system.",
"Convolutional neural networks are used for image classification.",
"RAG can reduce hallucinations by grounding responses in retrieved documents.",
"Transformers use self-attention mechanisms.",
"RAG allows knowledge to be updated without retraining the LLM."
]
generated_response = """
RAG provides external knowledge, reduces hallucinations,
and allows knowledge updates without retraining the LLM.
Transformers use self-attention mechanisms to process
relationships between tokens.
"""
def evaluate_noise(query, reference, contexts, generated_response):
prompt = f"""
You are an expert RAG evaluator.
Query:
{query}
Reference Answer:
{reference}
Retrieved Contexts:
{json.dumps(contexts, indent=2)}
Generated Response:
{generated_response}
Count:
1. The total factual claims in the generated response.
2. Claims caused by irrelevant or noisy context.
Return JSON only.
"""
response = client.models.generate_content(
model=MODEL_NAME,
contents=prompt,
config=types.GenerateContentConfig(
temperature=0,
response_mime_type="application/json",
response_schema={
"type": "OBJECT",
"properties": {
"total_claims": {"type": "INTEGER"},
"noise_induced_claims": {"type": "INTEGER"}
},
"required": [
"total_claims",
"noise_induced_claims"
]
}
)
)
if not response.text:
raise ValueError("Gemini returned an empty response.")
return json.loads(response.text)
result = evaluate_noise(
query,
reference,
retrieved_contexts,
generated_response
)
total_claims = result["total_claims"]
noise_induced_claims = result["noise_induced_claims"]
noise_sensitivity = (
noise_induced_claims / total_claims
if total_claims > 0
else 0.0
)
print("Evaluation:", result)
print(f"Noise Sensitivity: {noise_sensitivity:.3f}")
Evaluation: {'total_claims': 4, 'noise_induced_claims': 1}
Noise Sensitivity: 0.250
4. Response Relevancy (Generator Metric) #
Response Relevancy (also known as Answer Relevancy) measures whether the generated response directly addresses the user’s question, regardless of factual accuracy. It catches answers that are off-topic, incomplete, or evasive.
- Target Component: Generator (LLM)
- Core Focus: Topic Utility & Directness
- Core Question: Does the generated response directly answer what was asked?
- How It Works (Reverse-Engineering Approach):
- RAGAS passes the Generated Response to an LLM and instructs it to reverse-generate hypothetical questions that the response would answer.
- An embedding model generates vector representations for both the hypothetical questions and the original User Query.
- RAGAS calculates the average cosine similarity between the embeddings of the generated questions and the original query.
- Ideal Score: 1.0 (Higher is better). If the response goes off-topic or adds unnecessary fluff, the similarity score drops.
5. Faithfulness (Generator Metric) #
Faithfulness is the primary hallucination detection metric in RAGAS. It checks whether every factual claim made in the generated response can be directly traced back to and verified by the retrieved context.
- Target Component: Generator (LLM)
- Core Focus: Grounding & Hallucination Prevention
- Core Question: Is the response strictly grounded in the provided context without inventing outside facts?
- How It Works:
- An LLM extracts all individual factual claims from the Generated Response.
- Each claim is verified against the Retrieved Context to mark it as supported (True) or unverified (False).
- The score is the ratio of supported claims to total claims in the response.
- Ideal Score: 1.0 (Higher is better). A score of 1.0 indicates zero hallucinations.
How RAGAS Works: The Step-by-Step Evaluation Process #
Evaluating a RAG system with RAGAS follows a clear 5-step workflow:
- Step 1: Benchmark Dataset Curation: Prepare a set of test queries paired with human-written reference answers (ground truth).
- Step 2: RAG Pipeline Execution: Run the test queries through your RAG system to capture the Retrieved Contexts and final Generated Responses.
- Step 3: Asynchronous Evaluation Calls: RAGAS sends the collected data to an evaluator LLM using asynchronous API calls to extract claims, verify context matches, and calculate embeddings.
- Step 4: Score Calculation: RAGAS aggregates normalized scores (0.0 to 1.0) for each selected metric across all test samples.
- Step 5: Diagnostics & Optimization: Convert the results into a structured format (such as a pandas DataFrame or CSV) to identify weaknesses and adjust chunk sizes, top-k limits, or prompt templates.
Hands-On Python Implementation: Evaluating Your RAG Pipeline #
Let’s look at how to implement an automated evaluation pipeline using Python.
To keep your code clean and production-ready, structure your project into three modular files:
rag_pipeline.py: Defines document loading, vector storage, and the RAG generation chain.evaluate.py: Prepares the RAGAS dataset and executes the metrics scoring engine.main.py: The main entry point that coordinates building and evaluating the pipeline.
Step 1: Build the RAG Pipeline (rag_pipeline.py) #
First, set up your standard RAG chain using LangChain, ChromaDB, and OpenAI embeddings. This file returns both your generation chain and the retriever object.
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings, ChatOpenAI
from langchain_community.vectorstores import Chroma
def build_rag_chain(pdf_path=""):
# 1. Load and split documents into chunks
loader = PyPDFLoader(pdf_path)
docs = loader.load()
text_splitter = RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=150)
chunks = text_splitter.split_documents(docs)
# 2. Create vector database and retriever
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vectorstore = Chroma.from_documents(chunks, embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 4})
# 3. Initialize generation LLM
rag_chain = ChatOpenAI(model="gpt-4o-mini", temperature=0)
return rag_chain, retriever
Step 2: Create the RAGAS Evaluation Script (evaluate.py) #
Next, set up the evaluation engine. This script takes test queries, retrieves context chunks, generates responses, pairs them with human ground truth reference answers, and computes RAGAS scores.
import asyncio
import pandas as pd
from openai import AsyncOpenAI
from ragas import evaluate
from ragas.llms import llm_factory
from ragas.embeddings import embedding_factory
from ragas.dataset_schema import EvaluationDataset
from ragas.metrics.collections import (
ContextRecall,
ContextPrecision,
Faithfulness,
AnswerRelevancy,
NoiseSensitivity,
)
# Human-curated ground truth benchmark dataset
BENCHMARK_QA = [
(
"What are the three pillars of sustainable development?",
"The three pillars of sustainable development are economic growth, social inclusion, and environmental protection."
),
(
"What is the Paris Agreement temperature target?",
"The Paris Agreement aims to limit global warming to well below 2 degrees Celsius above pre-industrial levels."
),
]
async def run_evaluation(rag_chain, retriever):
# Initialize asynchronous OpenAI client for evaluator LLM
client = AsyncOpenAI()
eval_llm = llm_factory(model_name="gpt-4o-mini", client=client)
eval_embeddings = embedding_factory(provider="openai", model_name="text-embedding-3-small", client=client)
dataset_records = []
# Run queries through RAG pipeline to capture context and responses
for question, reference_answer in BENCHMARK_QA:
# Retrieve context chunks
docs = retriever.invoke(question)
contexts = [doc.page_content for doc in docs]
# Generate LLM response
response_obj = rag_chain.invoke(question)
response_text = response_obj.content if hasattr(response_obj, "content") else str(response_obj)
# Build sample record for RAGAS
dataset_records.append({
"user_input": question,
"retrieved_contexts": contexts,
"response": response_text,
"reference": reference_answer
})
# Convert records to RAGAS EvaluationDataset
eval_dataset = EvaluationDataset.from_list(dataset_records)
# Instantiate evaluation metrics
metrics = [
ContextRecall(llm=eval_llm),
ContextPrecision(llm=eval_llm),
Faithfulness(llm=eval_llm),
AnswerRelevancy(llm=eval_llm, embeddings=eval_embeddings),
NoiseSensitivity(llm=eval_llm),
]
# Execute evaluation and save results to CSV
results = evaluate(dataset=eval_dataset, metrics=metrics)
df_results = results.to_pandas()
df_results.to_csv("ragas_evaluation_results.csv", index=False)
print("Evaluation completed! Results saved to 'ragas_evaluation_results.csv'.")
Step 3: Run the Complete Pipeline (main.py) #
Finally, create a main driver script to execute building and evaluating your system with a single command.
import asyncio
from rag_pipeline import build_rag_chain
from evaluate import run_evaluation
if __name__ == "__main__":
print("Building RAG chain and initializing vector database...")
rag_chain, retriever = build_rag_chain("book.pdf")
print("Launching RAGAS evaluation suite...")
asyncio.run(run_evaluation(rag_chain, retriever))
Running this script produces a CSV file containing metric scores for every test query, giving you clear data to optimize your chunk sizes, retrieval limits, and prompt templates.
Real-World Examples #
To better understand these metrics, let me walk through four practical scenarios.
Example 1: Context Recall (Type 2 Diabetes Symptoms & Causes)
- User Query: “What are the causes and symptoms of Type 2 Diabetes?”
- Ground Truth Reference: “Type 2 Diabetes is caused by insulin resistance, obesity, and a sedentary lifestyle. Symptoms include frequent urination, excessive thirst, fatigue, blurred vision, and slow-healing sores.”
- Retrieved Context: Contains chunks describing frequent urination, excessive thirst, fatigue, and blurred vision, but misses insulin resistance and lifestyle causes.
- Evaluation: Out of 8 key facts in the reference answer, the retriever only brought chunks covering 4 facts.
- Context Recall Score: 0.50 (50% coverage).
Example 2: Context Precision & Ranking (Eiffel Tower Details)
- User Query: “Where is the Eiffel Tower located and how tall is it?”
- Top-4 Retrieved Chunks:
- Chunk 1: The Eiffel Tower is located in Paris, France. (Relevant)
- Chunk 2: The Louvre Museum attracts 9 million visitors per year. (Irrelevant)
- Chunk 3: The Eiffel Tower stands 330 meters tall. (Relevant)
- Chunk 4: The Arc de Triomphe is 50 meters tall. (Irrelevant)
- Evaluation: The retriever fetched the right information, but placed an irrelevant chunk about the Louvre at Rank 2 ahead of the height information at Rank 3.
- Context Precision Score: 0.75 (Penalized because irrelevant information was ranked higher than relevant context).
Example 3: Faithfulness & Hallucination (Albert Einstein Biography)
- Retrieved Context: “Albert Einstein was born on March 14, 1879, in Germany. He developed the theory of relativity and won the Nobel Prize in Physics in 1921.”
- Generated Response: “Albert Einstein was born on March 14, 1879, in Germany. He developed the theory of relativity, won the Nobel Prize in Physics in 1921, worked extensively on quantum mechanics, and had an IQ of 160.”
- Evaluation:
- Fact 1 (Born March 14, 1879) -> Supported by context (True)
- Fact 2 (Theory of relativity) -> Supported by context (True)
- Fact 3 (Nobel Prize 1921) -> Supported by context (True)
- Fact 4 (Quantum mechanics) -> Not in context (False)
- Fact 5 (IQ of 160) -> Not in context (False)
- Faithfulness Score: 3 / 5 = 0.60 (The LLM hallucinated facts from its pre-training memory instead of relying strictly on retrieved context).
Example 4: Response Relevancy (Boiling Point of Water)
- User Query: “What is the boiling point of water?”
- Generated Response: “Water is a vital resource found across the earth. It covers 71% of the planet’s surface and plays a central role in regulating global climate patterns.”
- Evaluation: While the response contains true scientific statements about water, it completely fails to answer the user’s specific question about boiling point.
- Response Relevancy Score: 0.20 (Significantly penalized for going off-topic).
Comparison Table of RAGAS Metrics
| Metric Name | Target Component | Core Focus | Ideal Score | Primary Goal |
|---|---|---|---|---|
| Context Recall | Retriever | Completeness | 1.0 (High) | Ensure no key facts are missed during retrieval |
| Context Precision | Retriever | Ranking & Relevance | 1.0 (High) | Ensure relevant chunks are ranked at the top |
| Noise Sensitivity | Generator (LLM) | Noise Robustness | 0.0 (Low) | Prevent LLM from including context noise in output |
| Response Relevancy | Generator (LLM) | Topic Utility | 1.0 (High) | Ensure output directly answers the query |
| Faithfulness | Generator (LLM) | Grounding & Truth | 1.0 (High) | Detect and eliminate LLM hallucinations |
Advantages and Limitations of RAGAS #
Advantages
- Component-Level Diagnostics: Pinpoints whether errors stem from the retriever or the generator.
- Semantic Understanding: Uses LLM-as-a-Judge reasoning rather than rigid word-for-word string matching.
- Automated & Scalable: Enables continuous integration testing for AI applications without manual inspection.
- Actionable Tuning: Helps measure the exact impact of changing parameters like chunk size, overlap, or top-k retrieval values.
Limitations
- API Cost and Latency: Running multiple evaluator LLM calls increases API token costs and execution time.
- Ground Truth Dependency: Metrics like Context Recall and Context Precision require well-curated reference answer datasets.
- Judge LLM Bias: The quality of evaluation depends on the reasoning capabilities of the judge LLM being used.
Real-World Applications #
RAGAS evaluation is widely used across production AI domains:
- Enterprise Document Q&A: Benchmark internal search assistants over financial reports, HR policies, and technical documentation.
- Legal & Compliance Search: Ensure legal contract summarizers maintain 100% Faithfulness to avoid incorrect interpretations.
- Medical Knowledge Retrieval: Verify high Context Recall so clinical search engines never omit essential medical context.
- Customer Support Chatbots: Test Response Relevancy to ensure bots deliver direct answers without fluff.
Important Points for Revision #
- RAG Evaluation Scope: Focuses primarily on evaluating the Retrieval Phase and Generation Phase.
- Retriever Metrics: Context Recall measures completeness; Context Precision measures ranking accuracy.
- Generator Metrics: Faithfulness detects hallucinations; Response Relevancy detects off-topic answers.
- Noise Sensitivity Metric: Measures LLM robustness to irrelevant context chunks—lower scores are better.
- Core Philosophy: Quality over quantity. It is better to systematically optimize 3–4 core metrics aligned with your goals than to use dozens of confusing benchmarks.
Technical Questions and Answers #
Q1: What is the difference between Context Recall and Context Precision?
Answer: Context Recall evaluates whether all required information was retrieved from the knowledge base (completeness). Context Precision evaluates whether the retrieved relevant chunks are ranked at the top of the context list rather than buried under irrelevant noise.
Q2: Why is a lower score better for Noise Sensitivity?
Answer: Noise Sensitivity measures how many claims in the final response were pulled from irrelevant context chunks. A score of 0.0 means the LLM successfully ignored all noise, whereas a higher score means the LLM was easily distracted.
Q3: Does Response Relevancy check for factual correctness?
Answer: No. Response Relevancy only evaluates whether the generated answer directly addresses the user’s query topic. Factual accuracy against context is evaluated separately by the Faithfulness metric.
Q4: Why should I use LLM-as-a-Judge instead of BLEU or ROUGE?
Answer: BLEU and ROUGE rely on exact keyword matches. Because LLMs express the same semantic meaning using different vocabulary and sentence structures, exact string matching often marks correct answers as failures. LLM-as-a-Judge evaluates semantic intent much like a human evaluator.
Quick Revision Summary #
The RAGAS framework provides a data-driven approach to evaluating Retrieval-Augmented Generation applications. By decoupling the pipeline into Retrieval and Generation stages, RAGAS allows developers to systematically diagnose errors:
- Use Noise Sensitivity to test LLM robustness against distracting context..
- Use Context Recall to verify context completeness.
- Use Context Precision to optimize re-ranking and chunk order.
- Use Faithfulness to eliminate LLM hallucinations.
- Use Response Relevancy to keep answers direct and concise.
RAGAS Quiz #
1. What does the acronym RAGAS stand for in RAG system evaluation?
Retrieval-Augmented Generation Automated Scoring
RAG Assessment
Robust AI Generation Analysis System
Recurrent Agentic Guidance and Assessment Standard
Explanation
RAGAS stands for RAG Assessment, an evaluation framework used to measure and score the performance of RAG systems.
2. Which two primary components of a RAG pipeline are targeted for evaluation in RAGAS?
Vector Database and Chunking Strategy
Embedding Model and Tokenizer
Retriever and Generator LLM
Document Loader and User Interface
Explanation
RAGAS specifically targets the Retriever and the Generator LLM as the two primary components to evaluate in a RAG pipeline.
3. Why does RAGAS adopt an 'LLM-as-a-Judge' approach rather than using traditional metrics like BLEU or ROUGE?
Traditional metrics are computationally too slow for real-time evaluations.
LLM-as-a-Judge uses semantic understanding and intent matching rather than rigid keyword matching.
BLEU and ROUGE require external web search access.
LLM-as-a-Judge operates without needing any reference answer or context.
Explanation
RAGAS uses an LLM as a judge because RAG evaluations rely on semantic meaning and intent rather than exact word-for-word string matching.
4. What does the Context Recall metric evaluate in RAGAS?
Whether the generated response contains hallucinations.
Whether the retriever fetched all necessary information from the knowledge base to answer the query.
How fast the vector database returns search results.
How well the LLM ignores irrelevant context noise.
Explanation
Context Recall evaluates the completeness of the retrieved context by checking if all required facts from the ground truth reference were fetched by the retriever.
5. How is the Context Recall score calculated in RAGAS?
Ratio of retrieved claims supported by the reference divided by total claims in the response.
Ratio of reference claims present in the retrieved context divided by total claims in the reference.
Average cosine similarity between original query and retrieved context.
Number of noisy chunks divided by total retrieved chunks.
Explanation
Context Recall is calculated as the number of claims in the reference answer supported by the retrieved context divided by the total number of claims in the reference answer.
6. Which metric evaluates both the relevance AND the ranking order of retrieved chunks?
Context Precision
Response Relevancy
Faithfulness
Noise Sensitivity
Explanation
Context Precision evaluates whether relevant chunks are retrieved and whether they are ranked higher in the retrieved context list.
7. In RAGAS, how does the Noise Sensitivity metric differ in its score interpretation compared to other metrics?
A higher score indicates better LLM performance.
A lower score is better, where 0 indicates zero noise included in the response.
It ranges from -1 to +1 instead of 0 to 1.
It only outputs discrete pass or fail labels.
Explanation
Unlike most RAGAS metrics where higher is better, Noise Sensitivity is better when lower; a score of 0 means the generated response contains no noisy or irrelevant claims.
8. What is the formula used to calculate Noise Sensitivity?
Number of claims in reference / Total retrieved chunks
Number of incorrect/noisy claims in response / Total number of claims in response
Number of true claims in context / Total claims in reference
Cosine similarity of generated queries / Original query embedding
Explanation
Noise Sensitivity is calculated by taking the number of incorrect/noisy claims in the response and dividing it by the total number of claims in the response.
9. Which metric serves as the primary hallucination detection metric in RAGAS?
Context Precision
Faithfulness
Context Recall
Response Relevancy
Explanation
Faithfulness is the primary hallucination detection metric in RAGAS, checking whether every claim in the generated response is grounded in the retrieved context.
10. How is the Faithfulness metric calculated?
Number of claims in response supported by retrieved context divided by total claims in response.
Total retrieved chunks divided by total claims in reference answer.
Cosine similarity between response embeddings and query embeddings.
Number of noisy chunks in context divided by total chunks.
Explanation
Faithfulness is calculated as the number of claims in the generated response that are supported by the retrieved context divided by the total number of claims in the response.
11. What smart technique does RAGAS use to calculate Response Relevancy?
It compares keyword frequencies between prompt and answer.
It reverse-engineers hypothetical questions from the response and measures cosine similarity with the user query.
It checks if the response matches a hardcoded human answer string.
It counts the number of citations inside the response.
Explanation
Response Relevancy reverse-engineers hypothetical questions from the generated response using an LLM, then measures average cosine similarity between their embeddings and the original user query embedding.
12. Does the Response Relevancy metric evaluate the factual correctness of the answer?
Yes, it verifies every fact against the ground truth answer.
No, it only evaluates whether the response directly addresses the query topic, regardless of factual accuracy.
Yes, but only for numerical claims and dates.
No, it only measures retrieval speed.
Explanation
Response Relevancy checks if the generated response is on-topic and directly addresses what was asked, without evaluating factual correctness.
13. In RAGAS terminology, what is a 'Sample'?
A full database dump of all vector embeddings.
A single test entry consisting of a test query, its context, and outputs used for evaluation.
A sub-segment of an LLM prompt template.
The execution time log of a vector search call.
Explanation
A sample in RAGAS refers to a single test query or test entry used to evaluate the RAG pipeline.
14. In RAGAS terminology, what defines an 'Experiment'?
Deleting and rebuilding the vector index from scratch.
A complete RAG pipeline setup where a single hyperparameter or knob is changed for comparison.
Running an LLM without any prompt template.
Testing the RAG system with zero internet access.
Explanation
An experiment in RAGAS represents a RAG pipeline run where a specific parameter or knob (e.g. chunk size or top-k) is tweaked to observe performance changes.
15. According to the core evaluation philosophy of RAGAS, what is the best strategy when choosing evaluation metrics?
Use as many metrics as possible (10+) to create a complex evaluation matrix.
Prioritize quality over quantity by selecting a focused set of 3 to 4 metrics suitable for your pipeline.
Rely solely on non-LLM exact string matching metrics.
Evaluate only the document ingestion phase while skipping the retrieval phase.
Explanation
RAGAS emphasizes quality over quantity, advocating for a focused set of 3 to 4 relevant metrics that clearly evaluate and guide the optimization of your RAG pipeline.