Retrieval-Augmented Generation (RAG) is one of the most important architectures used in modern Generative AI applications. It combines information retrieval with Large Language Models (LLMs) so that an AI system can retrieve relevant information from external knowledge sources before generating an answer.
RAG is widely used in AI chatbots, document question-answering systems, enterprise search, customer support, knowledge assistants and AI agents. Understanding retrieval is especially important because the quality of a RAG system depends heavily on how accurately relevant information is retrieved.
1. Explain a RAG Pipeline End-to-End #
Concept #
RAG stands for Retrieval-Augmented Generation. It is an architecture in which an application first retrieves relevant information from an external knowledge base and then provides that information to an LLM as context for generating the final answer.
Instead of asking an LLM to answer only from its pretrained knowledge, RAG
allows the model to use information from documents, databases, websites,
PDFs, company knowledge bases or other external sources.
Simple Example #
Imagine a company has 10,000 internal documents containing information about employees, products, policies and technical documentation.
An employee asks:
What is the company's work-from-home policy?
A normal LLM may not know the company’s private policy. A RAG system searches the company’s documents, finds the relevant policy section and sends that information to the LLM.
The LLM then generates an answer based on the retrieved information.
End-to-End RAG Pipeline #
Step 1: Document Loading #
The first step is to collect the knowledge that the application should use. The documents may come from PDFs, Word files, websites, databases, APIs, Google Drive, cloud storage or internal company systems.
documents = load_documents("company_documents/")
Step 2: Chunking #
Large documents are divided into smaller pieces called
chunks. Instead of embedding an entire 100-page document as one vector, we divide it into meaningful sections.
chunks = text_splitter.split_documents(documents)
Step 3: Embedding #
Each chunk is converted into a numerical vector using an embedding model.
The vector represents the semantic characteristics of the text.
embeddings = embedding_model.embed_documents(chunks)
Step 4: Store in Vector Database #
The vectors and their associated text are stored in a vector database such
as FAISS, Qdrant, Pinecone, Weaviate or Milvus.
Step 5: Convert User Query into an Embedding #
When a user asks a question, the question is converted into an embedding using the same embedding space.
query_vector = embedding_model.embed_query(
"What is the work from home policy?"
)
Step 6: Similarity Search #
The query vector is compared with stored document vectors. The system retrieves the chunks that are semantically closest to the query.
Step 7: Reranking #
In advanced RAG systems, retrieved documents can be passed through a reranker. The reranker evaluates the relationship between the query and retrieved documents and places the most relevant documents first.
Step 8: Context Construction #
The selected chunks are combined with the user’s question to create the
context given to the LLM.
Step 9: Generation #
The LLM receives the question and retrieved context and generates the final
answer.
User Question
+
Retrieved Context
↓
LLM
↓
Grounded Answer
Interview Answer #
A RAG pipeline has two major stages: indexing and query-time retrieval.
During indexing, documents are loaded, cleaned, split into meaningful chunks,
converted into embeddings and stored in a vector database along with useful
metadata. When a user asks a question, the query is converted into an
embedding and used to search the vector database for relevant chunks. The
retrieved documents can then be filtered or reranked before being combined
into context. Finally, the context and the original question are provided to
the LLM, which generates a grounded response.
2. What Happens During the Indexing Phase of RAG? #
Concept #
The indexing phase is the offline or preparation phase of a
RAG system. Its purpose is to transform raw documents into searchable
representations.
This process normally happens before the user sends a query.
1. Load Documents #
Documents are loaded from different sources.
PDF
DOCX
HTML
Web Pages
Database
CSV
Markdown
Company Knowledge Base
2. Extract Text #
If the source is a PDF, the system extracts the text from the PDF. If it is a
web page, the relevant content is extracted from HTML.
3. Clean the Text #
Unnecessary information such as repeated headers, excessive whitespace,
navigation text and corrupted characters may be removed.
4. Split Documents into Chunks #
Large documents are divided into smaller pieces so that retrieval can return
specific relevant sections rather than entire documents.
5. Add Metadata #
Metadata can include information such as document ID, title, author,
department, page number, date or category.
metadata = {
"source": "employee_handbook.pdf",
"page": 25,
"category": "HR"
}
6. Generate Embeddings #
Each chunk is converted into a numerical vector.
7. Store in Vector Database #
The vector, original text and metadata are stored together so that the
application can retrieve the actual content later.
Important Point #
The indexing phase is generally performed offline. However, it may need to be repeated when documents are added, updated or deleted.
Interview Answer #
During indexing, the RAG system prepares the knowledge base for retrieval. It loads documents from sources such as PDFs, websites or databases, extracts and cleans the text, divides the content into meaningful chunks, attaches metadata, generates an embedding for each chunk and stores the embeddings with the original content in a vector database. This preparation allows the system to perform fast and relevant retrieval when a user submits a query.
3. What Happens During the Retrieval / Query Phase? #
Concept #
The retrieval phase happens when a user asks a question. The system attempts to find the most relevant information from the indexed knowledge base.
Example #
User:
What is the company's annual leave policy?
The system converts this query into an embedding and searches the vector database for chunks that have similar semantic meaning.
Suppose the database returns:
Chunk 1 → Annual leave policy
Chunk 2 → Sick leave policy
Chunk 3 → Employee attendance
Chunk 4 → Holiday calendar
Chunk 5 → Work from home policy
The system can then rank these chunks and select the most relevant ones.
Retrieval Pipeline #
User Query
↓
Query Embedding
↓
Vector Search
↓
Top-K Documents
↓
Filtering / Reranking
↓
Relevant Context
↓
LLM
↓
Final Answer
Why Retrieval Is Important #
Even if the LLM is powerful, poor retrieval can produce a poor answer. If the correct information is not retrieved, the LLM may not have the necessary context to answer accurately.
Interview Answer #
During query time, the user’s question is first processed and converted into an embedding. The system then searches the vector database for semantically similar document chunks. The retrieved candidates can be filtered using metadata and reranked to identify the most relevant information. The selected chunks are assembled into context and passed to the LLM together with the original question. The LLM then uses that context to generate the final grounded response.
4. How Do You Decide the Right Chunk Size and Chunk Overlap? #
Concept #
Chunking determines how a large document is divided into smaller pieces. Choosing the correct chunk size is important because chunks that are too small may lose context, while chunks that are too large may contain unnecessary information.
Chunk Size Too Small #
Suppose we split the following sentence incorrectly:
Chunk 1:
The refund policy allows customers to request a
Chunk 2:
refund within 30 days of purchase.
The important meaning is separated across chunks. Retrieval may return only one chunk, resulting in incomplete context.
Chunk Size Too Large #
If every chunk contains several thousand tokens, the retrieved chunk may contain a large amount of irrelevant information. This can increase cost and reduce retrieval precision.
What Is Chunk Overlap? #
Chunk overlap means repeating a small portion of one chunk in the next chunk. It helps preserve context at chunk boundaries.
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=50
)
How to Choose Chunk Size #
There is no universally optimal chunk size. It depends on the document type, retrieval task, embedding model, query type and downstream LLM context window.
Practical Starting Point #
Start with:
Chunk Size → 300–800 tokens
Overlap → 10–20%
Then evaluate retrieval quality.
Interview Answer #
I would not select chunk size based on a fixed universal rule. I would consider the document structure, semantic boundaries, query complexity,embedding model, expected retrieval behavior and the context window of theLLM. As a starting point, I might test a moderate chunk size with a small overlap, but the final values should be determined through evaluation on realistic queries. I would compare retrieval metrics and inspect whether important information is being split across chunks or whether retrieved chunks contain too much irrelevant content.
5. Compare Fixed-Size, Recursive, Semantic and Parent-Child Chunking #
1. Fixed-Size Chunking #
The document is divided into chunks using a fixed number of characters or
tokens.
Document
↓
500 tokens
↓
500 tokens
↓
500 tokens
↓
500 tokens
It is simple and predictable but may split sentences or concepts in the
middle.
2. Recursive Chunking #
Recursive chunking attempts to split text using increasingly smaller separators while trying to preserve meaningful structure.
For general-purpose text RAG, recursive splitting is often a strong baseline.
3. Semantic Chunking #
Semantic chunking attempts to group text based on meaning rather than simply using character or token boundaries.
This can be useful when topic boundaries are more important than fixed lengths, although it may add computational complexity.
4. Parent-Child Chunking #
Parent-child chunking maintains a relationship between larger parent documents or sections and smaller child chunks.
The smaller child chunks can be used for precise retrieval, while the larger parent section can provide additional context to the LLM.
Comparison #
Fixed Size
→ Simple and predictable
Recursive
→ Preserves common text boundaries
Semantic
→ Groups content by meaning
Parent-Child
→ Precise retrieval + larger contextual information
Interview Answer #
Fixed-size chunking divides content according to a predefined character or token limit, making it simple and predictable but potentially breaking semantic boundaries. Recursive chunking uses separators such as paragraphs,
sentences and smaller text units to preserve structure and is a useful general-purpose approach. Semantic chunking uses meaning or topic changes to determine boundaries, while parent-child chunking uses smaller child chunks for precise retrieval while retaining access to a larger parent section for additional context.
6. Your RAG Application Is Returning Irrelevant Documents. How Would You Debug It? #
Concept #
Poor retrieval is one of the most common problems in RAG systems. I would debug the system systematically instead of immediately changing the LLM.
Step 1: Inspect the User Query #
The query itself may be ambiguous, too short or poorly formulated.
Step 2: Inspect Retrieved Chunks #
Before changing anything, print the actual retrieved documents and their
similarity scores.
results = vector_store.similarity_search_with_score(
query,
k=5
)
for document, score in results:
print("Score:", score)
print(document.page_content)
Step 3: Check Chunking #
If chunks are too large, too small or split important information, retrieval quality can suffer.
Step 4: Check Embedding Model #
The embedding model needs to represent the semantics of the application’s domain effectively.
Step 5: Check Similarity Metric #
Verify whether the vector database is configured correctly for cosine similarity, dot product or Euclidean distance.
Step 6: Check Metadata Filtering #
Metadata filters can prevent irrelevant documents from unrelated departments, products, dates or document types from being retrieved.
Step 7: Tune Top-K #
Retrieving too few documents can miss useful information. Retrieving too many can introduce noise.
Step 8: Add Reranking #
A reranker can improve the ordering of retrieved candidates by evaluating query-document relevance more directly.
Step 9: Evaluate Retrieval Separately #
Do not evaluate only the final LLM answer. First evaluate whether the correct document was retrieved.
Question
↓
Did retrieval find the correct document?
↓
YES → Evaluate generation
NO → Debug retrieval
Interview Answer #
I would debug irrelevant retrieval systematically rather than immediately changing the LLM. First, I would inspect the original query and the actual retrieved chunks along with their similarity scores. Then I would check the
chunking strategy, embedding model, similarity metric, metadata filters and top-K value. If the initial retrieval produces a reasonable candidate set, I would consider adding a reranker. Finally, I would create a representative evaluation dataset and measure retrieval separately using metrics such as Recall@K, Precision@K, MRR and NDCG. This helps determine whether the problem is in retrieval or in the generation stage.
7. What Is an Embedding? How Does an Embedding Model Represent Text? #
Concept #
An embedding is a numerical representation of data such as text, images or
audio. In RAG, text embeddings represent the semantic characteristics of
text as vectors in a high-dimensional mathematical space.
Simple Example #
"How can I reset my password?"
↓
Embedding Model
↓
[0.12, -0.43, 0.81, 0.27, ...]
The actual embedding may contain hundreds or thousands of dimensions,
depending on the model.
Semantic Similarity #
Consider these two sentences:
Sentence A:
How do I reset my password?
Sentence B:
I forgot my password. How can I create a new one?
Although the words are not exactly the same, their meanings are similar.
A good embedding model places their vectors relatively close together.
Example in Python #
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
text = "How do I reset my password?"
vector = model.encode(text)
print(vector)
print(len(vector))
Interview Answer #
An embedding is a dense numerical vector representation of data. In text retrieval, an embedding model converts text into a vector space where semantically related texts tend to have similar representations. For example,questions about resetting a password can be represented by vectors that are close to other texts discussing password recovery. These vectors can then be compared using measures such as cosine similarity or dot product to retrieve relevant information.
8. Explain Cosine Similarity, Dot Product and Euclidean Distance #
Why Do We Need Similarity Measures? #
After converting text into vectors, we need a mathematical method to compare those vectors. Similarity or distance metrics determine how close two vectors are.
1. Cosine Similarity #
Cosine similarity measures the angle between two vectors rather than their
absolute magnitude.
cosine_similarity(A, B)
=
(A · B) / (||A|| ||B||)
A value closer to 1 generally indicates that the vectors point in similar
directions, while a value closer to -1 indicates opposite directions.
2. Dot Product #
The dot product is calculated by multiplying corresponding dimensions and
summing the results.
A · B = A1B1 + A2B2 + A3B3 + ...
Dot product similarity is commonly used in vector search systems. Its
interpretation depends on how the vectors are normalized.
3. Euclidean Distance #
Euclidean distance measures the straight-line distance between two points in
vector space.
d(A,B) =
√((A1-B1)² + (A2-B2)² + ...)
Visual Comparison #
Comparison #
Cosine Similarity
→ Focuses on vector direction
Dot Product
→ Measures vector interaction and magnitude
Euclidean Distance
→ Measures geometric distance
Interview Answer #
Cosine similarity measures the angle between two vectors and is commonly used to compare semantic representations. Dot product calculates the sum of element-wise products and can be used as a similarity measure, particularly when vectors are normalized. Euclidean distance measures the geometric distance between two vectors. In a RAG system, the appropriate metric depends on the embedding model, vector normalization and the configuration supported by the vector database.
9. What Is a Vector Database? How Does Vector Similarity Search Work? #
Concept #
A vector database is a system designed to store, index and search numerical
vectors efficiently.
In RAG, document chunks are converted into embeddings and stored in a vector database along with their original content and metadata.
Examples #
FAISS
Qdrant
Pinecone
Weaviate
Milvus
Chroma
How Similarity Search Works #
Example #
Query:
"How many annual leave days do employees receive?"
↓
Query Vector
↓
Vector Database
↓
Chunk 1 → High Similarity
Chunk 2 → Medium Similarity
Chunk 3 → Low Similarity
↓
Return Top-K Chunks
Why Not Use a Normal Database? #
A traditional relational database is excellent for structured exact-match
queries.
SELECT *
FROM employees
WHERE department = 'Engineering';
Vector databases are designed for approximate nearest-neighbor or similarity
search over high-dimensional vectors.
ANN Search #
Large vector databases may contain millions or billions of vectors. Comparing a query against every vector can be expensive. Approximate Nearest Neighbor (ANN) indexing methods are used to make vector search much faster.
Interview Answer #
A vector database is designed to store, index and search high-dimensional
vector representations efficiently. In a RAG system, document chunks are
converted into embeddings and stored with their original content and
metadata. At query time, the user’s question is converted into an embedding
and compared with the stored vectors using a similarity or distance metric.
The vector index efficiently identifies the nearest candidates, ranks them
and returns the top-K document chunks for the next stage of the RAG pipeline.
10. How Would You Improve Retrieval Quality in a Large-Scale RAG System? #
Concept #
At small scale, basic vector similarity search may work reasonably well.
However, large production RAG systems require multiple retrieval and
optimization techniques.
1. Improve Chunking #
Use chunking strategies appropriate for the document type. Preserve headings,
paragraph relationships and semantic boundaries wherever possible.
2. Use a Better Embedding Model #
Embedding quality directly affects semantic retrieval. Evaluate multiple embedding models using representative queries rather than choosing a model only because it is popular.
3. Use Hybrid Search #
Hybrid search combines semantic vector retrieval with traditional keyword
retrieval.
Dense retrieval is useful for semantic similarity, while keyword search can
be especially useful for exact terms, product names, IDs, error messages and
technical terminology.
4. Query Rewriting #
The user’s original query can sometimes be transformed into a clearer search
query before retrieval.
Original:
"How does it work?"
Rewritten:
"How does the company's employee reimbursement process work?"
5. Query Expansion #
The system can generate alternative formulations or related terms to improve
recall.
6. Metadata Filtering #
Metadata can dramatically reduce the search space.
filter = {
"department": "engineering",
"document_type": "technical"
}
7. Reranking #
Instead of directly sending the first retrieved documents to the LLM, retrieve
a larger candidate set and use a reranker to identify the most relevant
documents.
Vector Search
↓
Top 50 Candidates
↓
Reranker
↓
Top 5 Documents
↓
LLM
8. Parent-Child Retrieval #
Retrieve a precise child chunk but provide its larger parent section to the
LLM when additional context is necessary.
9. Diversify Results #
Returning five nearly identical chunks is not always useful. Retrieval
strategies such as Maximum Marginal Relevance (MMR) can help balance
relevance with diversity.
10. Evaluate Retrieval #
A production RAG system should have an evaluation dataset containing
representative questions and expected relevant documents.
Important Retrieval Metrics #
Recall@K
→ Did we retrieve the relevant document within top K?
Precision@K
→ How many retrieved documents were relevant?
MRR
→ How highly was the first relevant document ranked?
NDCG
→ How good was the overall ranking of relevant documents?
Interview Answer #
For a large-scale RAG system, I would improve retrieval at multiple levels
rather than relying only on basic vector search. I would optimize chunking
and evaluate embedding models using real application queries. I would also
consider hybrid retrieval, metadata filtering, query rewriting, query
expansion, reranking, parent-child retrieval and result diversification.
Finally, I would create a representative evaluation dataset and measure
retrieval quality using metrics such as Recall@K, Precision@K, MRR and NDCG.
This allows retrieval improvements to be measured objectively instead of
depending only on the quality of the final LLM response.
Complete RAG Architecture #
The following diagram combines the indexing and query phases into one
complete RAG architecture.
Real-World RAG Example #
Problem #
Suppose a company has thousands of technical documents and wants to build an
AI assistant for its developers.
Knowledge Base #
API Documentation
Database Documentation
Deployment Guides
Authentication Documentation
Internal Engineering Policies
Troubleshooting Guides
User Question #
How do I authenticate with the payment API?
RAG Process #
Why RAG Is Useful Here #
The company’s documentation may change frequently. Instead of retraining the LLM every time a document changes, the company can update the RAG knowledge base.
This is one of the major advantages of RAG: the external knowledge source can
be updated independently of the model’s pretrained parameters.
RAG Interview Quick Revision #
RAG
│
├── Indexing
│ ├── Load Documents
│ ├── Extract Text
│ ├── Clean Text
│ ├── Chunk Documents
│ ├── Add Metadata
│ ├── Generate Embeddings
│ └── Store in Vector Database
│
└── Query
├── Receive User Query
├── Process Query
├── Generate Query Embedding
├── Similarity Search
├── Retrieve Top-K
├── Metadata Filtering
├── Reranking
├── Build Context
├── Send Context to LLM
└── Generate Answer
How to Explain RAG in an Interview #
If the interviewer asks you to explain RAG in a few minutes, avoid giving
only the definition. Explain the architecture and clearly distinguish the
offline indexing phase from the online query phase.
"RAG stands for Retrieval-Augmented Generation.
It allows an LLM to use external knowledge instead of relying
only on its pretrained knowledge.
During the indexing phase, we load documents, clean them,
split them into chunks, generate embeddings and store those
embeddings with metadata in a vector database.
When a user asks a question, we convert the query into an
embedding and perform similarity search to retrieve relevant
chunks.
For better production retrieval, we can use metadata filtering,
hybrid search, query rewriting and reranking.
Finally, the retrieved context is passed to the LLM along with
the user's question, and the LLM generates a grounded answer.
The important point is that RAG separates knowledge retrieval
from language generation. Therefore, we can update the external
knowledge base without retraining the LLM every time the
underlying documents change."
11. What Is Reranking? Why Do We Need a Reranker? #
Concept #
Reranking is a second-stage retrieval technique used to reorder the documents returned by an initial retriever. A vector database can quickly retrieve a set of candidate documents, but the initial ranking may not always place the
most relevant document at the top. A reranker examines the relationship between the user’s query and each retrieved document more carefully and produces a better relevance ordering.
Why Do We Need Reranking? #
Suppose a user asks:
How can I reset my company VPN password?
The vector database may retrieve several documents that are semantically
related to passwords, VPNs and account management.
Document 1 → General password policy
Document 2 → VPN installation guide
Document 3 → VPN password reset procedure
Document 4 → Employee account management
Document 5 → Network security policy
The initial retriever may return all of these documents, but the exact VPN
password reset procedure should receive the highest relevance score. A
reranker can analyze the query-document relationship more precisely and
reorder the candidates.
Two-Stage Retrieval #
First Stage: Retrieval #
The first-stage retriever is optimized for speed and recall. It may retrieve
20, 50 or even 100 candidate documents from a large knowledge base.
candidates = vector_store.similarity_search(
query,
k=20
)
Second Stage: Reranking #
The reranker receives the original query and the retrieved candidates and
assigns a more detailed relevance score to each candidate.
20 Retrieved Documents
↓
Reranker
↓
5 Most Relevant Documents
↓
LLM
Advantages #
Initial Retriever
→ Fast candidate retrieval
Reranker
→ Better relevance ordering
Final Context
→ Higher-quality information for the LLM
Trade-Off #
Reranking improves retrieval quality but adds additional computation and
latency. Therefore, a common production architecture is to retrieve a
moderately large candidate set quickly and then rerank only those candidates.
Interview Answer #
Reranking is a second-stage retrieval process that reorders documents returned by an initial retriever according to their relevance to the user’s query. I would typically retrieve a larger candidate set using vector or hybrid search and then pass those candidates through a reranker to select the most relevant documents before sending them to the LLM. This improves precision and context quality, although it adds some computational cost and latency.
12. What Is Hybrid Search? Why Combine Keyword and Vector Search? #
Concept #
Hybrid search combines traditional keyword-based retrieval with semantic
vector search. Keyword search is effective when exact words, identifiers,
product names, error codes or technical terms matter, while vector search is
useful when the query and document use different words but have similar
meaning.
Example #
Consider the query:
What causes HTTP 504 Gateway Timeout?
Keyword search can strongly match the exact term
HTTP 504 Gateway Timeout. Vector search can also retrieve
documents discussing concepts such as upstream server timeouts even if the
exact wording is different.
How Hybrid Search Works #
Keyword Search #
Keyword retrieval searches for exact or related terms. Technologies such as
BM25 are commonly used for lexical retrieval.
Query:
"FAISS IndexFlatL2"
Keyword Search
↓
Documents containing:
FAISS
IndexFlatL2
L2
Vector Search #
Vector search converts the query into an embedding and retrieves documents
that are semantically similar.
Query
↓
Embedding
↓
Vector Similarity
↓
Semantically Similar Documents
Why Combine Them? #
Keyword Search
→ Exact terms
→ Product IDs
→ Error codes
→ Names
→ Technical identifiers
Vector Search
→ Semantic meaning
→ Paraphrases
→ Similar concepts
→ Natural language questions
Example #
A developer asks:
Why am I getting error ORA-00942?
Keyword search is highly useful because the exact error code
ORA-00942 is important. However, vector search can also retrieve documents explaining that the problem may occur when a referenced database table or view does not exist or is inaccessible.
Interview Answer #
Hybrid search combines lexical keyword retrieval with semantic vector retrieval. Keyword search is useful for exact terms such as product names, IDs, error codes and technical terminology, while vector search is useful for semantic similarity and paraphrased queries. Combining both approaches can improve retrieval robustness because the system benefits from both exact matching and semantic understanding. In a production system, I would typically combine the results and optionally apply reranking.
13. How Would You Implement Metadata Filtering in RAG? #
Concept #
Metadata filtering restricts retrieval based on structured information
associated with each document or chunk. Metadata can include department,
document type, author, date, product, access level, language, customer or
other attributes.
Example #
Suppose a company has documents from multiple departments:
Engineering
HR
Finance
Marketing
Legal
A user asks:
What is our backend deployment process?
If the application knows that the question belongs to the Engineering
department, it can filter the search to Engineering documents before or
during vector retrieval.
Metadata Structure #
metadata = {
"department": "engineering",
"document_type": "technical",
"product": "backend",
"year": 2026,
"access_level": "internal"
}
Metadata Filtering Flow #
Example Filter #
filter = {
"department": "engineering",
"document_type": "technical"
}
The retriever can then search only documents matching these metadata
conditions.
Why Metadata Filtering Is Important #
Metadata filtering can reduce irrelevant results, improve retrieval
precision, reduce the search space and support access-control requirements.
It is especially useful in enterprise RAG systems where users may have
different permissions or where the knowledge base contains many unrelated
documents.
Example: Date Filtering #
User Query:
What is the latest API authentication process?
Filter:
document_type = "technical"
year >= 2026
Important Security Consideration #
Metadata filtering should not be treated as the only security mechanism.
For sensitive enterprise data, access control must be enforced at the
application and data-access layers so that users cannot retrieve documents
they are not authorized to access.
Interview Answer #
I would attach structured metadata to every document or chunk during indexing and use that metadata as a retrieval filter during query time. For example, I could filter by department, document type, date, product or access level before performing vector search. This reduces irrelevant candidates and can improve retrieval precision and efficiency. For sensitive enterprise systems, I would also enforce authorization separately rather than relying only on retrieval filters.
14. What Is Query Rewriting and When Would You Use It? #
Concept #
Query rewriting is the process of transforming a user’s original question
into one or more clearer search queries before retrieval. The purpose is to
make the query easier for the retrieval system to understand and improve the
chance of finding relevant documents.
Example #
Consider a conversation:
User:
What is LangChain?
User:
How does it handle memory?
The second question is ambiguous because the user says
“it”. A query rewriting component can transform it into:
How does LangChain handle memory?
Query Rewriting Flow #
When Is Query Rewriting Useful? #
1. Ambiguous Questions
2. Follow-up Questions
3. Very Short Queries
4. Conversational Queries
5. Poorly Formulated Queries
6. Domain-Specific Terminology
7. Complex Search Requirements
Example: Short Query #
Original:
"deployment"
Rewritten:
"What is the production deployment process for the backend API?"
Example: Conversational Query #
Previous:
How does authentication work?
Current:
What about refresh tokens?
The retrieval system may benefit from rewriting this as:
How are refresh tokens handled in the authentication system?
Multi-Query Rewriting #
A system can also generate multiple search queries for one user question.
This can improve recall when the same concept can be expressed in different
ways.
Original Query
↓
┌───────────────┐
│ Query Rewriter│
└───────────────┘
↓
┌────────┬────────┬────────┐
Query 1 Query 2 Query 3
↓ ↓ ↓
Retrieval from Knowledge Base
↓
Combined Results
Trade-Off #
Query rewriting can improve retrieval but adds an additional model call and
therefore may increase latency and cost. It should be evaluated rather than
added automatically to every RAG system.
Interview Answer #
Query rewriting transforms the user’s original question into a clearer or more retrieval-friendly query before searching the knowledge base. I would use it for ambiguous questions, conversational follow-ups, short queries or cases where the original wording does not match the terminology used in the documents. For example, a follow-up such as “How does it handle memory?” can be rewritten using conversation history into “How does LangChain handle memory?” This can improve retrieval recall and precision, but it adds some latency and cost.
15. What Is Multi-Hop RAG? #
Concept #
Multi-hop RAG is a retrieval approach used when answering a question requires information from multiple retrieval steps or multiple pieces of evidence. Instead of performing only one search, the system retrieves information, uses that information to determine the next retrieval step, and continues until enough evidence has been collected.
Simple Example #
Consider the question:
Who is the CEO of the company that acquired Company A?
To answer this, the system may need to determine which company acquired
Company A and then retrieve the CEO of that acquiring company.
Multi-Hop Process #
Example #
Question:
Which framework was used in the project created by the
developer who wrote the authentication module?
Hop 1:
Find the developer who wrote the authentication module.
↓
Hop 2:
Find the project created by that developer.
↓
Hop 3:
Find the framework used by that project.
↓
Final Answer
Why Normal RAG May Not Be Enough #
A single similarity search may retrieve documents containing parts of the
question but may not establish the relationships between those pieces of
information. Multi-hop retrieval explicitly performs multiple retrieval
steps to connect the evidence.
Challenges #
• Error propagation
• Increased latency
• More retrieval calls
• More complex orchestration
• Need for intermediate reasoning
• Greater evaluation complexity
Interview Answer #
Multi-hop RAG is a retrieval approach where answering a complex question requires multiple retrieval steps. The system retrieves an initial piece of evidence, uses that evidence to formulate the next query, retrieves additional information and finally combines the evidence to generate the answer. It is useful for questions involving relationships across multiple documents or entities, but it increases latency and system complexity compared with single-hop retrieval.
16. How Would You Retrieve Information Across Multiple Documents? #
Concept #
In many real-world RAG systems, the answer is not contained in a single
document. Important information may be distributed across multiple files,
pages or knowledge sources. The retrieval system should therefore be able
to retrieve multiple relevant chunks and combine them before generation.
Example #
Suppose a company stores information in:
employee_handbook.pdf
engineering_policy.pdf
deployment_guide.pdf
security_policy.pdf
The user asks:
What security requirements must developers follow before deploying
an application to production?
The answer may require information from both the engineering deployment guide
and the security policy.
Multi-Document Retrieval #
Approach 1: Top-K Retrieval #
Retrieve the top K chunks from the entire knowledge base rather than
searching each document independently.
results = vector_store.similarity_search(
query,
k=10
)
Approach 2: Retrieve by Source #
If the system knows that multiple sources may be required, it can retrieve
candidate chunks from different document groups and then combine them.
Approach 3: Reranking #
When many chunks are retrieved, a reranker can prioritize the strongest
evidence and remove low-quality candidates.
Approach 4: Parent-Child Retrieval #
A small child chunk can be used to identify the relevant location while the
larger parent section can be supplied as context.
Important Consideration #
The system should avoid blindly sending a large number of documents to the
LLM. Too much retrieved information can increase token usage, latency and
noise. The goal is to retrieve sufficient evidence while keeping the final
context relevant.
Interview Answer #
I would treat the entire knowledge base as a searchable collection and retrieve the top relevant chunks across all documents rather than assuming the answer exists in one file. I would then use metadata filtering, deduplication and reranking to select the strongest evidence from multiple sources. If the question requires information that is distributed across documents, I could also use multi-hop retrieval or query decomposition to retrieve each required piece before combining the evidence for the LLM.
17. What Is GraphRAG? When Would You Choose It Over Traditional RAG? #
Concept #
GraphRAG is a retrieval architecture that combines retrieval-augmented
generation with a knowledge graph or graph-based representation of
relationships between entities and concepts.
Traditional vector RAG primarily retrieves semantically similar chunks.
Graph-based retrieval can additionally exploit relationships between entities,
such as people, organizations, products, locations and events.
Traditional RAG vs GraphRAG #
Example #
Suppose a company has information about thousands of employees, projects,
teams and managers.
Employee A
↓ works on
Project X
↓ owned by
Team Y
↓ managed by
Manager Z
A question such as:
Which projects are managed by the manager responsible for
the team working on Project X?
may require following several relationships. A graph representation can make
these relationships explicit.
Knowledge Graph Example #
Alice
↓ works_on
Project A
↓ belongs_to
Engineering Team
↓ managed_by
Bob
When Traditional RAG Is Often Sufficient #
Traditional vector RAG is often suitable for straightforward document
question answering, semantic search, FAQs and situations where the relevant
answer exists within a small number of text chunks.
When Graph-Based Retrieval Can Help #
• Complex entity relationships
• Multi-hop questions
• Organizational knowledge
• Research literature
• Supply-chain relationships
• Large interconnected knowledge bases
• Questions requiring relationship traversal
Trade-Offs #
Graph-based systems can require additional data modeling, entity extraction, relationship extraction, graph construction and maintenance. Therefore, GraphRAG should be selected based on the retrieval problem rather than simply because it is more sophisticated.
Interview Answer #
GraphRAG combines retrieval-augmented generation with a graph representation of entities and their relationships. Traditional RAG mainly retrieves semantically similar text chunks, while GraphRAG can use explicit relationships and graph traversal to answer questions that require connecting multiple entities or documents. I would consider GraphRAG for complex multi-hop and relationship-heavy knowledge bases, while traditional vector RAG is often sufficient for straightforward document retrieval and question answering.
18. How Would You Build RAG for PDFs Containing Tables, Charts and Images? #
Concept #
PDF-based RAG becomes more challenging when documents contain more than plain text. Important information may be stored inside tables, charts, diagrams, images or page layouts. A robust PDF RAG pipeline should therefore preserve the different content types instead of extracting only plain text.
Basic PDF RAG #
Step 1: Extract Text #
Text should be extracted while preserving useful information such as page
numbers, headings and section structure.
Step 2: Extract Tables #
Tables should be parsed into structured representations rather than treating
the table as an arbitrary sequence of characters.
Table:
Product Revenue Growth
A $10M 12%
B $15M 18%
C $8M 10%
The extracted table can then be converted into structured text or another
representation suitable for retrieval.
Step 3: Handle Images and Charts #
Images and charts may contain information that is not present in the PDF’s text layer. A multimodal model or image understanding pipeline can be used to describe or extract relevant information from those visual elements.
Example #
Suppose a financial report contains a revenue chart. A user asks:
Which quarter had the highest revenue?
If the answer exists only in the chart, text-only extraction may fail. The
chart must be processed as visual information.
Page-Level Metadata #
metadata = {
"source": "annual_report.pdf",
"page": 42,
"content_type": "table",
"section": "Revenue Analysis"
}
Multimodal Retrieval #
A production system can maintain separate representations for text, tables
and images or use a multimodal representation depending on the application’s
requirements.
Important Considerations #
• Preserve page numbers
• Preserve section headings
• Extract tables separately
• Process images when necessary
• Preserve relationships between captions and figures
• Store metadata
• Keep source references for citations
• Evaluate extraction quality
Interview Answer #
For PDFs containing tables, charts and images, I would not rely only on plain-text extraction. I would use a document parsing pipeline that preserves text, tables, images, page numbers and section metadata. Tables can be converted into structured text or structured data, while charts and images can be processed using appropriate vision or multimodal models when their content is required. I would store the extracted content with metadata and source references so the retriever can return the correct evidence and the final answer can be grounded in the original PDF.
19. How Would You Handle Scanned PDFs and OCR Errors? #
Concept #
A scanned PDF may contain images of pages instead of machine-readable text. In that situation, normal text extraction may return little or no useful content. Optical Character Recognition (OCR) is required to convert the visual page content into machine-readable text.
Scanned PDF Pipeline #
Step 1: Detect Scanned Pages #
First determine whether the PDF already contains a usable text layer. If it
does not, render the pages as images and run OCR.
Step 2: OCR #
OCR systems identify characters from page images and convert them into text.The quality depends on scan resolution, font, language, page layout and image quality.
Common OCR Errors #
O → 0
I → 1
rn → m
cl → d
missing punctuation
incorrect line breaks
incorrect table structure
incorrect words
Example #
Original:
Employee ID: EMP-1025
OCR Output:
Employee lD: EMP-1O25
These errors can affect retrieval quality, especially when the document
contains names, IDs, numbers, dates, formulas or technical terminology.
OCR Cleaning Pipeline #
OCR Output
↓
Remove Noise
↓
Fix Line Breaks
↓
Normalize Characters
↓
Validate Important Fields
↓
Chunking
↓
Embedding
Confidence Scores #
Where available, OCR confidence scores can help identify text that may require additional validation. Low-confidence regions can be sent through additional processing or reviewed using the original page image.
Keep the Original Source #
It is important to retain the original PDF and page-level metadata. This
allows the application to trace retrieved text back to the source page and
helps with verification and citations.
Example Metadata #
metadata = {
"source": "scanned_contract.pdf",
"page": 17,
"ocr": True,
"ocr_confidence": 0.91
}
Interview Answer #
For scanned PDFs, I would first detect whether a usable text layer exists. If not, I would render the pages as images and apply OCR. Because OCR can introduce errors, I would clean and normalize the extracted text, preserve page-level metadata and validate important fields such as names, IDs, dates and numbers. Where possible, I would use OCR confidence information to identify uncertain regions and retain the original page image for verification. After preprocessing, the cleaned text can be chunked, embedded and indexed like other RAG content.
20. How Would You Evaluate a RAG System? #
Concept #
Evaluating a RAG system is more complex than evaluating only the final generated answer. A RAG pipeline contains multiple components, including document processing, retrieval, reranking and generation. Each component
can introduce errors, so evaluation should measure retrieval quality and answer quality separately.
RAG Evaluation Pipeline #
1. Build an Evaluation Dataset #
Start with a representative collection of questions that real users are
likely to ask. Where possible, define the expected relevant documents or
evidence for each question.
Question:
What is the company's leave policy?
Expected Evidence:
employee_handbook.pdf, page 25
Expected Answer:
Employees receive...
2. Evaluate Retrieval #
First determine whether the retriever is finding the correct information.
Useful retrieval metrics include Recall@K, Precision@K, MRR and NDCG.
Recall@K #
Recall@K measures whether the relevant document appears within the top K
retrieved results.
Recall@K =
Relevant documents retrieved in Top-K
---------------------------------------
Total relevant documents
Precision@K #
Precision@K measures how many of the retrieved results are relevant.
Precision@K =
Relevant retrieved documents
----------------------------
Total retrieved documents
MRR #
Mean Reciprocal Rank measures how highly the first relevant result appears in the ranking. If the correct result appears very early, the reciprocal rank is higher.
NDCG #
Normalized Discounted Cumulative Gain evaluates the quality of the overall
ranking, giving more importance to highly ranked relevant documents.
3. Evaluate the Generated Answer #
Retrieval quality alone is not enough. The generated answer should also be
evaluated for correctness, relevance and whether it is supported by the
retrieved context.
Answer Evaluation
├── Correctness
├── Relevance
├── Completeness
├── Faithfulness
└── Groundedness
4. Faithfulness / Groundedness #
A grounded answer should be supported by the retrieved context rather than
introducing unsupported claims.
Example #
Retrieved Context:
Employees receive 20 days of annual leave.
Generated Answer:
Employees receive 20 days of annual leave.
→ Supported by context
If the model instead says that employees receive 30 days when the retrieved
context says 20 days, the answer is not grounded in the retrieved evidence.
5. End-to-End Evaluation #
6. Evaluate Different Components Separately #
If the final answer is incorrect, the first question should be whether the
retriever found the correct information.
Wrong Answer
↓
Was correct evidence retrieved?
↓
┌───────────────┐
│ │
NO YES
↓ ↓
Fix Retrieval Check LLM
↓
Check Prompt / Context
7. Production Monitoring #
After deployment, evaluation should continue using production feedback,
representative queries, latency measurements, retrieval statistics, user
feedback and failure analysis.
Important Metrics #
Retrieval
→ Recall@K
→ Precision@K
→ MRR
→ NDCG
Generation
→ Correctness
→ Relevance
→ Faithfulness
→ Groundedness
System
→ Latency
→ Token Usage
→ Cost
→ Failure Rate
Interview Answer #
I would evaluate a RAG system at both the retrieval and generation levels. First, I would create a representative evaluation dataset containing user questions and expected relevant evidence. For retrieval, I would measure metrics such as Recall@K, Precision@K, MRR and NDCG. For generated answers, I would evaluate correctness, relevance, faithfulness and groundedness. I would also monitor latency, token usage and cost in production. If an answer is wrong, I would first determine whether the correct evidence was retrieved; if retrieval was correct, I would then investigate the prompt, context construction and generation stage.
RAG Interview Quick Revision: Questions 11–20 #
11. Reranking
→ Reorders retrieved candidates based on query-document relevance.
12. Hybrid Search
→ Combines keyword retrieval + vector retrieval.
13. Metadata Filtering
→ Restricts retrieval using structured metadata.
14. Query Rewriting
→ Converts an unclear or conversational query into a better
retrieval query.
15. Multi-Hop RAG
→ Uses multiple retrieval steps to answer complex questions.
16. Multi-Document Retrieval
→ Retrieves and combines evidence from multiple documents.
17. GraphRAG
→ Uses graph relationships for relationship-heavy and
multi-hop retrieval.
18. PDF RAG
→ Handles text, tables, charts, images and page-level metadata.
19. Scanned PDF / OCR
→ Converts page images into text and handles OCR errors.
20. RAG Evaluation
→ Evaluates retrieval, generation and production performance.
How to Answer Advanced RAG Questions in an Interview #
For advanced RAG interview questions, do not only define the technology.
Explain where it fits in the pipeline, why it is needed, how you would
implement it, what trade-offs it introduces and how you would evaluate it.
For example, when discussing reranking, explain that the initial retriever
provides candidates while the reranker improves their ordering. When
discussing hybrid search, explain why keyword and semantic retrieval
complement each other. For production systems, also mention practical
concerns such as latency, cost, metadata, security, evaluation and
observability.
One-Line Interview Summary #
A production RAG system is not simply an embedding model connected to a vector database; it is a retrieval pipeline that may include document processing, intelligent chunking, embeddings, metadata filtering, hybrid retrieval, query rewriting, reranking, multi-document or multi-hop retrieval, multimodal document processing and systematic evaluation before the final context is provided to the LLM.