AI Engineering (RAG) Interview Questions
Reviewed by Mark Dickie · Last updated
Retrieval-augmented generation (RAG) is a pattern that feeds a language model relevant text retrieved from an external knowledge source at inference time, so the model's answers are grounded in documents rather than parametric memory alone. RAG interviews test whether you can build and debug the full pipeline: chunking documents, picking an embedding model, querying a vector store, reranking results, and writing prompts that keep the model within retrieved context. Expect questions about chunk size and overlap tradeoffs, hybrid (dense + sparse) retrieval, embedding dimensionality, and how to measure whether outputs are faithful to the source. You should also be able to explain where RAG fails (stale indices, poor chunk boundaries, retrieval that misses the right passage) and what you would change.
What does a RAG interview test?
| Pipeline stage | Key decision | Common failure mode |
|---|---|---|
| Document chunking | Chunk size, overlap, split strategy | Chunks cut mid-sentence or too large to fit the context window |
| Embedding model | Model choice, dimensionality, domain fit | Embeddings that do not capture domain-specific vocabulary |
| Vector store | Index type (HNSW, IVF), similarity metric | Slow queries at scale or wrong distance metric |
| Retrieval | Top-k value, dense vs. sparse vs. hybrid | Low recall: the right passage never surfaces |
| Reranking | Cross-encoder vs. bi-encoder, rerank depth | Most relevant chunk ranked below the top-k cutoff |
| Generation | Prompt template, context window management | Model hallucinates beyond retrieved text or ignores context |
How do you evaluate a RAG system?
- Measure retrieval quality with recall@k and precision@k to check whether the right passages are being surfaced.
- Score generation with faithfulness (is every claim supported by retrieved text?) and answer relevance (does the answer address the question?).
- Track context precision: of the passages you retrieved, how many were actually needed to answer.
- Run an LLM-as-judge evaluation on a held-out set of queries and compare against human ratings for calibration.
- Monitor latency end-to-end, since retrieval plus reranking plus generation adds up and matters for production systems.
What are the most common RAG failure modes?
Stale indices are the first thing to break: if documents change and the vector store is not re-embedded, answers reference outdated information. Poor chunk boundaries are just as common; splitting on a fixed token count can tear a key sentence or table in half. Retrieval can return semantically close but contextually wrong passages, which is where reranking helps. And the generation step can still hallucinate, especially when the prompt does not explicitly instruct the model to answer only from the provided context.
Key facts
- Tarmac has 89 AI Engineering interview questions on this topic, 10 of them on this page, at difficulty 1–5 of 5.
- Tarmac tracked 2,298 job postings asking for AI Engineering in September 2026.
- Roles asking for AI Engineering advertise a median base salary of US$179,900, across 791 job postings as of September 2026.
- Tarmac last reviewed these AI Engineering interview questions on 28 September 2026.
At a glance
| Questions | 10 shown · 89 in the bank |
|---|---|
| Difficulty | 1–5 of 5 |
| Formats | Multiple answer, Fill in the blank, Code output, Multiple choice, Design exercise |
What you'll review
- rag basics
- rag
Practice questions
Try one before you open the answer. Pick an option and press Check; it's marked on the spot.
In a production Retrieval-Augmented Generation (RAG) pipeline, which of the following techniques directly improve retrieval quality (i.e., the relevance of retrieved chunks), as opposed to improving generation quality or system throughput?#
Options
Pick every one that applies.
Show answer
The techniques that directly improve retrieval quality in a RAG pipeline are: hybrid search with BM25 + dense vectors + RRF, query rewriting / HyDE, and cross-encoder re-ranking. Hybrid search combines lexical and semantic signals; HyDE/query rewriting closes the embedding gap between query and relevant documents; and cross-encoder re-ranking refines the candidate list before generation. Temperature and context-window size affect generation, not retrieval.
Hybrid search (dense + BM25 + RRF) improves recall and precision by combining lexical and semantic signals — this directly affects which chunks are retrieved. Query rewriting and HyDE improve retrieval by making the query embedding closer to relevant document embeddings in vector space. Re-ranking with a cross-encoder is a post-retrieval step that re-scores candidates for relevance before passing them to the LLM — still part of the retrieval quality pipeline. Increasing LLM temperature affects generation diversity, not retrieval. Reducing the context window affects inference cost/latency, not retrieval quality.
The acronym RAG stands for _____-Augmented Generation, a technique that supplements a language model's parametric knowledge with externally retrieved information at inference time.#
Show answer
The acronym RAG stands for Retrieval-Augmented Generation, a technique that supplements a language model's parametric knowledge with externally retrieved information at inference time.
RAG stands for Retrieval-Augmented Generation. The retrieval step fetches relevant passages from an external knowledge source (e.g., a vector database) and includes them in the prompt context so the model can ground its answer in up-to-date, source-specific information.
In a basic RAG (Retrieval-Augmented Generation) pipeline, source documents are first split into smaller passages through a process called _____. Each passage is then converted into a numerical vector using an _____ model and stored in a vector database so that relevant passages can be found via similarity search at query time.#
Show answer
In a basic RAG (Retrieval-Augmented Generation) pipeline, source documents are first split into smaller passages through a process called chunking. Each passage is then converted into a numerical vector using an embedding model and stored in a vector database so that relevant passages can be found via similarity search at query time.
The two foundational preprocessing steps in a RAG pipeline are chunking (dividing documents into manageable passages) and embedding (converting each passage into a dense vector). These vectors are what the vector database indexes and compares against the query vector at retrieval time.
In a RAG system, cosine similarity is commonly used to rank document chunks against a user query. Trace the following Python code that computes cosine similarity between a document vector and a query vector. What is printed?#
import numpy as np
def cosine_similarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
doc = np.array([3, 4, 0])
query = np.array([3, 0, 0])
print(round(cosine_similarity(doc, query), 4))Show answer
0.6
The dot product of [3,4,0] and [3,0,0] is 33 + 40 + 0*0 = 9. The norm of the doc vector is sqrt(3²+4²+0²) = sqrt(25) = 5, and the norm of the query vector is sqrt(3²+0²+0²) = 3. Cosine similarity = 9 / (5 * 3) = 9/15 = 0.6. Rounding to 4 decimal places gives 0.6.
A minimal RAG retrieval step computes cosine similarity between a query vector and a list of document vectors, then returns the top-k document texts. Trace the code below and determine what it prints.#
import math
def cosine_sim(a, b):
dot = sum(x*y for x, y in zip(a, b))
na = math.sqrt(sum(x*x for x in a))
nb = math.sqrt(sum(x*x for x in b))
return dot / (na * nb)
def retrieve(query, docs, k=2):
scored = [(cosine_sim(query, d['vec']), d['text']) for d in docs]
scored.sort(key=lambda x: x[0], reverse=True)
return [text for score, text in scored[:k]]
docs = [
{'text': 'alpha', 'vec': [1, 0, 0]},
{'text': 'beta', 'vec': [0, 1, 0]},
{'text': 'gamma', 'vec': [1, 1, 0]},
]
query = [1, 1, 0]
print(retrieve(query, docs, k=2))Show answer
['gamma', 'alpha']
Cosine similarities with query [1,1,0]: gamma = 2/(√2·√2) = 1.0; alpha = 1/(√2·1) ≈ 0.707; beta = 1/(√2·1) ≈ 0.707. Sorting descending puts gamma first, then alpha and beta tie at ≈0.707. Python's sort is stable, so alpha (originally before beta) stays ahead. Taking k=2 yields ['gamma', 'alpha'].
In a two-stage RAG pipeline, a fast first-stage retriever (such as a bi-encoder) pulls a broad candidate set of N documents, then a _____ model re-scores those candidates and narrows them to the final top-k. The first stage is optimized for _____ (casting a wide net), while the second stage is optimized for _____ (surfacing the most relevant results at the top).#
Show answer
In a two-stage RAG pipeline, a fast first-stage retriever (such as a bi-encoder) pulls a broad candidate set of N documents, then a cross-encoder model re-scores those candidates and narrows them to the final top-k. The first stage is optimized for recall (casting a wide net), while the second stage is optimized for precision (surfacing the most relevant results at the top).
Two-stage retrieval is a standard RAG optimization. The first stage uses a cheap, high-throughput model (bi-encoder or BM25) to maximize recall — ensuring relevant passages enter the candidate pool. The second stage uses a more expensive but more accurate cross-encoder (a reranker) to re-score candidates and maximize precision, so only the most relevant passages reach the LLM's context window.
You are tuning a production RAG pipeline that uses HNSW (Hierarchical Navigable Small World) as the ANN index in your vector database (e.g., FAISS HNSW, Qdrant, pgvector with HNSW). During evaluation you observe that recall@10 is too low but query latency is well within budget. Which HNSW parameter should you increase to directly improve query-time recall at the cost of higher latency?#
Options
Show answer
Increase ef_search. In HNSW, ef_search controls the size of the dynamic candidate list explored during layer traversal at query time — a larger value improves recall by examining more neighbors at the cost of additional distance computations and higher latency. ef_construction and M are build-time parameters that shape the graph structure, and ml governs layer-assignment probability; none of them can be tuned at query time to trade recall for speed.
ef_search (called ef in FAISS) is the parameter that controls how many candidates HNSW explores in the dynamic list during each layer traversal at query time. Increasing it widens the search frontier, improving recall at the cost of more distance computations and higher latency. ef_construction affects index build quality and the resulting graph structure, but once the index is built it does not change query-time behavior. M (max connections per node per layer) also affects graph connectivity but is set at build time; while a low M can limit achievable recall, the parameter you tune at query time to trade recall for latency is ef_search. ml (the level normalization factor) controls the exponential probability distribution for layer assignment and is not a query-time tunable.
A RAG retrieval pipeline uses Maximal Marginal Relevance (MMR) to diversify retrieved chunks. Trace the following MMR selection code and determine the exact output printed to stdout.#
import numpy as np
def mmr_select(query_sim, doc_sim, k=3, lambda_=0.5):
selected = []
candidates = list(range(len(query_sim)))
for _ in range(k):
best = None
best_score = -1
for d in candidates:
relevance = lambda_ * query_sim[d]
if selected:
redundancy = (1 - lambda_) * max(doc_sim[d][s] for s in selected)
else:
redundancy = 0
score = relevance - redundancy
if score > best_score:
best_score = score
best = d
selected.append(best)
candidates.remove(best)
return selected
query_sim = [0.9, 0.8, 0.3, 0.7]
doc_sim = [
[1.0, 0.6, 0.1, 0.2],
[0.6, 1.0, 0.1, 0.5],
[0.1, 0.1, 1.0, 0.1],
[0.2, 0.5, 0.1, 1.0],
]
print(mmr_select(query_sim, doc_sim, k=3, lambda_=0.5))Show answer
[0, 3, 1]
The MMR score for each candidate is λ·Sim(d,query) − (1−λ)·max_{s∈selected} Sim(d,s). With λ=0.5:
Round 1 (selected=[], no redundancy term): scores are d0=0.45, d1=0.40, d2=0.15, d3=0.35. Select d0. selected=[0].
Round 2 (selected=[0], redundancy = 0.5·doc_sim[d][0]):
- d1: 0.40 − 0.5·0.6 = 0.40−0.30 = 0.10
- d2: 0.15 − 0.5·0.1 = 0.15−0.05 = 0.10
- d3: 0.35 − 0.5·0.2 = 0.35−0.10 = 0.25
Select d3 (highest). selected=[0,3].
Round 3 (selected=[0,3], redundancy = 0.5·max(doc_sim[d][0], doc_sim[d][3])):
- d1: 0.40 − 0.5·max(0.6,0.5) = 0.40−0.30 = 0.10
- d2: 0.15 − 0.5·max(0.1,0.1) = 0.15−0.05 = 0.10
Tie at 0.10; the code uses strict > so the first candidate encountered (d1) wins. Select d1. selected=[0,3,1].
The function returns [0, 3, 1].
In an HNSW (Hierarchical Navigable Small World) index used for dense vector retrieval in a RAG pipeline, which statement correctly describes the roles of the three key parameters M, ef_construction, and ef_search?#
Options
Show answer
M, ef_construction, and ef_search in HNSW serve distinct roles: M sets the maximum number of graph connections per node per layer, governing memory and navigability; ef_construction sets the candidate list size during index building, governing graph quality and build time; ef_search sets the candidate list size during querying, governing the recall-vs-latency tradeoff. The number of layers is determined probabilistically, not by any of these three parameters.
In HNSW, M is the maximum number of bidirectional connections per node per layer, directly impacting memory usage (more edges = more storage) and graph navigability. ef_construction governs the size of the dynamic candidate list used while inserting nodes during index build—higher values produce a better-quality graph at the cost of slower build times. ef_search governs the same dynamic list during query traversal—higher values improve recall at the cost of higher query latency. The option that swaps these roles is wrong; the options that assign each parameter to an incorrect concern are also wrong. The number of graph layers in HNSW is determined by a probabilistic decay function (exponential level assignment), not by M directly, so the option claiming M controls the number of graph layers is wrong on that point as well.
Design a production RAG system for a multinational law firm with the following requirements:#
Show answer
We would build a multi-stage RAG pipeline with the following components:
Chunking & Citation Provenance: Documents are parsed into a hierarchical structure (document → section → paragraph). Each paragraph becomes a chunk with metadata: {doc_id, doc_version, paragraph_id, section_path, jurisdiction, matter_id, client_id, language, effective_date}. Chunks are embedded using a legal-domain-tuned embedding model. At generation time, the LLM is prompted to include [doc_id, paragraph_id] citations for every claim. A post-processing validation step checks that every cited chunk was actually in the retrieved context set and is accessible to the user—if a citation references a chunk not in the retrieved set, the answer is flagged for review.
Access Control: Access control is enforced at the vector store level using metadata pre-filters. Each query includes the user's authorized matter_id and client_id list as filter predicates on the ANN search (supported by Pinecone, Weaviate, Milvus via metadata filtering). We use a single HNSW index with metadata filtering rather than separate indexes per matter, since the matter count is large and many matters share reference documents (e.g., public statutes). Unauthorized documents are never returned by the search—they are excluded by the filter before the ANN traversal produces candidates.
Document Versioning: When a document is updated, new chunks are embedded and inserted with an incremented version number and a new effective_date. Old version chunks are soft-deleted (marked with tombstone=true and superseded_by= new_version) rather than immediately removed, preserving the audit trail. At query time, a metadata filter (tombstone=false AND version=latest) ensures only current versions are retrieved. A background compaction job periodically rebuilds index shards to physically remove tombstoned chunks and reclaim memory.
Scale & Retrieval Quality: At 50M+ documents, we shard the index by jurisdiction (reducing per-shard size and enabling jurisdiction-aware routing—queries specifying a jurisdiction hit only the relevant shard). Retrieval uses hybrid search: dense HNSW retrieval (top-100) fused with BM25 sparse retrieval (top-100) via reciprocal rank fusion. A legal-domain fine-tuned cross-encoder re-ranks the fused top-50 down to the final top-10 chunks passed to the LLM. HNSW parameters (M=32, ef_construction=400, ef_search=64) are tuned against a held-out recall@10 benchmark set of 5,000 annotated legal queries, with ef_search adjusted dynamically based on shard size.
Multi-Jurisdiction & Multilingual: We use a multilingual embedding model (e.g., multilingual-e5-large) so queries in one language retrieve documents in another without a translation step. Jurisdiction is stored as chunk metadata and used as an optional explicit filter when the user specifies a jurisdiction. For cross-jurisdiction comparison queries (e.g., 'Compare force-majeure clauses in NY and German law'), the query is routed to multiple jurisdiction shards in parallel; results from each jurisdiction are fetched separately (top-k per jurisdiction) and merged before cross-encoder re-ranking, ensuring balanced representation from each jurisdiction in the final context.
This rubric evaluates whether the candidate can architect a RAG system that addresses the five hardest production concerns in a legal-domain setting: citation provenance, pre-retrieval access control, document versioning, retrieval quality at scale, and cross-lingual/multi-jurisdiction handling. Each criterion tests a distinct architectural decision; an answer that omits any one leaves a critical production gap. The sample answer satisfies all five: hierarchical chunking with paragraph metadata (c1), metadata pre-filtering at the vector store (c2), version increment with tombstoning and query-time filtering (c3), hybrid search + cross-encoder re-ranking + jurisdiction sharding + HNSW parameter tuning (c4), and multilingual embeddings with parallel multi-jurisdiction routing (c5).
Related interview questions
Job market
See ai-engineering salaries and hiring demand from live job postings.
The other 79 questions
This page shows 10 and marks what you pick. That's as far as a page can go. A free account opens the other 79 and keeps every answer. What you miss comes back until it's right: after a day, then at longer gaps.
Free · the whole bank · 100 marked answers per 30 days · written feedback on the paid plan