AI Engineering Interview Questions: RAG Evaluation & Safety Metrics
Reviewed by Mark Dickie · Last updated
4 questions on this page you can answer and see marked on the spot. Go to the first one
RAG evaluation metrics are quantitative measures that assess whether a retrieval-augmented generation system produces grounded answers, stays relevant to the query, and avoids hallucinated or unsafe output. For interview purposes, you should know the core metric families (faithfulness/groundedness, answer relevance, context precision and recall, plus safety scoring), along with the tooling used to compute them such as RAGAS, TruLens, and DeepEval. Be ready to explain how context-level and answer-level metrics differ and how to catch ungrounded claims with an LLM-as-judge; you should also know where human review still adds value. Safety-specific evaluation layers refusal accuracy and jailbreak resistance on top of toxicity scoring.
What does an AI Engineering interview test on RAG evaluation?
Interviews on this topic tend to split into three areas: metric definitions, pipeline debugging, and safety red-teaming.
| Metric | What it measures | Common tooling |
|---|---|---|
| Faithfulness / Groundedness | Whether the generated answer is supported by retrieved context, not invented | RAGAS, DeepEval, LangSmith |
| Answer Relevance | Whether the answer actually addresses the question asked | RAGAS, TruLens |
| Context Precision | Whether the top-ranked retrieved chunks are the useful ones | RAGAS, custom rank eval |
| Context Recall | Whether all information needed to answer is present in retrieved context | RAGAS |
| Refusal Accuracy | Whether the model correctly refuses unsafe or out-of-scope prompts | DeepEval, custom safety eval |
| Toxicity Score | Probability that output contains harmful, biased, or profane content | Perspective API, OpenAI Moderation |
How do you evaluate RAG safety and grounding?
- Extract each claim from the generated answer and verify it against the retrieved context. If a claim has no supporting passage, it counts as ungrounded. This is the faithfulness check, often implemented with an LLM-as-judge that scores each claim individually.
- Measure answer relevance by checking whether the response maps back to the query intent, not just the query keywords.
- Score context precision and recall to confirm the retrieval step is pulling the right chunks. A faithful answer is meaningless if the context was wrong.
- Run safety evaluations: refusal accuracy on harmful prompts, toxicity scoring on outputs, and jailbreak resistance on adversarial inputs.
- Log everything (query, retrieved chunks, answer, per-claim scores) so failures are traceable to either the retriever or the generator.
Key facts
- Tarmac has 21 AI Engineering interview questions on this topic, 10 of them on this page, at difficulty 1–5 of 5.
- Tarmac tracked 2,298 job postings asking for AI Engineering in September 2026.
- Roles asking for AI Engineering advertise a median base salary of US$179,900, across 791 job postings as of September 2026.
- Tarmac last reviewed these AI Engineering interview questions on 5 October 2026.
At a glance
| Questions | 10 shown · 21 in the bank |
|---|---|
| Difficulty | 1–5 of 5 |
| Formats | Ordering, Design exercise, Code output, Find the bug, Flashcard, Multiple choice, Multiple answer, True / false, Short answer, Fill in the blank |
What you'll review
- rag eval metrics
Practice questions
Try one before you open the answer. Pick an option and press Check; it's marked on the spot.
RAG evaluation metrics assess different stages of the retrieve-then-generate pipeline. Arrange these three metrics in order from the one that evaluates the earliest pipeline stage to the one that evaluates the full end-to-end pipeline.#
Put these in order
Show answer
Context relevance, faithfulness, and answer correctness are ordered by the pipeline scope each evaluates: retrieval first, then generation, then end-to-end. Context relevance assesses whether retrieved passages match the query (retrieval stage). Faithfulness checks whether the generated answer is grounded in retrieved context (generation stage). Answer correctness compares the final answer to ground-truth labels, evaluating the full pipeline end-to-end.
Context relevance evaluates only the retrieval stage — it needs just the query and the retrieved passages, so it can be assessed as soon as retrieval is complete (earliest). Faithfulness evaluates the generation stage — it checks whether the generated answer is grounded in the retrieved context, which requires the generation step to be complete. Answer correctness evaluates the entire end-to-end pipeline — it compares the final answer against ground-truth labels, measuring the combined effect of both retrieval and generation quality (broadest scope, last). The progression retrieval stage → generation stage → end-to-end pipeline gives a → b → c.
You are building a retrieval-augmented generation (RAG) system that answers questions over a product documentation corpus. Your manager asks you to set up an evaluation pipeline that can score how well the system is performing. Design an evaluation approach that covers at least two distinct RAG-specific metrics (one measuring retrieval quality and one measuring generation quality). For each metric, describe: (1) what it measures, (2) what inputs you need to compute it, and (3) a simple way to compute or estimate it.#
Show answer
I would evaluate the RAG pipeline along two axes: retrieval quality and generation quality.
For retrieval quality, I would use context recall. Context recall measures whether all the information needed to answer the question is present in the retrieved chunks. To compute it, I need: (1) the user question, (2) a ground-truth answer, and (3) the retrieved context chunks. A simple way to estimate it is to use an LLM-as-judge: given the ground-truth answer and the retrieved context, ask the LLM to classify each claim in the ground-truth answer as supported or not supported by the context, then compute the fraction of supported claims. A high context recall means the retriever is surfacing the right documents.
For generation quality, I would use faithfulness. Faithfulness measures whether the generated answer is fully supported by the retrieved context — i.e., the model is not hallucinating. To compute it, I need: (1) the generated answer and (2) the retrieved context chunks. The computation involves breaking the generated answer into individual statements and checking each one against the context. I can use an LLM-as-judge to verify each statement, then report the ratio of supported statements to total statements. A faithfulness score of 1.0 means every claim in the answer is grounded in the retrieved context.
Together, these two metrics give a basic but informative picture: context recall tells me whether the retriever is finding the right information, and faithfulness tells me whether the generator is using that information without making things up.
This is a junior-level recall and applied-basics question about the standard RAG evaluation metrics. A difficulty-1 design exercise checks whether the candidate can name and describe foundational RAG metrics — one for retrieval (context recall/precision/hit rate) and one for generation (faithfulness/answer relevance) — and identify the inputs each requires. The rubric rewards naming a metric from each category and correctly describing its inputs and computation method. Criterion c1 was broadened so the description fits any of the listed retrieval metrics rather than prescribing a single definition that only applies to context recall/precision.
In RAG evaluation, a simple context recall metric measures what fraction of the ground-truth answer's tokens are covered by the retrieved context. What does the following Python code print?#
def context_recall(retrieved_context, ground_truth_answer):
"""Fraction of ground-truth answer tokens present in retrieved context."""
context_tokens = set(retrieved_context.lower().split())
answer_tokens = ground_truth_answer.lower().split()
covered = sum(1 for t in answer_tokens if t in context_tokens)
return round(covered / len(answer_tokens), 2)
context = "Paris is the capital of France"
answer = "The capital of France is Paris and has a population exceeding 2 million"
print(context_recall(context, answer))Show answer
0.46
The context string "Paris is the capital of France" is lowercased and split into the set {"paris", "is", "the", "capital", "of", "france"}. The answer string "The capital of France is Paris and has a population exceeding 2 million" is lowercased and split into 13 tokens: ["the", "capital", "of", "france", "is", "paris", "and", "has", "a", "population", "exceeding", "2", "million"]. Each answer token is checked for membership in the context set. The covered tokens are: the, capital, of, france, is, paris — 6 tokens. 6 / 13 = 0.4615…, which round(..., 2) rounds to 0.46. The remaining 7 tokens ("and", "has", "a", "population", "exceeding", "2", "million") are absent from the context set, lowering recall.
The function below is intended to compute the faithfulness metric for a RAG system: the fraction of claims in the generated answer that are supported by (i.e., appear in) the retrieved context. Both arguments are sets of atomic claim strings. A faithfulness score of 1.0 means every answer claim is grounded in the context; 0.0 means none are. Find the buggy line.#
def faithfulness(answer_claims, context_claims):
"""Faithfulness = fraction of answer claims grounded in context.
answer_claims: set of atomic claims extracted from the generated answer
context_claims: set of atomic claims extracted from the retrieved context
"""
if not answer_claims:
return 1.0
return len(answer_claims & context_claims) / len(answer_claims | context_claims)Show answer
The bug is on line 9.
Line 9 uses len(answer_claims | context_claims) (the size of the union) as the denominator. This computes a Jaccard-like overlap ratio, not faithfulness. Faithfulness is defined as the fraction of answer claims that are grounded in the context, so the denominator must be len(answer_claims). With the union as the denominator, adding more irrelevant claims to the context inflates the denominator and depresses the score even when all answer claims are fully supported, yielding an incorrect metric. The correct line should read: return len(answer_claims & context_claims) / len(answer_claims).
How do RAG-specific eval frameworks like RAGAS measure quality, and why score components instead of just the final answer?#
Show answer
They score the pipeline per stage. Retrieval: context precision (are the retrieved passages relevant, with relevant ones ranked high?) and context recall (were all the passages needed to answer fetched? — usually checked against a reference). Generation: faithfulness (is every claim in the answer grounded in the retrieved context, not hallucinated?) and answer relevance (does the answer address the question?). Many of these are LLM-judged and reference-free — faithfulness and answer relevance need no gold answer, though context recall typically does. Why per-component: a wrong answer might be retrieval's fault (needed passage never fetched → low recall) or generation's fault (evidence was there but ignored → low faithfulness). Stage-level metrics localize the failure so you know what to fix.
RAGAS-style evals decompose RAG into retrieval (context precision/recall) and generation (faithfulness, answer relevance), mostly via an LLM judge. The point is attribution: an end-to-end score says the answer is bad; component metrics say which stage broke, and therefore what to fix.
A RAGAS-style evaluation decomposes a RAG system into a retrieval stage and a generation stage. Which mapping of metric to the stage it primarily measures is correct?#
Options
Show answer
Context precision and context recall measure retrieval; faithfulness and answer relevance measure generation. The retrieval metrics judge the chunks the retriever returned — precision is whether the ranked passages are relevant, recall is whether all the needed passages were fetched. The generation metrics judge what the model did with that context — faithfulness is whether every claim is supported by the retrieved context, answer relevance is whether the answer addresses the question. The split exists so you can attribute a wrong answer to the right stage.
RAGAS-style frameworks score the two stages separately. Retrieval metrics judge the chunks the retriever returned: context precision (are the retrieved/ranked passages actually relevant — signal vs. noise) and context recall (did retrieval fetch all the passages needed to answer, typically checked against a reference answer). Generation metrics judge what the model did with that context: faithfulness (is every claim in the answer supported by the retrieved context — grounded, not hallucinated) and answer relevance (does the answer actually address the question). The point of the split is attribution: a wrong answer might be retrieval's fault (the needed passage was never fetched → low context recall) or generation's fault (the passage was present but the model ignored or contradicted it → low faithfulness) — a single end-to-end score can't tell you which.
You're building a component-wise RAG eval. Which of the following are primarily retrieval-stage signals (as opposed to generation-stage)? Select all that apply.#
Options
Pick every one that applies.
Show answer
The retrieval-stage signals are context recall (whether all the passages needed to answer were retrieved), context precision (whether the retrieved passages are relevant and rank highly), and classic IR metrics like hit rate or recall@k of the gold passage in the top-k. Faithfulness and answer relevance are generation-stage: they judge whether the produced answer is grounded in context and addresses the question, and can be poor even when retrieval was perfect.
Retrieval-stage signals judge what the retriever returned, independent of what the model writes: context recall (did we fetch everything needed), context precision (are the returned/ranked passages relevant), and classic IR metrics like recall@k / hit rate of the gold passage in the top-k. Faithfulness and answer relevance are generation-stage — they evaluate the produced answer (is it grounded in the context, does it address the question) and can be poor even when retrieval was perfect. Keeping the two sets separate is what lets you attribute a failure to the right stage rather than guessing.
If a RAG answer scores perfectly on faithfulness, it is guaranteed to be a correct and complete answer to the user's question.#
Options
Show answer
False. Faithfulness measures only that every claim in the answer is supported by the retrieved context — that the model did not fabricate beyond its evidence. It says nothing about whether that context was the right or complete evidence, or whether the answer is on-topic. A perfectly faithful answer can still be irrelevant (faithfully summarizing the wrong passages) or incomplete (faithful to context that was missing the key fact because retrieval failed). That is why evals score faithfulness alongside answer relevance and retrieval metrics.
Faithfulness measures only that every claim in the answer is supported by the retrieved context — i.e. the model didn't fabricate beyond its evidence. It says nothing about whether that context was the right or complete evidence, nor whether the answer is on-topic. A perfectly faithful answer can still be (1) irrelevant — faithfully summarizing the wrong passages instead of answering the question (low answer relevance), or (2) incomplete or wrong — faithful to context that was itself missing the key fact because retrieval failed (low context recall). That's exactly why RAGAS-style evals score faithfulness alongside answer relevance and the retrieval metrics, instead of trusting any single one.
Why do RAG evaluation frameworks (e.g. RAGAS) score the retrieval stage and the generation stage with separate metrics instead of one end-to-end answer-quality score?#
Show answer
Because a single end-to-end score tells you the answer was bad but not why, and the two stages fail for different reasons with different fixes. Retrieval metrics like context recall and context precision check whether the right, relevant passages were fetched at all; generation metrics like faithfulness and answer relevance check whether the model grounded its answer in that context and actually addressed the question. Decomposing localizes the fault: if context recall is low, the needed passage was never retrieved, so you fix chunking, the embedding model, or reranking; if recall is fine but faithfulness is low, the evidence was present but the model ignored or contradicted it, so you fix the prompt or the model. Separate metrics turn 'the answer is wrong' into an actionable diagnosis.
A blended score can't separate 'we never fetched the right evidence' (retrieval — low context recall) from 'we had the evidence and still got it wrong' (generation — low faithfulness). Those have different fixes — chunking/embeddings/reranking versus prompt/model — so component metrics convert a failing answer into a specific stage you can act on.
In component-wise RAG evaluation, two retrieval-stage metrics are context _____ (are the retrieved passages relevant?) and context _____ (were all the passages needed to answer actually retrieved?). On the generation side, _____ measures whether every claim in the answer is supported by the retrieved context.#
Show answer
In component-wise RAG evaluation, two retrieval-stage metrics are context precision (are the retrieved passages relevant?) and context recall (were all the passages needed to answer actually retrieved?). On the generation side, faithfulness measures whether every claim in the answer is supported by the retrieved context.
Context precision = relevance/signal of the retrieved passages (relevant ones ranked high); context recall = coverage (did retrieval fetch everything needed, usually checked against a reference answer). Faithfulness (a.k.a. groundedness) = the share of the answer's claims entailed by the retrieved context — the anti-hallucination signal. The fourth common metric, answer relevance, checks that the answer actually addresses the question. Splitting retrieval metrics from generation metrics is what lets you attribute a failure to the right stage.
Related interview questions
Job market
See ai-engineering salaries and hiring demand from live job postings.
The other 11 questions
This page shows 10 and marks what you pick. That's as far as a page can go. A free account opens the other 11 and keeps every answer. What you miss comes back until it's right: after a day, then at longer gaps.
Free · the whole bank · 100 marked answers per 30 days · written feedback on the paid plan