TechNova is a fictional company used as a running example throughout this series.
The Team That Blamed the Model
TechNova’s support assistant told a customer they could return a refurbished WH-1000 within 15 days. The real limit for refurbished items bought from the TechNova Outlet is 7 days.
The team’s first guess was the model. They started testing newer, more expensive ones.
The model was fine. The answer was wrong because of what reached it, and one look at the retrieved chunks would have shown that.
A wrong answer can come from three places, and each one has a different fix.
| Layer | What it does | What to inspect |
|---|---|---|
| Retrieval | Finds the chunks that should answer the question | The retrieved chunks and their ranks |
| Assembly | Builds the prompt from those chunks | The final prompt the model received |
| Generation | Writes the answer from the prompt | The answer, compared with the prompt |
To the user, all three look the same. Each one is a confident wrong answer. This article shows how to tell them apart, and how to measure each one so the next failure is caught before a customer sees it.
We’ll use the same question throughout this article.
“Can I return a refurbished WH-1000 I bought from the TechNova Outlet 10 days ago?”
In this example, the return policy is split into two chunks. C1 says new products can be returned within 15 days of delivery, in the original packaging. C2 says refurbished products bought from the TechNova Outlet have a 7-day window. The correct answer is no, and it needs only C2. C1 is related, because it’s the general rule, but it can’t answer this question.
Start with the Diagnostic Spine
When one answer is wrong, check the three layers in order, and stop at the first one that failed.
- Are the right chunks retrieved? Look at the retrieved chunks and their ranks.
- Did they reach the model intact? Look at the final assembled prompt.
- Did the model use them correctly? Read the answer next to the prompt.

Checks 2 and 3 only matter once check 1 passes. If the right chunks never came back, no prompt change or model upgrade can fix the answer.
| Layer that failed | Typical fixes |
|---|---|
| Retrieval | Chunking, ranking, search method |
| Assembly | Token budget, prompt template, chunk order |
| Generation | Instructions, or a model that follows them more closely |
Here is why the order matters. Our assistant answered “Yes, within 15 days.” That one wrong answer can come from any of the three layers. The diagram follows C2 through each case.

In our case it was cause 1, because C2 ranked 4th. The model answered from C1, the rule for new products. Only the trace can tell the three cases apart, which is why you check them in order.
Two definitions keep the spine precise.
- Retrieval here means the chunks passed to assembly, after any ranking or filtering.
- The spine tells you where a failure shows up. The fix can sit further upstream. If the right content isn’t in the index at all, the problem is ingestion (Part 3). If the index holds an old version of a document, it shows up as a retrieval problem, and the fix is re-indexing (Part 8).
Symptoms point to a likely layer before you open the trace.
| What the user sees | Likely layer | First thing to check |
|---|---|---|
| “It says it doesn’t know, but the answer is in the docs.” | Retrieval, sometimes generation | Is the needed chunk in the results? If yes, did the model refuse anyway? |
| A detailed, confident, wrong answer | Retrieval, then assembly, then generation | Run all three checks in order |
| A vague, hedging answer (“the policy may vary”) | Retrieval, because the chunks are too broad | The retrieved chunks. Are they specific enough? (Part 4) |
| A correct answer to a different question | Retrieval (nearby content), or generation (wrong chunk used) | Was the right chunk retrieved? If yes, the model picked the wrong one |
| Different answers to the same question over time | Stale and current chunks both in the index | Look for two versions of the same document (Part 8) |
The spine is for one wrong answer. When many answers get worse after a change, start from what changed and from your measurements. That’s the rest of this article.
From One Answer to Many
The spine debugs one answer. Measurement tells you whether a change helped or hurt across many. You need both. The spine finds a cause, and metrics prove a fix worked without breaking something else.
Each layer gets its own measurement.
Retrieval Metrics
Retrieval metrics answer two questions. Did we find the chunks we needed, and did we rank them high enough to be used?
All of them need the same input, which is a set of test questions and, for each one, the chunks that should come back. (Building that set comes later.)
Here is our example before and after a fix. The system puts the top 3 chunks in the prompt, and the only chunk this question needs is C2.

The metrics put numbers on the picture.
| Metric | Plain question | Before | After |
|---|---|---|---|
| Recall@3 | Of the chunks we needed, how many are in the top 3? | 0 of 1 = 0% | 1 of 1 = 100% |
| Precision@3 | Of the top 3, how many did we need? | 0 of 3 = 0% | 1 of 3 = 33% |
| Reciprocal rank | How high is the first needed chunk? (1 divided by its rank) | Rank 4, so 0.25 | Rank 1, so 1.0 |
| Recall@10 | Of the chunks we needed, how many are in the top 10? | 1 of 1 = 100% | 1 of 1 = 100% |
Recall and precision fail in different ways. Low recall means something you needed is missing, so the model can’t answer correctly. Low precision means you got noise along with it, so the model has to read past chunks that don’t matter. Missing content is worse than noise, which is why you check recall first.
Mean reciprocal rank (MRR) is the reciprocal rank averaged over all your test questions. It tells you how high the first needed chunk usually lands. It only looks at the first needed chunk, so for questions that need several chunks, rely on recall.
Compare recall at two depths to see what kind of problem you have, as the diagram shows.
- If recall@10 is low, the chunk isn’t being found. Change chunking, add hybrid search or query expansion, or test another embedding model.
- If recall@10 is high but recall@3 is low, the chunk is found but ranked too low. Add a reranker, send more chunks to the prompt, or add hybrid search. Each one costs something (Part 4).
- If recall@3 is high but precision is low, the right chunk is there, with some noise. That’s usually fine, as in our “after” list.
Assembly Checks
Between retrieval and the model, check that what was retrieved is what the model received.
Assembly fails quietly, because the retrieved chunks look right in the logs. Three checks catch most of it.
- Every retrieved chunk appears in the final prompt. Compare the chunk IDs retrieved with the chunk IDs in the prompt. Any difference is a dropped chunk.
- Truncation is logged. When the prompt is cut to fit the token budget, record which chunks were cut. Conversation history can also crowd out retrieved chunks as a chat grows.
- Chunk order is logged. Models don’t read long context evenly, so order can change which chunk the answer relies on.
Cause 2 in the three-causes diagram is this failure. C2 ranks 3rd, the two chunks ahead of it are long, and the 600-token budget cuts it. Recall@3 is 100%, yet the answer is still “15 days.” Only check 1 here catches it.
Generation Metrics
Once the right chunks reach the model, three questions decide whether the answer is good.
- Faithfulness asks whether each statement in the answer can be traced to the prompt.
- Answer relevance asks whether the answer responds to the question that was asked.
- Completeness asks whether the answer includes the details the user needs, such as which rule applies and why.
Here are four answers to our question, each from a prompt that contains both C1 and C2.
| Answer | Verdict | Which check catches it |
|---|---|---|
| “No. Refurbished products bought from the TechNova Outlet have a 7-day return window, so 10 days is too late. New products have 15 days.” | Correct | All three pass |
| “No, refurbished Outlet items have a 7-day window. You can also exchange it within 30 days.” | Wrong | Faithfulness. No chunk mentions a 30-day exchange. It may sound like a typical retail policy, but it isn’t TechNova’s. |
| “The WH-1000 is covered by TechNova’s warranty against manufacturing defects.” | Off-topic | Answer relevance. True, but it answers a different question |
| “No, it’s too late to return it.” | Correct but incomplete | Completeness. The conclusion is right and nothing is invented, but it leaves out the 7-day refurbished rule, so the customer can’t see why or check it |
The second row is a hallucination, a claim the context doesn’t support.
The answer from the opening, “Yes, within 15 days”, isn’t in this table on purpose. It isn’t incomplete; it’s wrong. If C2 was missing from the prompt, it’s a retrieval or assembly failure. If C2 was in the prompt and the model still said 15 days, the model ignored relevant evidence, and check 3 of the spine catches it.
A quick test shows whether the model uses the context at all. Swap the retrieved chunks for unrelated ones and ask the same question. If the answer barely changes, the model is answering from what it already knows, not from your documents. The usual fix is in the instructions, for example “Answer only from the documents provided. If they don’t contain the answer, say so.”
Lowering the temperature makes wording more consistent. It rarely fixes an unsupported claim.
Two notes on names. Scores that count word overlap with a reference answer, such as BLEU and ROUGE, come from translation and summarization. They can’t tell a grounded answer from a fluent wrong one, so this article doesn’t use them. And frameworks such as RAGAS use some of the same metric names with different formulas. Some weight precision by rank, for example. Expect the idea to match, not the number.
Start by Reading Answers
Before you choose metrics, read real answers and write down what went wrong. Generic metrics give you a starting point. Reading your own failures tells you what your system actually needs measured.
- Collect 30 to 50 real questions and answers, with their traces, from logs or a test run.
- For each wrong answer, write one line about the first thing that went wrong. Use the spine to name the layer, then say what happened. Note only the first failure, because later problems often follow from it.
- Group the notes and count them. The biggest groups are what to measure and fix first.
Here are example notes from TechNova.
| Note | Layer |
|---|---|
| Refurbished rule ranked below the cut-off | Retrieval |
| Firmware changelog from last year retrieved instead of the current one | Retrieval (stale index) |
| Long warranty chunk pushed the returns chunk out of the prompt | Assembly |
| Added a 30-day exchange that no chunk mentions | Generation |
| Refurbished rule ranked below the cut-off (Outlet headsets) | Retrieval |
Two of five notes are the same problem. That group becomes your first evaluation questions and your first metric to watch, which is recall@3 on conditional policy questions.
Building an Evaluation Set
An evaluation set is a fixed list of questions with known answers that you re-run after every change. Every metric in this article needs one.
Start with 20 to 50 questions, written by hand from the groups you found. For each one, record the question, the facts a correct answer must contain, and the chunks that should be retrieved.
- id: R07
question: "Can I return a refurbished WH-1000 I bought from the TechNova Outlet 10 days ago?"
required_facts:
- "Refurbished Outlet products have a 7-day return window"
- "10 days is past that window"
expected_chunks: [returns-policy#refurbished]
A known-good answer is a list of facts, not exact wording. Any answer that contains the required facts and nothing the context doesn’t support passes, however it’s phrased.
Mix question shapes, because each one tests a different part of the pipeline.
| Question shape | TechNova example | What it tests |
|---|---|---|
| Simple lookup | “What is the warranty period on the WH-1000?” | Finding one obvious chunk |
| Condition or exception | “Can I return a refurbished WH-1000 after 10 days?” | Ranking the exception chunk, and the model honoring it |
| Combines two chunks | “What’s covered under warranty if I bought it refurbished?” | Recall across two documents |
| Version-sensitive | “What does firmware v3.2 fix?” | Whether the index holds the current version |
A few questions of each shape find more failures than fifty versions of one shape.
Five rules keep the set useful.
- Curated first, generated second. Tools such as RAGAS can generate questions from your documents. Use them to add coverage, and have a person check the expected answers.
- Keep it outside the corpus. If expected answers sit in indexed documents, retrieval can find them, and the test measures nothing.
- Update it when documents change. A new policy or firmware version changes the expected answers. A stale evaluation set gives false failures, and worse, false confidence.
- Remap after a chunking change. Chunk IDs change when chunking changes. Keep the required facts and source sections as the stable target.
- Run it after every change. Compare recall before and after. If recall dropped, you lost something. If precision dropped, you added noise.
Common starting points include RAGAS, LangSmith, and the evaluation features in cloud platforms such as Amazon Bedrock and Vertex AI. The habits above apply to all of them.
LLM-as-a-Judge
A judge is a second model that checks answers for you, so you can run the generation checks on every evaluation question, not just the few you read.
You give the judge the question, the context the model received, and the answer. It returns PASS or FAIL for one check, with a one-line reason.

A faithfulness judge prompt is short.
You are checking an answer against the context it was given.
Question: {question}
Context the model received: {context}
Answer: {answer}
PASS if each statement in the answer can be traced to the context.
FAIL if any statement is absent from the context or contradicts it.
Return PASS or FAIL, then one sentence naming any unsupported claim.
Why PASS or FAIL instead of a 1 to 5 score? Take the answer “No, refurbished Outlet items have a 7-day window. You can also exchange it within 30 days.” It contradicts nothing, but it adds an unsupported claim. On a 1 to 5 scale, a judge might give it a 2 on one run and a 4 on the next. Against “FAIL if any statement is absent from the context”, the decision is clear. One clear decision per answer also makes it easier for people to agree, and for you to check the judge. Write a separate judge for each check, one each for faithfulness, relevance, and completeness.
Check the judge before you trust it.
- Label 30 to 50 answers yourself, PASS or FAIL.
- Run the judge on the same answers.
- Count two things. First, how many of your FAILs the judge caught. Second, how many of your PASSes it wrongly failed.
If both numbers aren’t high, for example 90 percent or more, improve the judge prompt with clearer definitions and a few PASS and FAIL examples, and check again. Re-check whenever you change the judge prompt or model.
Judges fail in predictable ways.
| Problem | What to do |
|---|---|
| Prefers longer answers | Judge against a written PASS/FAIL definition, not “which is better” |
| Prefers the first answer when comparing two | Randomize the order, or run both orders |
| Prefers answers from its own model family | Use a different model as the judge than the one that wrote the answer |
| Noise on a small set. On 30 questions, one flipped verdict moves the score by about 3 points | React to changes that last across runs, and read the questions that flipped |
A judge can check thousands of answers in the time a person checks ten. It still misses subtle unsupported claims. Use it to watch trends, and use people to build the evaluation set and investigate what the judge flags. Thumbs-down feedback in the product is the cheapest source of new evaluation questions.
Apply This Tomorrow
- Log the retrieved chunks and the final prompt for every request. Without them, you can’t run the spine.
- Read 30 real answers. Note the first failure in each, then group the notes.
- Turn the biggest groups into 20 evaluation questions, with required facts and expected chunks.
- Run them after every change. Watch recall for retrieval and PASS rates for answers.
- When one answer is wrong, run the spine. Check the chunks, then the prompt, then the answer.
Keep Going
Hands-on lab. Prefer learning by doing? Try the RAG Debugging Lab, where several policy questions stopped finding the right document. Trace the evidence, find where it breaks, and prove your fix. Runs locally with Python.
Next, Part 8 covers what breaks after a RAG system reaches production, and how to run it safely over time.
References & artifacts
- RAG in Practice examples repository
- Evals for AI Engineers by Shreya Shankar and Hamel Husain (O’Reilly, early release). Further reading on error analysis, judges, and evaluating RAG.