Technical explanations · Drawbly
How to evaluate a RAG answer: check retrieval and the answer separately
By Drawbly ·
A RAG answer can be wrong in two different places. Search may miss the passage that contains the rule, or the answer model may have the right passage and still misstate it. One small test case can show which happened.

Start with a question and a verified answer
Use a fictional API retry guide, version 2. Its 400 Bad Request section says, “Fix the request; do not retry it unchanged.” A separate 429 Too Many Requests section permits a delayed retry up to three attempts. Our test question is: “Should the client retry an unchanged HTTP 400?” Before running any model, a reviewer records the expected answer: “No. Fix the request first.” The expected supporting source is the current guide's 400 section, with its document version.
The chunk-boundary guide shows how splitting this same fictional source can detach a status code from its instruction. Here we assume the complete source exists in the index and inspect what retrieval and answer generation did with it. A vector search explanation covers how candidate passages are found.
Run the same question through two failure cases
- Run A: the 400 section is missing. The top retrieved context contains the 429 retry passage but not the 400 instruction. If the answer says “retry up to three times,” the first problem is retrieval coverage. Check query wording, filters, chunk boundaries, index freshness and the returned passage IDs. Changing only the answer prompt cannot supply a rule that search omitted.
- Run B: the 400 section is present. The retrieved context includes “do not retry it unchanged,” but the generated answer still says “retry up to three times,” perhaps borrowing from the adjacent 429 section. Retrieval found the required evidence; inspect the answer prompt, how passages were assembled, the actual model input and claim-to-source matching. Even a citation pointing to the 400 section would not make the contradictory claim correct.
These are two imagined outputs, not a benchmark of a particular model or search service. A real run might fail in both places. Preserve the question, expected source, retrieved chunk IDs and versions, final answer and cited passages for each test, while keeping private document text out of public analytics.
What should you check separately?
Retrieval check: did the current 400 section appear among the passages actually supplied to the answer model, with the “do not retry” condition intact? For this one question, mark yes or no and record its rank. A relevant 429 passage alone does not cover the expected answer.
Answer check: does the answer reject retrying an unchanged 400, explain the required fix, and cite the 400 section rather than a neighboring rule? Check factual correctness against the verified source and whether each claim is supported by the cited context.
Do not collapse these into one “RAG accuracy” number. Amazon Bedrock's RAG evaluation metrics separate retrieve-only context relevance and coverage from retrieve-and-generate correctness, faithfulness and citation checks. Microsoft's end-to-end evaluation guidance likewise recommends examining grounding data and the generated response. Their automated scores are useful signals, but a reviewer still needs to verify important source rules and inspect disputed cases.
Turn one example into a useful test set
Add paraphrases such as “Can I resend the same 400 request?” and a neighboring positive case, “When can I retry a 429?” Include an unanswerable question so the system can show that it lacks evidence. For each case, write the expected source section and acceptable answer criteria before seeing model output. Include changed document versions and access rules where those matter to the application.
When a test fails, label the earliest broken boundary: source missing or stale, relevant passage not retrieved, passage omitted from the model input, or answer unsupported by the supplied passage. Then change one part of the system and rerun the same cases. A single passing example proves only that this example passed under those settings; it cannot establish overall RAG quality.
Draw the failed boundary in your system
Open the editable wide drawing and replace the fictional 400 rule with a real, approved source. Put the expected passage next to what search returned, then write the actual answer below it. Mark the first mismatch. The portrait drawing can be opened from Drawbly's Files menu. Drawbly helps explain a test case; it does not run a RAG evaluation or verify your source.
Amazon Bedrock: RAG evaluation metrics and Microsoft Learn: RAG end-to-end evaluation. The API guide, retrieved passages and model outputs in this article are fictional.