Drawbly

Technical explanations · Drawbly

How vector search finds context for a RAG answer

By Drawbly ·

To use vector search, an embedding model turns a question and stored passages into numerical representations. A search index then retrieves passages with nearby embeddings. In a RAG system, the application gives those passages to a model as context for an answer. The search step finds candidates; it does not establish that a candidate is correct, current or allowed for this user.

A late parcel question is embedded and searched against policy passages. The delayed delivery passage is selected as a candidate, then its source and version are checked before it can support an answer.
One hypothetical support question. The diagram starts at retrieval; the earlier RAG pipeline guide shows when documents enter the index.

A question and three policy passages

Imagine a support assistant is asked, “My parcel was late. Can I get the shipping fee back?” Its policy collection includes: A, “Delayed delivery: shipping fee credit”; B, “Return shipping label for damaged items”; and C, “Late payment charges.” These are fictional passages, not a retailer's policy. A matches the question's intent even though the wording is different. B shares “shipping” and C shares “late,” but neither answers the whole question.

A keyword search can be strong when the user supplies an exact phrase, order code or product name. A vector search can surface A when “late parcel” and “delayed delivery” are represented as similar by the chosen embedding model. It returns a ranked candidate list, usually the top k passages. The ranking depends on the model, index, chunk boundaries and similarity metric; it is not a guarantee that A wins every time.

What goes into the vector index?

During preparation, split each policy document into passages that can stand alone, create an embedding for each one and store the vector alongside its readable text and source metadata. Keep a stable passage ID, policy version and link back to the document. At question time, embed the question with a compatible model and compare it with indexed vectors. A vector database or vector-capable search index performs this retrieval; the language model writes the eventual answer.

Chunk size changes what can be found. A passage that says “a shipping fee may be credited” without the preceding “only when the carrier confirms a delay” is easy to misread. Preserve the condition in the same retrievable unit or fetch neighboring context after the match. Amazon Bedrock's chunking guide explains fixed, hierarchical and semantic options; none removes the need to inspect actual retrieved passages.

When is vector search alone a poor fit?

Similarity can miss exact identifiers, negation and narrow exceptions. If the question names policy DLV-204, lexical matching may be essential. A hybrid query runs text and vector searches together and merges their rankings. Azure AI Search's vector overview describes this pattern. Its hybrid ranking guide explains one implementation using reciprocal rank fusion; other search systems can combine results differently.

Also apply the right metadata and authorization constraints. An old policy or another customer's private document might be semantically close. Similarity cannot decide whether that passage is current or permitted. Filter candidates to the appropriate tenant and version in the retrieval design, then verify the source before quoting it. Whether filtering happens before or after nearest-neighbor search depends on the system and configuration.

How do you test whether retrieval helped?

  1. Label a small question set. For real support questions, record which passage and version would justify an answer, including questions with no valid passage.
  2. Inspect the retrieved list. Check whether the supporting passage appears in the top k, whether a condition was cut off and whether stale or unauthorized passages appear.
  3. Inspect the answer separately. Confirm that each policy claim points to a passage that actually supports it. If retrieval is right but the answer is wrong, changing the vector index alone will not fix the generation step.

The worked example has no measured retrieval scores. It shows what to examine, not a benchmark or a claim that vector search always beats keywords. Microsoft's RAG design guide separates retrieval quality from answer quality.

Draw the search step in your system

Open the editable drawing and replace the example question and passages with your own. Label the embedding model, index, filters, source/version check and the point where context reaches the answer model. Keep the wide file or download the phone drawing for a vertical explanation. Drawbly is a drawing tool; this diagram does not run retrieval.

Technical references

Azure AI Search: vector search overview, vector relevance and ranking, hybrid scoring, and Amazon Bedrock: content chunking. Product details differ; the three policy passages above are illustrative.