New:AI replies grounded in your own documentation.
All articles
Engineering14 July 20269 min read

Building a RAG pipeline that actually answers the question

Dense vectors alone miss the exact phrasing customers use. Here's how we combined semantic search, keyword search and reciprocal rank fusion — then added a reranker — to get support answers grounded in the right paragraph.

The first version of our retrieval was the one everybody builds: embed the article chunks, embed the question, take the top five by cosine similarity, put them in the prompt. It demos beautifully. It also fails in a very specific, very embarrassing way — it is at its worst exactly when a customer is most annoyed.

Where a single vector search falls down

Three failure modes showed up almost immediately, and none of them are exotic.

  • Exact tokens. A customer pastes RATE_LIMIT_EXCEEDED or the name of a specific setting. Embeddings are built to blur surface form — that is the entire point — so the chunk containing the literal string doesn't reliably win.
  • Follow-ups. The second message in a conversation is usually a fragment: “and on the free plan?” On its own it means nothing, and its embedding is a cloud of noise.
  • Plausible neighbours. Ask about refunds and you get the cancellation policy, the billing FAQ and the returns page. All related, none of them the answer.

The tempting fix is a bigger model or a longer prompt. Neither addresses the actual problem, which is that the wrong paragraph is being handed over in the first place. Generation quality is capped by retrieval quality, and no amount of cleverness downstream recovers a passage you never fetched.

Step one: make the question searchable

Before anything is searched, we rewrite the customer's latest message against the conversation so far. “And on the free plan?” becomes “Is CSV export available on the free plan?” This is a small, cheap model call and it is the single highest-leverage part of the pipeline, because every later stage inherits its output.

It also has a pleasant side effect: the condensed question is a much better cache key and a much better thing to log when you're trying to work out why an answer went sideways.

Step two: search twice

We run a dense vector search and a keyword full-text search in parallel over the same corpus. The vector side handles paraphrase — “why am I being throttled” finding the rate limit docs. The keyword side handles the literal — error codes, flag names, endpoints, the words your product actually uses.

The obvious question is how to merge two ranked lists whose scores mean completely different things. Cosine similarity and text-search rank are not on a common scale, and normalising them is guesswork that quietly breaks whenever your corpus changes.

Step three: fuse by rank, not by score

Reciprocal rank fusion sidesteps the problem entirely. Each result contributes 1 / (k + rank) from each list it appears in, with k = 60, and the fused score is the sum. Nothing depends on the magnitude of either score — only on position.

score(doc) = Σ  1 / (60 + rank_in_list)

The behaviour this produces is exactly what you want. A passage that both searches like rises to the top. A passage only one search likes still survives into the shortlist, which is the whole reason for running two searches. And a passage neither ranks highly stays where it belongs.

Step four: rerank the shortlist

Fusion gives a good shortlist, not a good answer set. The last step scores each surviving candidate directly against the question — a cross-encoder pass, which is far more accurate than comparing two independently-produced embeddings, and affordable precisely because it only sees a handful of candidates.

This is where the plausible neighbours die. The billing FAQ looks close in embedding space and shares vocabulary with the question, so it clears both earlier stages; asked directly “does this passage answer this question”, it doesn't survive.

What reaches the model

Only the top passages after reranking, fenced in the prompt and explicitly marked as untrusted reference material — never as instructions. If you import a documentation site, you are importing text nobody on your team wrote line by line, and it will eventually contain something that looks like a command.

If retrieval comes back empty, the correct behaviour is to escalate. An agent that invents a refund policy costs more than one that says “let me get someone”.

What we'd tell you to do first

If you're building this and can only do one thing, condense the question. If you can do two, add keyword search alongside your vectors and fuse by rank. Reranking is a genuine step up in quality, but it's the third improvement, not the first — and none of it matters if your chunks are bad. Retrieval can only find what someone bothered to write down clearly.

Talk to us

Questions about this post?A person reads every message.

Get answers