RAG that works: chunking, retrieval, and reranking
Retrieval-augmented generation looks simple on a diagram: search your documents, paste the best matches into the prompt, let the model answer. In practice, a system built exactly like that often gives confident, wrong answers, and the model usually isn't the culprit. The search step returned the wrong text, and the model did its best with what it got.
That's the most useful thing to know about RAG. Quality is decided mostly upstream of the model, in how you split documents, how you search them, and how you check the results. This guide walks through those pieces in the order they appear in a pipeline.
The pipeline, end to end
A typical system has two phases. Ingestion happens offline: you collect documents from wherever they live, clean them by stripping navigation, boilerplate, and duplicated headers, split them into chunks, compute embeddings for each chunk (plus, optionally, a keyword index), and store everything with metadata such as source, title, section, date, and permissions.
Query time happens online. You take the user's question and possibly rewrite it, retrieve candidate chunks with vector search, keyword search, or both, optionally rerank the candidates with a stronger model, build a prompt containing the best chunks and the question, and generate an answer with citations.
Each step can fail on its own, which is why you'll want to measure them separately, as we'll get to below. If you're still deciding whether RAG fits your situation, long context or RAG is the place to start.
Clean before you chunk
Garbage in, garbage out applies harder here than almost anywhere. If your PDF extraction glues table columns together, or every page carries a header and footer that gets embedded thousands of times, retrieval quality suffers before you've made a single design choice.
Spend time on the unglamorous parts. Extract text with structure preserved, so headings stay headings and tables stay tables. Remove repeated boilerplate such as cookie banners, nav menus, and page numbers, and deduplicate near-identical documents, or old versions will compete with new ones. Record real metadata too: title, URL, last-updated date, product or version, language.
Then open twenty random chunks and read them. If they're hard for you to understand out of context, they'll be hard for the model too.
Chunking: the choice that matters most
A chunk is the unit you retrieve. Too big, and each chunk covers several topics, so the embedding is a blurry average and the prompt fills with irrelevant text. Too small, and a chunk lacks the context needed to make sense.
Some guidance that holds up in practice:
Split on structure first. Headings, sections, and paragraphs are natural boundaries. Splitting a document into fixed 500-character windows regardless of content cuts sentences and tables in half.
Aim for a moderate size. Somewhere from a couple hundred to several hundred tokens works for most prose. There's no universal best value, so treat it as a parameter to tune against your own test questions.
Use overlap sparingly. A small overlap between neighboring chunks reduces the chance that an answer straddles a boundary and gets lost. Heavy overlap bloats the index and returns near-duplicates.
Keep the heading trail. Prepend each chunk with its document title and section path: "Billing > Refunds > Annual plans." That gives both the embedding and the model the context that the chunk itself lacks.
Handle tables and code specially. Keep a table together with its caption and header row, and keep code blocks whole, with the surrounding explanation.
Consider adding chunk-level context. A technique Anthropic described as contextual retrieval has a model write a short note situating each chunk within its document ("This section is from the 2025 refund policy and covers annual subscriptions") and prepends it before embedding. It's a one-time indexing cost that helps chunks that would otherwise be ambiguous. Read the original write-up for the reported gains and the details.
Embeddings and vector search
An embedding model turns text into a vector, a list of numbers, such that similar meanings land near each other. Vector search finds chunks whose vectors sit closest to the question's vector.
It's excellent at meaning, weak at exact strings. A search for "error 0x80070005" might surface general troubleshooting text and miss the page containing that precise code, because the code carries little semantic weight. The same goes for product names, IDs, and rare acronyms.
A few practical points:
- Use the same embedding model for indexing and querying.
- Pick a model that suits your language and domain, and test more than one on your questions.
- Store vectors somewhere you can filter by metadata. For many projects, Postgres with the pgvector extension is enough, and it keeps your data in one place. Dedicated vector databases earn their place at larger scale.
- Re-embed when you change models, since vectors from different models aren't comparable.
Hybrid search: use both
Keyword search (often BM25) is the mirror image of vector search: great at exact terms, blind to synonyms. Combining them covers both failure modes, and it's one of the highest-value upgrades you can make.
The usual recipe is to run both searches, then merge the ranked lists. Reciprocal rank fusion is a simple and dependable way to do that:
def reciprocal_rank_fusion(rankings, k=60):
"""Merge several ranked lists of doc ids into one."""
scores = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)
candidates = reciprocal_rank_fusion([vector_results, keyword_results])[:50]
A document that ranks well in either list rises, and one that ranks well in both rises further. There's no score normalization to fuss over, which is why people like it.
Rerank before you prompt
First-stage retrieval optimizes for speed over many candidates, so it's rough. A reranker is a slower, more accurate model that reads the question and each candidate together and scores how well the chunk answers it. The pattern is to retrieve generously, say fifty candidates, rerank, and keep the top five to ten.
Reranking tends to help most when the top-k from the first stage contains the answer but not at the top, which happens often. It also lets you retrieve wide without stuffing the prompt, since only the best few survive. The cost is added latency and another component to run, so measure whether it earns its place with your questions.
Improve the query, not just the index
Users don't type queries the way documents are written, and a few cheap tricks close that gap. A model can rewrite a conversational follow-up ("what about for annual plans?") into a standalone query ("refund policy for annual subscription plans") using the chat history. You can generate two or three phrasings and merge the results, which helps when vocabulary differs between question and source. If the question mentions a product or a date, filter to that product or period before searching; extracting those filters with a model works well. And you can hand the model a search tool and let it try again with a different query when the first results look wrong, the agentic approach mentioned in context engineering.
Write the generation prompt carefully
Retrieval gets the right text into the prompt. The prompt then has to make good use of it.
- Tell the model to answer only from the provided sources, and to say so when they don't contain the answer.
- Label each chunk with an id and source, and ask for citations by id.
- Ask for supporting quotes for factual claims, then verify them programmatically by checking that the quoted text appears in the chunk.
- Put the question after the documents.
- Say what to do with conflicts: prefer the newest source, or flag the disagreement.
A short template:
Answer the question using only the sources below. Cite sources as [id].
If the sources don't contain the answer, say "I couldn't find that in the documents."
<sources>
<source id="1" title="Refund policy" updated="2026-01-12">...</source>
<source id="2" title="Billing FAQ" updated="2025-08-03">...</source>
</sources>
Question: {question}
For prompt structure more broadly, see prompt engineering that still works.
Measure retrieval on its own
This is where teams most often go wrong. They test the whole pipeline end to end, see a bad answer, and start editing the prompt, when the real problem was that the right chunk never got retrieved.
Evaluate retrieval separately:
- Build a test set. Write 50 or more realistic questions, each with the document or chunk that answers it. Real user questions beat invented ones.
- Measure recall@k. For each question, did the correct chunk appear in the top k results? Try k of 5, 10, and 20. If recall@20 is low, no prompt will save you.
- Measure ranking quality. Is the right chunk near the top, or buried? This tells you whether reranking would help.
- Change one thing at a time, such as chunk size, hybrid weights, or embedding model, and compare.
Then evaluate generation on top: given the right chunks, does the model answer correctly and cite properly? Splitting the problem this way turns "the RAG is bad" into "recall is 62 percent at k=10, and here are the misses." The general method in how to evaluate an AI agent carries over.
Common failure modes
| Symptom | Likely cause | Try |
|---|---|---|
| Right document exists, never retrieved | Chunk too broad or too small; vocabulary mismatch | Better chunking, hybrid search, query rewriting |
| Exact IDs and codes are missed | Pure vector search | Add keyword search |
| Answer mixes old and new policy | Stale or duplicate documents | Dedupe, filter on date, prefer newest |
| Confident answer with no support | Prompt allows guessing | "Only from sources," require quotes |
| Answer cites the wrong source | Poor chunk labels | Clear ids, verify quotes in code |
| Great on easy questions, bad on comparisons | Multi-hop needs many chunks | Retrieve more, agentic search, or long context |
| Slow responses | Large k, heavy reranker | Cache, reduce candidates, parallelize |
Access control and freshness
Two production concerns deserve early thought.
Permissions. If users can see different documents, filter at retrieval time using metadata stored with each chunk, and never rely on the model to withhold something it was given. Whatever reaches the prompt can appear in the answer.
Freshness. Decide how updates flow in: webhooks, a nightly crawl, or on-demand refresh. Track document versions so that deleted or superseded content actually leaves the index. Serving an answer from a policy that was retired last month is worse than serving no answer.
A sensible build order
If you're starting from scratch, this sequence gets you to something useful fastest:
- Clean the data and chunk on structure, with heading trails.
- Embed and store with metadata.
- Build a small test set of real questions.
- Measure recall@k on plain vector search. That's your baseline.
- Add keyword search and fuse. Re-measure.
- Add a reranker if the answer is in the top twenty but not the top five.
- Write the grounded prompt with citations.
- Add query rewriting for multi-turn chat.
- Add monitoring: log queries, retrieved chunks, and user feedback, and review the failures weekly.
Each step is small, and each one is measured. That discipline is what separates RAG that demos well from RAG that keeps working after the second week.