Skip to content

Long context or RAG? How to choose

By SunnyKumar Jonwal 9 min read

Every time a model ships with a larger context window, someone declares RAG dead. And every time, teams building real products keep using retrieval anyway. Both sides are partly right, because the two approaches solve overlapping but different problems.

Here's a working way to decide, without ideology. Start with your data and your questions, then work backward to the design.

The two approaches in one paragraph each

Long context means you put the material directly into the prompt and let the model read it all. No search step, no index. The model sees the whole handbook, contract, or codebase and answers from it.

RAG (retrieval-augmented generation) means you keep the material in an external store, search it for each question, and put only the most relevant pieces into the prompt. The model answers from those pieces. RAG that works covers how to build the retrieval side well.

Both end at the same place: some text in the prompt and a model that reads it. The difference is who decides what goes in, you up front (long context) or a retrieval system per query (RAG).

What long context does well

  • Synthesis across a whole document. "What are the recurring themes in these forty interview transcripts?" needs everything at once. Retrieval returns fragments, and themes live in the gaps between them.
  • No retrieval failures. If the answer is in the prompt, the model can find it. RAG can fail before the model even sees the right text, when the search misses it.
  • Simplicity. No chunking, embeddings, vector store, or index refresh to maintain. For prototypes and internal tools that's a real advantage.
  • Questions you can't predict. When you don't know what will be asked, having everything available beats hoping the search finds it.
  • Whole-picture reasoning about structure, such as how modules in a codebase depend on each other, or how a contract's clauses interact.

What long context does badly

The costs show up quickly. Reading 500,000 tokens on every question is slow and expensive, and although caching helps a lot when the same material is queried repeatedly (see prompt caching), the first call and any cache expiry still hurt. Accuracy can suffer too. Models can miss details in very long inputs, especially facts buried in the middle or several similar passages competing for attention, so "find one specific fact" tasks tend to do worse than when the relevant paragraph is handed over directly.

Then there are the hard limits. Your data might not fit: a company wiki, a decade of tickets, or a large monorepo can dwarf any window. If different users may see different documents, putting everything in the prompt breaks that boundary, whereas retrieval can filter by permission before anything reaches the model. And a big prompt assembled once goes out of date. Rebuilding and re-caching it is possible, but it's work.

What RAG does well

  • Scale. The corpus can be millions of documents. Only a few reach the prompt.
  • Freshness. Update the index when documents change, and answers reflect it. No retraining, no giant prompt rebuild.
  • Cost per query. Small prompts are cheap and fast.
  • Citations. You know exactly which chunks supported an answer, so you can link to sources and audit results.
  • Permissions. Filter by user, tenant, or classification at retrieval time.
  • Focus. Giving the model just the relevant passages often improves precision on factual questions.

What RAG does badly

Retrieval is a lossy step, and that's the dominant failure mode. If the right chunk isn't retrieved, the answer is wrong, and the model may not know anything is missing. Chunking destroys context as well: a paragraph gets split from its heading, a table is cut in half, a clause is separated from the definitions it relies on. Good chunking reduces this without eliminating it.

Multi-hop and global questions are awkward, because "compare policy A across all regions" or "summarize everything about project X" need many pieces at once. On top of that comes the machinery. Ingestion, embeddings, indexes, rerankers, freshness jobs, and evaluation all need building and maintaining, and chunk size, top-k, hybrid weighting, and query rewriting all affect quality without any universal best value.

A comparison at a glance

Question Long context RAG
Data fits comfortably in the window? Good fit Optional
Data is large or growing? Doesn't fit Good fit
Need to find one specific fact? Can miss in long inputs Strong if retrieval works
Need synthesis across the whole set? Strong Weak without extra work
Per-query cost High, lower with caching Low
Freshness Rebuild prompt Update index
Per-user permissions Awkward Natural
Citations Manual Built in
Engineering effort Low Medium to high

A decision guide

Work through these questions in order.

  1. How big is the corpus, in tokens? If it comfortably fits in the window with room for the question and answer, long context is on the table. If it's several times bigger than any window, you need retrieval of some kind. The dividing line isn't only the hard limit: prompts approaching the maximum tend to be slower, costlier, and less accurate than mid-sized ones.

  2. How often will it be queried? A one-off analysis of a single document is a natural fit for long context. A document queried thousands of times a day favors either RAG or long context with caching, so run the numbers for both.

  3. How does it change? Static material suits a cached long prompt. Data that changes hourly favors an index you can update incrementally.

  4. What do the questions look like? Pinpoint lookups ("what's the notice period in clause 14?") favor retrieval. Big-picture questions ("what are the biggest risks across this contract?") favor long context. If you see both, consider a hybrid.

  5. Who's allowed to see what? Any per-user or per-tenant restrictions push you toward retrieval with filters, or at least toward building the prompt per user.

  6. Do you need citations and audit trails? RAG makes these straightforward. With long context you can ask the model to quote passages, and you should verify the quotes.

  7. What can you afford to build and run? If you have a day, long context. If you have a team and a roadmap, RAG is a reasonable investment.

Four quick cases

A 60-page employee handbook, asked about by staff all day. That's perhaps 40,000 tokens, small enough to sit in the prompt. Put it in a cached system prompt and skip retrieval entirely. Simple, accurate, cheap after the first call.

A 200-page contract, reviewed once by a lawyer. Long context. The questions are about the document as a whole, it's queried a handful of times, and you want the model to see cross-references between clauses. Ask it to quote clause numbers and check them.

A support knowledge base with 8,000 articles, updated daily. RAG. It's too large for a window, changes constantly, and answers need citations. Add hybrid search and reranking.

A large monorepo. Neither in pure form. Coding agents typically search and read files on demand, which is retrieval done by the agent itself using tools such as grep and file reads. That's covered in the next section.

The hybrid patterns worth knowing

Most serious systems end up mixing the two.

Retrieve, then read generously. Use search to narrow the corpus to a few dozen relevant documents, then put whole documents, not tiny chunks, into a long-context prompt. You get the scale of RAG and the coherence of long context.

Agentic retrieval. Give the model search and read tools and let it decide what to fetch, iterating as it learns. This is the approach coding agents use, and it fits the just-in-time pattern in context engineering. It costs extra round trips and gains adaptability, since the agent can refine its query when the first attempt misses.

Cached core plus retrieved extras. Keep the essential, stable material (glossary, policies, schema) in a cached prefix, and retrieve the long tail per query.

Summarize, then drill down. Store hierarchical summaries: a short summary per document, and a section-level breakdown below it. Start at the top and expand where needed. This helps with big corpora and questions that need a global view.

Test it on your own questions

General advice only goes so far, and the honest answer for your data is empirical. A workable process:

  1. Collect 30 to 50 real questions, including hard and ambiguous ones, with reference answers.
  2. Build the simplest version of each approach that could work.
  3. Score correctness, citation accuracy, cost per query, and latency.
  4. Read the failures. Are they retrieval misses? Buried details? Contradictions in the sources?
  5. Fix the biggest failure category and rerun.

The technique in evaluating an AI agent applies here as well. One team's clear winner is another's disaster, because the answer depends on your documents.

Cost math you can do on a napkin

Numbers settle arguments faster than opinions, so estimate before you build. Take a corpus of 300,000 tokens and 2,000 questions a day.

With naive long context, each question re-reads the corpus: 2,000 × 300,000 = 600 million input tokens a day. Even at modest per-token prices that's a serious bill, and each answer waits on a very long prompt.

With caching, most of those reads become cheap cache hits, and the daily cost falls dramatically, provided the cache stays warm. Two thousand questions a day is roughly one every 43 seconds, so a short-lived cache stays warm through the working day. If traffic were one question an hour, it wouldn't, and the economics flip.

With RAG, each question sends perhaps 3,000 to 8,000 tokens of retrieved context: about 10 million tokens a day. That's sixty times less input than the naive approach, at the price of building and running the retrieval system.

Neither number is the answer on its own. What it shows is where the money goes: repeated reading of the same material is the expense, and both caching and retrieval attack it from different directions. Plug in your own volumes, current prices, and cache behavior, then decide. The arithmetic in tokens, context windows, and LLM costs will help.

A note on the "RAG is dead" argument

Windows keep growing, and some tasks that once needed retrieval now don't. That's real progress, and it's worth reevaluating your design when a new model arrives. But three constraints don't go away with bigger windows: cost, access control, and freshness at scale. As long as those matter, retrieval stays useful, whether it's a vector index, a keyword search, or an agent that runs grep.

The practical stance is to treat the window as a budget. Spend it deliberately. Pick the cheapest design that meets your accuracy bar, and be ready to change your mind when a new model arrives with a bigger window, a lower price, or better recall over long inputs, because any of those can tip a decision that looked settled six months ago. Sometimes that's a single cached prompt. Sometimes it's a full retrieval pipeline. Often it's something in between, and you'll only find out by measuring on your own documents with your own questions.