Somebody in your architecture review is reading your AI vendor’s post announcing its newest model. The headline number is the context window, a million tokens, enough to fit your whole document library in one question. Then comes the question: If the model can read everything, why still pay engineers to maintain your retrieval-augmented generation (RAG) layer?
That layer splits your documents into chunks, has an embedding model store each one in a searchable index, then hands the model only the chunks that match a question. Does that still earn its keep?
A window that big is genuinely useful. Feed the model your entire library on every question, though, and you stop controlling what an answer costs or how much you can trust it. And you’re stuck with your current vendor.
Start with what the model vendors get right. Gemini has offered million-token windows since 2024 and later two million. Claude introduced one in beta in 2025. Those numbers measure how much text fits. How much of it the model uses is a separate question.
The advertised size isn’t the number to design around. How much a model can hold and how much it can answer accurately from are different lengths. NVIDIA’s RULER test feeds a model steadily longer prompts and checks whether it can still find facts buried inside them.
The length where accuracy holds up, called effective context, is often a fraction of the advertised window. Stanford researchers mapped the failure, noting a model reads the start and end of a long prompt well and gets measurably worse in the middle.
A model never warns you when it has passed that point. Hand an agent a 100-page contract and it will quote pages 2 and 99 with total confidence, then talk straight past the indemnity clause on page 45. Nothing about the reply looks wrong. It just comes back missing the paragraph your legal team cared about.
Your engineers will argue accuracy. Your CFO will argue cost, and the money at stake is larger than most teams expect.
AI providers bill by the token, input and output alike, but the two approaches differ on input, that’s how much text rides along with every question. A user asks your agent something. Retrieval searches the index and sends only the passages that come back. A full-context prompt sends every document you own, then does it again on the next question. Google’s Gemini API pricing charges $1.25 per million input tokens up to 200,000 on Gemini 2.5 Pro and $2.50 above that. Send one percent of what you store and you move a hundred times fewer tokens, and the premium on long prompts pushes that closer to 200.
Prompt caching takes some of that back, when the provider keeps your text and charges far less when the same text arrives again. The price per token drops, but you still ship the entire library every time, and holding more of it costs more.
Now multiply that per-question difference by the turns an agent takes. Agents rarely answer in one shot: it plans, calls a tool, reads the result and plans again, resending the whole conversation each turn. That cycle is the agent loop, and after enough turns, the price difference becomes its own line item on the infrastructure bill.
Pro Tip: Price the agent loop as well as the model. Each turn resends the history, so by turn twenty a request has billed its opening context twenty times over. Caching gives much of that back, so estimate on the cached price.
Cost and accuracy you can measure and revisit. Vendor neutrality you decide once, when you pick an architecture, and a retrieval layer that pulls only the passages answering a question is what keeps the option open.
A RAG setup runs on two models, and they are not equally easy to replace. Swapping the model that writes the answer costs little. Swapping the embedding model, the one that turned your documents into searchable form, means rebuilding the index from scratch, because the numbers one produces mean nothing to another.
Caching deepens the commitment. Anthropic’s prompt caching takes roughly 90% off repeated reads, but only when the opening of the prompt is identical character for character, and only for about five minutes after its last use. Put a per-request user ID at the top and the discount disappears. Every vendor caches differently, so prompts built around one vendor’s discount have to be rebuilt for the next one.
That only matters if you have reason to leave, and the model market moves faster than any architecture gets revised. Z.ai’s GLM-5.3 ships a million-token window, and Kimi API Platform lists Kimi K3 at $3.00 per million input tokens. Neither was on anyone’s shortlist a year ago. The retrieval layer keeps that door open: it doesn’t care which model reads its passages. And the choice isn’t always yours. Procurement can write a competing vendor into a contract, with a deadline.
Staying free to switch vendors points to one design, in two steps. First, retrieval narrows a question down to the passages that answer it. Second, the model reasons across those passages, with the large window giving it room to do so.
One case makes that design the wrong choice: what you store is small and stable, and the answer connects clues scattered across all of it. Load the whole thing there, because retrieval can miss the right passage and hand back a confident wrong answer with no record of what it skipped. Most enterprise agents are the other case: a library that changes weekly and takes thousands of queries a day.
Everywhere else, retrieval can do better than raw chunks. RAPTOR summarizes groups of related chunks, then summarizes those summaries, so a broad question pulls one summary while a narrow one pulls the exact chunk.
What that two-step design gets you:
A change of AI vendor that costs a config edit and a round of answer-quality testing.
A record of which passages produced each answer, and a bill sized to what you sent.
That is why platforms like Progress Agentic RAG treat a bigger window as more room for what retrieval returns rather than a replacement for it.
Keep retrieval separate from the model that writes the answers, so either one can be replaced without touching the other. Note how a question goes in and what comes back, and have every passage record where it came from and which embedding model indexed it. When you swap models, those records tell you which indexes you have to rebuild and which you can keep.
The retrieval layer’s other job is access control. Put everything in the window and anyone who asks a question reaches anything the model can see. Your existing systems already track who may read what, so the retrieval layer should check those permissions rather than copy them.
None of that is real until you test it. Ask what would break if you changed AI vendors next quarter. Editing one setting in the retrieval layer is fine; re-processing every document you own and rechecking every answer your agents gave is not.
So when the next AI model announcement doubles the window again, the answer you give in that architecture review doesn’t change. The window sets how much the model can read at once. Your retrieval layer still decides what it reads, what that costs and whose model does the reading.
Fitting and reading accurately are two different things. The length a model handles reliably is often far shorter than its advertised window, so documents that fit can still be too long to answer from reliably.
With a retrieval layer, only two things change: the code that calls the model and the prompt formatting, both easy to test. If you rely on the full window, the breakage lands where you can’t schedule it, because every vendor caches differently and your prompt was tuned to how one model reads long text.
Founder & Principal Consultant
Subscribe to get all the news, info and tutorials you need to build better business apps and sites