Do Agents Still Need RAG with Million-Token Context Windows?

Million-Token Context Windows | Agentic RAG
by Adam Bertram Posted on September 30, 2026

Somebody in your architecture review is reading your AI vendor’s post announcing its newest model. The headline number is the context window, a million tokens, enough to fit your whole document library in one question. Then comes the question: If the model can read everything, why still pay engineers to maintain your retrieval-augmented generation (RAG) layer?

That layer splits your documents into chunks, has an embedding model store each one in a searchable index, then hands the model only the chunks that match a question. Does that still earn its keep?

A window that big is genuinely useful. Feed the model your entire library on every question, though, and you stop controlling what an answer costs or how much you can trust it. And you’re stuck with your current vendor.

What Does a Million-Token Window Actually Deliver?

Start with what the model vendors get right. Gemini has offered million-token windows since 2024 and later two million. Claude introduced one in beta in 2025. Those numbers measure how much text fits. How much of it the model uses is a separate question.

The advertised size isn’t the number to design around. How much a model can hold and how much it can answer accurately from are different lengths. NVIDIA’s RULER test feeds a model steadily longer prompts and checks whether it can still find facts buried inside them.

The length where accuracy holds up, called effective context, is often a fraction of the advertised window. Stanford researchers mapped the failure, noting a model reads the start and end of a long prompt well and gets measurably worse in the middle.

A model never warns you when it has passed that point. Hand an agent a 100-page contract and it will quote pages 2 and 99 with total confidence, then talk straight past the indemnity clause on page 45. Nothing about the reply looks wrong. It just comes back missing the paragraph your legal team cared about.

What Are the Unit Economics of a Million-Token Window?

Your engineers will argue accuracy. Your CFO will argue cost, and the money at stake is larger than most teams expect.

AI providers bill by the token, input and output alike, but the two approaches differ on input, that’s how much text rides along with every question. A user asks your agent something. Retrieval searches the index and sends only the passages that come back. A full-context prompt sends every document you own, then does it again on the next question. Google’s Gemini API pricing charges $1.25 per million input tokens up to 200,000 on Gemini 2.5 Pro and $2.50 above that. Send one percent of what you store and you move a hundred times fewer tokens, and the premium on long prompts pushes that closer to 200.

Prompt caching takes some of that back, when the provider keeps your text and charges far less when the same text arrives again. The price per token drops, but you still ship the entire library every time, and holding more of it costs more.

Now multiply that per-question difference by the turns an agent takes. Agents rarely answer in one shot: it plans, calls a tool, reads the result and plans again, resending the whole conversation each turn. That cycle is the agent loop, and after enough turns, the price difference becomes its own line item on the infrastructure bill.

Pro Tip: Price the agent loop as well as the model. Each turn resends the history, so by turn twenty a request has billed its opening context twenty times over. Caching gives much of that back, so estimate on the cached price.

Why Does Vendor Neutrality Still Favor Retrieval?

Cost and accuracy you can measure and revisit. Vendor neutrality you decide once, when you pick an architecture, and a retrieval layer that pulls only the passages answering a question is what keeps the option open.

A RAG setup runs on two models, and they are not equally easy to replace. Swapping the model that writes the answer costs little. Swapping the embedding model, the one that turned your documents into searchable form, means rebuilding the index from scratch, because the numbers one produces mean nothing to another.

Caching deepens the commitment. Anthropic’s prompt caching takes roughly 90% off repeated reads, but only when the opening of the prompt is identical character for character, and only for about five minutes after its last use. Put a per-request user ID at the top and the discount disappears. Every vendor caches differently, so prompts built around one vendor’s discount have to be rebuilt for the next one.

That only matters if you have reason to leave, and the model market moves faster than any architecture gets revised. Z.ai’s GLM-5.3 ships a million-token window, and Kimi API Platform lists Kimi K3 at $3.00 per million input tokens. Neither was on anyone’s shortlist a year ago. The retrieval layer keeps that door open: it doesn’t care which model reads its passages. And the choice isn’t always yours. Procurement can write a competing vendor into a contract, with a deadline.

What Is the Two-Step Reference Pattern?

Staying free to switch vendors points to one design, in two steps. First, retrieval narrows a question down to the passages that answer it. Second, the model reasons across those passages, with the large window giving it room to do so.

One case makes that design the wrong choice: what you store is small and stable, and the answer connects clues scattered across all of it. Load the whole thing there, because retrieval can miss the right passage and hand back a confident wrong answer with no record of what it skipped. Most enterprise agents are the other case: a library that changes weekly and takes thousands of queries a day.

Everywhere else, retrieval can do better than raw chunks. RAPTOR summarizes groups of related chunks, then summarizes those summaries, so a broad question pulls one summary while a narrow one pulls the exact chunk.

What that two-step design gets you:

A change of AI vendor that costs a config edit and a round of answer-quality testing.

A record of which passages produced each answer, and a bill sized to what you sent.

That is why platforms like Progress Agentic RAG treat a bigger window as more room for what retrieval returns rather than a replacement for it.

What Belongs in the Reference Architecture?

Keep retrieval separate from the model that writes the answers, so either one can be replaced without touching the other. Note how a question goes in and what comes back, and have every passage record where it came from and which embedding model indexed it. When you swap models, those records tell you which indexes you have to rebuild and which you can keep.

The retrieval layer’s other job is access control. Put everything in the window and anyone who asks a question reaches anything the model can see. Your existing systems already track who may read what, so the retrieval layer should check those permissions rather than copy them.

None of that is real until you test it. Ask what would break if you changed AI vendors next quarter. Editing one setting in the retrieval layer is fine; re-processing every document you own and rechecking every answer your agents gave is not.

So when the next AI model announcement doubles the window again, the answer you give in that architecture review doesn’t change. The window sets how much the model can read at once. Your retrieval layer still decides what it reads, what that costs and whose model does the reading.

FAQ

Does a Document Set That Fits in the Window Still Need Retrieval?

Fitting and reading accurately are two different things. The length a model handles reliably is often far shorter than its advertised window, so documents that fit can still be too long to answer from reliably.

What Actually Breaks When You Swap Generation Providers?

With a retrieval layer, only two things change: the code that calls the model and the prompt formatting, both easy to test. If you rely on the full window, the breakage lands where you can’t schedule it, because every vendor caches differently and your prompt was tuned to how one model reads long text.

Does This Mean Million-Token Windows Are Not Worth Paying For?

They are absolutely worth paying for. The model needs room to reason across retrieved passages, and a 4,000-token window may be too small once you count the instructions and history in every prompt. The trouble starts when the big window becomes your substitute for searching at all.

Adam Bertram

Founder & Principal Consultant

Adam Bertram is a 25+ year IT veteran, former Microsoft MVP, and self-employed consultant who helps organizations replace repetitive manual work with generative AI automation and agent-based workflows. He’s a successful blogger, consultant, trainer, published author and freelance writer for dozens of technology publications.
More from the author

Related Products:

Agentic RAG

Progress Agentic RAG transforms scattered documents, video, and other files into trusted, verifiable answers accelerating AI adoption, reducing hallucinations, and improving AI-driven outcomes.

Get in Touch

Related Tags

Related Articles

What Is AI Token Economics - Tokenomics?
AI token economics connects model usage to the real cost of running an AI application. In this article, we’ll trace token consumption through a RAG pipeline and explore how teams can estimate and manage those costs.

Hassan Djirdeh September 30, 2026
Why AI Pilots Don’t Ship: The Trust Ceiling No One Talks About
AI pilot failures cluster around integration and a learning gap, how the system connects to real data and real workflows, not how well it reasons. In this post, we'll discuss why so many AI pilots fail to pass the trust ceiling when the governance to defend them don't exist.
Why AI Costs Spike After the First Use Cases
In this blog, we take a look at why AI costs accelerate so quickly after initial implementation success and how a modular approach to Agentic RAG can transform isolated pilots into a scalable, sustainable foundation for enterprise AI.
Prefooter Dots
Subscribe Icon

Latest Stories in Your Inbox

Subscribe to get all the news, info and tutorials you need to build better business apps and sites

Loading animation