AI token economics (i.e, AI tokenomics) connects the amount of text an AI system processes and produces to the real cost of running an application. It accounts for tokens used to prepare a knowledge base as well as the tokens consumed by model-backed retrieval and answer generation.
A token is a unit of text that a model can process. Depending on the model’s tokenizer (i.e., how a model splits text into tokens), a token may represent a complete word or only part of one. Spaces and punctuation can also affect the count. The same paragraph may therefore contain a different number of tokens when processed by different models.
In this article, we’ll trace token consumption through a RAG pipeline and turn that usage into a practical cost model.
AI token economics describes the relationship between model usage and the cost of delivering an AI feature. It begins with two measurements:
The first step in any inference is tokenization. Model providers use tokenization to convert blocks of text into a sequence of tokens where each token is represented by a unique identifier. OpenAI’s Tokenizer provides a visible example of this process. A user sees a sentence, while the model receives a sequence of tokens whose length depends on the text and the tokenizer being used.
It’s good to keep in mind that token count only keeps track of one side of the equation. We also need to consider the model rate and token type, as well as the fact that hosted services may charge different rates for input and output tokens. Those rates can also vary depending on where tokens are consumed in the pipeline. Embedding models usually have their own rates, while reranking may be billed by tokens or by request. Some services also distinguish cached input from text that must be processed again.
Lastly, volumetric pricing makes the unit costs add up to a budget. It’s easy to underestimate the value savings from reducing 200 input tokens in a single request, but the same reduction in requests made 1 million times eliminates 200 million tokens from the overall cost. Token economics allows us to visualize these cost savings before we launch an AI service to the general public.
In RAG, token consumption begins prior to the model generating a response. To estimate cost accurately, we need to account for ingestion first and then follow each request through retrieval and generation.
During ingestion, a RAG system extracts text from its source documents and divides that text into chunks. Each chunk is then passed to an embedding model, which converts it into a vector for semantic search. Embedding APIs accept tokenized text as input and report usage for the request, as shown in the OpenAI embeddings guide.
// example response from OpenAI's CreateEmbedding endpoint
{
"object": "list",
"data": [
{
"object": "embedding",
"embedding": [
0.0023064255,
-0.009327292,
.... (1536 floats total for ada-002)
-0.0028842222,
],
"index": 0
}
],
"model": "text-embedding-ada-002",
"usage": {
"prompt_tokens": 8,
"total_tokens": 8
}
}
The initial embedding cost grows with the amount of source text, while updates to the knowledge base introduce additional ingestion costs as new or changed content is embedded.
Parsing and chunking do not necessarily use model tokens. They still consume compute and storage, but those costs belong elsewhere in the system budget unless an LLM performs the work.
At query time, semantic search usually embeds the user’s question so it can compare the query vector with stored document vectors. That embedding is generally small, though it happens on every request.
The database lookup itself does not always use AI tokens. Keyword search and vector similarity can run without an LLM, although the database and search infrastructure still cost money. This is why “retrieval cost” and “retrieval token consumption” are related but not interchangeable.
Retrieved passages become tokens when the application places them in the prompt. Returning ten passages instead of four may have little effect on the search operation, yet it can materially increase the input sent to the generation model.
The generation request includes more than the user’s question. Its input can contain system instructions and retrieved passages as well as conversation history and tool definitions. The output tokens cover the answer and may also include structured tool arguments.
In a simple RAG flow, one retrieval step leads to one generation call. On the other hand, an agentic RAG flow may rephrase the question and search more than one source before it generates an answer. It may then call another model to check the result. Each call has its own input and output count, so the complete task would often cost more than the final response suggests.
RAG can reduce generation-side token consumption by retrieving a focused set of chunks instead of sending an entire document to the model. Only the passages most relevant to the question become prompt context, which lowers the number of input tokens the model must process for each answer.
For example, consider a 200-page PDF used as a knowledge source. Sending the entire PDF with every question would require the model to process all of that text each time. A RAG system might instead retrieve two paragraphs that are directly relevant to the question. This uses far fewer input tokens and gives the model a smaller, more focused set of information to work with.
The savings depend on how retrieval is configured. Returning more context than the answer needs can erase much of the benefit, while returning too little may weaken the result. We therefore need to tune the number and size of chunks against answer quality so that each passage sent to the model earns its place in the prompt.
Most hosted model providers publish separate input and output rates per million tokens. Once we calculate the total input and output tokens that we’ve consumed for a period of time (e.g., a month), we can apply the applicable rates:
monthly generation cost =
(input tokens / 1,000,000 x input rate)
+ (output tokens / 1,000,000 x output rate)
Ingestion is normally calculated separately because it follows the volume and update frequency of the knowledge base:
initial embedding cost =
document tokens / 1,000,000 x embedding rate
Suppose an application answers 100,000 questions each month. Across generation and validation with the same model, one completed request averages 3,180 input tokens and 240 output tokens. The monthly workload is therefore 318 million input tokens and 24 million output tokens. If the provider's per-million rates are represented by `I` and `O`, those calls cost `318I + 24O` before query embeddings or reranking are added.
This calculation separates the one-time cost of embedding the initial corpus from the recurring cost of serving questions. A frequently updated knowledge base also has recurring ingestion costs because changed documents must be processed again.
Token charges are only part of the production total because vector storage and database queries also have a price. Network transfer and monitoring may add another layer. If a provider charges for a tool call or managed search operation, that fee belongs in the same per-request model even though it’s not measured in tokens.
Self-hosting changes the above equation but not in the way one might expect. There may be no vendor invoice for each token, but the team pays for GPU capacity and power as well as deployment and operations. Dividing the allocated monthly hosting and operating cost by the number of tokens served gives us an effective token rate that can be compared with a hosted option. However, that rate will rise when expensive hardware sits idle and fall as utilization improves.
When it comes to managing token costs, the goal isn’t to minimize every prompt. Instead, we want to reduce unnecessary work while preserving the quality needed for the application. We can find a balance by keeping some of the points below in mind:
Every reduction should be tested against answer quality. Removing a relevant passage may lower the token count while increasing failed answers, which raises the cost of each successful outcome. The useful target is the least expensive configuration that still passes the application’s evaluations.
The Progress Agentic RAG solution handles ingestion and retrieval through Knowledge Boxes, then assembles the selected context for generation. Its token consumption documentation explains how the prompt and retrieved context contribute to input consumption. The user’s question is counted as input, while the generated answer contributes output tokens.
To receive consumption details from LLM-backed endpoints, we can include the X-SHOW-CONSUMPTION: true header. The response separates normalized input and output consumption from usage associated with a customer’s own model key. This gives teams a consistent measurement for comparing calls across supported providers.
AI token economics connects model activity to the cost of running an AI application. In a RAG system, that means measuring the initial corpus embedding separately from the query embeddings and model calls that recur with each question. It also means counting every call in an agentic workflow rather than looking only at the final answer.
For more details on measuring and controlling consumption, see the Progress Agentic RAG token consumption guide. You can also book a live demo or start a free trial to evaluate the pipeline with your own documents.
There is no fixed conversion because tokenization depends on the model and the text. As a rough guide for common English text, OpenAI estimates that 100 tokens represent about 75 words. Code and other languages can produce a different ratio, so applications should count tokens with the tokenizer used by their selected model.
No. A keyword lookup or vector database search can run without an LLM. Query embeddings and model-based reranking may consume tokens, and any retrieved passages sent to the generation model become input tokens. Teams should separate the cost of search infrastructure from the token consumption created by model-backed retrieval steps.
Self-hosting can eliminate a provider’s per-token fee, but it doesn’t make inference free. The organization still pays for hardware and power as well as the people and software needed to operate the model. Tokens remain useful for measuring throughput and calculating an effective cost for each workload.
Subscribe to get all the news, info and tutorials you need to build better business apps and sites