How much AI costs in a product
What you pay for when a product uses a language model: tokens, what one request is made of, five scenarios of a monthly bill, how to calculate your own and eight ways to cut it.
In short
The cost of AI in a product is made of two parts: the one-off work of building it and the running cost of the model, which is paid per token — pieces of text the model reads and writes. For a typical assistant answering from a knowledge base, one answer costs from a fraction of a cent to a couple of cents, and a thousand answers a day — from tens to several hundred dollars a month. The bill depends less on the model’s price list than on how the requests are built: a long instruction sent with every question, unneeded chat history and a flagship model where a fast one would do can multiply the cost several times. Caching, a fast model for simple tasks and short answers cut it back.
What you pay for: the basics
Seven facts that explain any bill for a language model.
- Unit of payment
- A token — about four characters of English text; prices are given per million tokens
- Other languages
- Text in many languages takes more tokens than the same text in English
- Input and output
- What the model writes costs about five times more than what it reads
- Cache
- A repeated beginning of the request is read many times cheaper
- Batches
- Non-urgent tasks sent in a batch cost about half
- Tiers
- A fast model is several times cheaper than a balanced one, a flagship — more expensive
- Who pays
- The owner of the API key — best when the keys are registered to the client
What one request is made of
A typical answer of an assistant on a knowledge base. The question itself is the smallest part of the bill.
| Part | Tokens | Repeats | How to reduce |
|---|---|---|---|
| Instruction to the model | 2 000 | in every request | cache, shorter wording |
| Found pieces of the knowledge base | 800 | different each time | fewer and shorter pieces |
| Question and history | 200 | grows during the dialogue | a summary instead of the whole chat |
| Answer of the model | 300 | different each time | a length limit, a format |
Five scenarios for one assistant
The same assistant, 1,000 answers a day, the request from the table above. Example prices per million tokens: fast model $1 in / $5 out, balanced $3 / $15, flagship $5 / $25; reading from the cache — 10% of the input price.
| Scenario | One answer | Month | Compared with the first |
|---|---|---|---|
| Flagship, no cache | $0.0225 | $675 | 100% |
| Balanced, no cache | $0.0135 | $405 | 60% |
| Balanced, instruction in cache | $0.0081 | $243 | 36% |
| Fast, no cache | $0.0045 | $135 | 20% |
| Fast, instruction in cache | $0.0027 | $81 | 12% |
How to calculate the bill for your product
-
1. Collect real requests
30–50 typical questions or documents — not invented ones, but from mail, chats and the CRM.
-
2. Run them through the prototype
The provider returns the number of input and output tokens with every answer.
-
3. Find the average
Average tokens in and out per request, and separately the share that repeats and can be cached.
-
4. Multiply by the volume
Requests per day × 30, with a margin for peaks and growth.
-
5. Compare tiers
The same set on a fast and a balanced model: if the quality is equal, the difference is pure saving.
-
6. Set a limit
A monthly spending limit in the provider’s account protects from loops and abuse.
8 ways to cut the bill
-
01
Cache the instruction
The repeating beginning of the request goes into the cache — the biggest saving with no loss of quality.
-
02
A fast model for simple work
Sorting, extraction and short answers rarely need more.
-
03
Route by difficulty
Simple questions go to the fast model, complex ones to the strong one.
-
04
Limit the answer
The answer is the most expensive part: a length limit and a clear format.
-
05
Fewer pieces of context
Three good pieces are cheaper and more accurate than ten average ones.
-
06
A summary instead of history
A long dialogue is replaced by a short summary, not sent in full every time.
-
07
Batches for non-urgent work
Product descriptions or nightly reports at half the price.
-
08
Ready answers to frequent questions
A question that was already answered is not sent to the model again.
What else costs money besides tokens
The model is the most visible line in the bill, but not the only one.
| Item | How it is paid | Share of the budget |
|---|---|---|
| Building the product | once, by hours of work | the largest at the start |
| Vectors for search | per token, once per document | small |
| Storage of vectors | inside the existing database | usually nothing extra |
| Server | monthly | small for cloud models |
| Review of the log | hours per month | decides the quality |
| Updating the knowledge base | as documents change | small if automated |
Common mistakes in AI budgets
-
Counting only the question
The instruction and the context are ten times longer than the question itself.
-
Forgetting the output price
A long answer costs more than the whole request that produced it.
-
No spending limit
One loop in an agent or a bot attack can eat a monthly budget in a night.
-
The whole chat with every message
The cost of each new message grows with the length of the dialogue.
-
Keys on the contractor
The client does not see the real spending and depends on someone else’s account.
-
Saving on the wrong thing
A cheaper model that answers wrongly costs more in lost customers than it saves.
Questions about the cost of AI
How much does an AI assistant cost per month?
The model for a thousand answers a day — from tens to several hundred dollars, depending on the tier, cache and length of requests.
What is a token?
A piece of text the model works with — on average about four characters of English text.
Why is the answer more expensive than the question?
Writing takes the model much more computation than reading, so output tokens cost several times more.
How much does caching save?
In the example above — 40% of the bill; the longer the repeating instruction, the bigger the saving.
Is a local model cheaper?
At large volumes of simple work — yes; at small volumes the server costs more than the API.
How to avoid a surprise bill?
A monthly limit at the provider, limits per user and an alert at half of the budget.
Who should own the API keys?
On the client: the spending is visible, and the product does not depend on the contractor’s account.
Online form
An AI budget
before the start
I calculate the running cost on your real requests before development starts, and build the product so the bill stays predictable. The keys are registered to you, and the model is paid to the provider directly. Tell me about the task — I answer within one working day.