How much AI costs in a product

What you pay for when a product uses a language model: tokens, what one request is made of, five scenarios of a monthly bill, how to calculate your own and eight ways to cut it.

Artificial intelligence Updated

In short

The cost of AI in a product is made of two parts: the one-off work of building it and the running cost of the model, which is paid per token — pieces of text the model reads and writes. For a typical assistant answering from a knowledge base, one answer costs from a fraction of a cent to a couple of cents, and a thousand answers a day — from tens to several hundred dollars a month. The bill depends less on the model’s price list than on how the requests are built: a long instruction sent with every question, unneeded chat history and a flagship model where a fast one would do can multiply the cost several times. Caching, a fast model for simple tasks and short answers cut it back.

What you pay for: the basics

Seven facts that explain any bill for a language model.

Unit of payment
A token — about four characters of English text; prices are given per million tokens
Other languages
Text in many languages takes more tokens than the same text in English
Input and output
What the model writes costs about five times more than what it reads
Cache
A repeated beginning of the request is read many times cheaper
Batches
Non-urgent tasks sent in a batch cost about half
Tiers
A fast model is several times cheaper than a balanced one, a flagship — more expensive
Who pays
The owner of the API key — best when the keys are registered to the client

What one request is made of

A typical answer of an assistant on a knowledge base. The question itself is the smallest part of the bill.

PartTokensRepeatsHow to reduce
Instruction to the model 2 000 in every request cache, shorter wording
Found pieces of the knowledge base 800 different each time fewer and shorter pieces
Question and history 200 grows during the dialogue a summary instead of the whole chat
Answer of the model 300 different each time a length limit, a format

Five scenarios for one assistant

The same assistant, 1,000 answers a day, the request from the table above. Example prices per million tokens: fast model $1 in / $5 out, balanced $3 / $15, flagship $5 / $25; reading from the cache — 10% of the input price.

ScenarioOne answerMonthCompared with the first
Flagship, no cache $0.0225 $675 100%
Balanced, no cache $0.0135 $405 60%
Balanced, instruction in cache $0.0081 $243 36%
Fast, no cache $0.0045 $135 20%
Fast, instruction in cache $0.0027 $81 12%

How to calculate the bill for your product

  1. 1. Collect real requests

    30–50 typical questions or documents — not invented ones, but from mail, chats and the CRM.

  2. 2. Run them through the prototype

    The provider returns the number of input and output tokens with every answer.

  3. 3. Find the average

    Average tokens in and out per request, and separately the share that repeats and can be cached.

  4. 4. Multiply by the volume

    Requests per day × 30, with a margin for peaks and growth.

  5. 5. Compare tiers

    The same set on a fast and a balanced model: if the quality is equal, the difference is pure saving.

  6. 6. Set a limit

    A monthly spending limit in the provider’s account protects from loops and abuse.

8 ways to cut the bill

  1. 01

    Cache the instruction

    The repeating beginning of the request goes into the cache — the biggest saving with no loss of quality.

  2. 02

    A fast model for simple work

    Sorting, extraction and short answers rarely need more.

  3. 03

    Route by difficulty

    Simple questions go to the fast model, complex ones to the strong one.

  4. 04

    Limit the answer

    The answer is the most expensive part: a length limit and a clear format.

  5. 05

    Fewer pieces of context

    Three good pieces are cheaper and more accurate than ten average ones.

  6. 06

    A summary instead of history

    A long dialogue is replaced by a short summary, not sent in full every time.

  7. 07

    Batches for non-urgent work

    Product descriptions or nightly reports at half the price.

  8. 08

    Ready answers to frequent questions

    A question that was already answered is not sent to the model again.

What else costs money besides tokens

The model is the most visible line in the bill, but not the only one.

ItemHow it is paidShare of the budget
Building the product once, by hours of work the largest at the start
Vectors for search per token, once per document small
Storage of vectors inside the existing database usually nothing extra
Server monthly small for cloud models
Review of the log hours per month decides the quality
Updating the knowledge base as documents change small if automated

Common mistakes in AI budgets

  1. Counting only the question

    The instruction and the context are ten times longer than the question itself.

  2. Forgetting the output price

    A long answer costs more than the whole request that produced it.

  3. No spending limit

    One loop in an agent or a bot attack can eat a monthly budget in a night.

  4. The whole chat with every message

    The cost of each new message grows with the length of the dialogue.

  5. Keys on the contractor

    The client does not see the real spending and depends on someone else’s account.

  6. Saving on the wrong thing

    A cheaper model that answers wrongly costs more in lost customers than it saves.

Questions about the cost of AI

How much does an AI assistant cost per month?

The model for a thousand answers a day — from tens to several hundred dollars, depending on the tier, cache and length of requests.

What is a token?

A piece of text the model works with — on average about four characters of English text.

Why is the answer more expensive than the question?

Writing takes the model much more computation than reading, so output tokens cost several times more.

How much does caching save?

In the example above — 40% of the bill; the longer the repeating instruction, the bigger the saving.

Is a local model cheaper?

At large volumes of simple work — yes; at small volumes the server costs more than the API.

How to avoid a surprise bill?

A monthly limit at the provider, limits per user and an alert at half of the budget.

Who should own the API keys?

On the client: the spending is visible, and the product does not depend on the contractor’s account.

Online form

An AI budget
before the start

I calculate the running cost on your real requests before development starts, and build the product so the bill stays predictable. The keys are registered to you, and the model is paid to the provider directly. Tell me about the task — I answer within one working day.

Or write to [email protected]