How Much Does AI Cost Per Month? Run Rate by Product Shape

The feature is scoped, the model is picked, and someone in finance wants a monthly number before the build starts. The vendor’s pricing page times an estimated token count produces one. It is wrong by the end of the first quarter, because the estimate froze two of the three inputs that move and the token price was never the one that mattered.

Ask how much AI costs per month and the honest answer is a product of three numbers, only one of which is on any pricing page: active users, requests per active user, and cost per request. On the production AI products I have measured, the first two moved more in six months than model prices moved in two years.

The direct version: monthly AI cost is active users times requests per active user times cost per request, plus the line items outside the model bill (retrieval, evals, provisioned capacity). For a product in production, the second input doubled within two quarters of launch and the third moved by more than half after a prompt change, so a monthly budget set at launch is half the real number by year end.

How much AI costs per month: the three inputs

InputWhere it comes fromHow it moved in production
Active usersProduct analytics, the definition finance sharesActivation climbs after launch; a feature at 20 percent activation can be cheaper per seat than one everyone uses
Requests per active user per monthUsage export joined to analyticsDoubled within two quarters of shipping
Cost per requestTokens, cache, retries, tool calls, from the usage exportMoved by more than half in a quarter on the same model after a prompt change

The estimate most teams make holds cost per request at the pricing-page number and requests per user at whatever the pilot showed. Both are the wrong number by the first invoice.

Monthly AI cost by product shape

Same profile for each: what drives the monthly number, the share of cost per request that is not tokens, and the caveat. Shares are from production, rounded to the range.

Chat assistant (conversational AI, copilot chat)

An output problem. Output tokens ran 45 to 60 percent of cost per request, and output length crept up every quarter nobody capped it. The monthly number is driven by session length: products that replay full history saw tokens per session grow quadratically with turn count.

Caveat: the median session is short and the p95 session is where the monthly cost lives.

Summarizer (document summarization, meeting notes)

An input problem. Input tokens ran 55 to 70 percent of cost per request, so the monthly number follows document length, not user count. Caching helped almost nothing because each document is read once.

Caveat: the prompt you tested with is shorter than the documents your users upload.

RAG search (retrieval-augmented generation, enterprise search)

A fixed-cost problem. Retrieval was under 5 percent of cost per request but a large share of the fixed monthly infrastructure: the embedding refresh, the vector database, the reranker. Monthly cost has a floor that does not move with usage.

Caveat: re-embedding on every index refresh is the line that surprises finance.

Agent or copilot (multi-step, tool-using)

A call-count problem. One user request fans out into up to ten model calls; tool calls plus retrieval ran 15 to 30 percent of cost per request and retries another 10 to 20. Cost per request is a distribution, so the monthly number depends on the p90 run, not the mean.

Caveat: a workflow that gets more capable next quarter makes more calls per run, and the monthly bill moves with no change in users.

Classifier or extraction (structured output, tagging)

A volume problem. Cheap per request, run on everything. Monthly cost follows event volume rather than user count, and it is the shape most likely to move down a model tier without a quality loss.

Caveat: moving down a tier raised retry rates on structured output, which is how the cheaper model cost more per resolved task.

The line items outside the model bill

Retrieval infrastructure (vector database, embeddings)

Sits in a different account from the model, which is how it goes missing from the monthly number. Fixed and usage components both.

Caveat: cheap per query, expensive per refresh.

Eval overhead (LLM-as-judge, quality monitoring)

A second model grading the first. Sampled at 5 to 10 percent of traffic it stayed at a few percent of cost per request; run on every request it nearly doubled it.

Caveat: charged to the platform team, so it never lands on the product’s number unless someone allocates it.

Provisioned throughput (PTU, reserved capacity)

A fixed monthly commitment shared across applications. At low utilization, a request on a provisioned pool ran two to three times the pay-per-token equivalent; at high utilization, below it.

Caveat: the monthly bill looks flat while usage climbs, then steps when the pool is resized.

Fine-tuning and embedding refresh

Periodic rather than monthly, and they land as spikes. Training runs and inference traffic, not the surrounding data pipeline, were the dominant drivers on the platforms I have measured.

Caveat: a training run scheduled by an engineer at 3 a.m. can cost more than a month of inference before anyone is awake.

Observability and logging

Prompt and response logging at scale is a storage bill. Small next to inference, large next to the estimate that forgot it.

Caveat: retention policy is a cost decision nobody makes on purpose.

A worked example, for shape only

Illustrative numbers, not vendor prices; the point is the arithmetic. Take a chat assistant at two cents per request, 60 requests per active user per month, and 5,000 active users: 6,000 dollars a month at the token estimate, 7,800 to 10,800 once the ride-alongs (retries, tool calls, retrieval, evals) add their 30 to 80 percent. Now let requests per active user double, as it did in production: 15,600 to 21,600. Same users, same model, same price sheet. The LLM cost calculator runs this with your own inputs at three usage tiers.

Pay-per-token versus provisioned: which monthly bill you get

Pay-per-token gives a monthly bill that tracks usage, which finance can pace. Provisioned capacity gives a flat bill that tracks nothing until the resize. Below steady, predictable throughput, pay per token. Above it, provision and allocate the pool by throughput share so each feature sees its slice.

Pre-build versus in production: where the monthly number comes from

Before launch the number comes from a model: token split, requests per user, an assumed cache hit rate. After launch it comes from the usage export joined to product analytics, the only version fit for a forecast. Swapping one for the other is how forecasts miss. The sibling piece on what one AI request actually costs covers the nine drivers behind the third input.

The mistake: budgeting the API bill

The mistake is taking last month’s vendor invoice, adding a growth percentage, and calling it the AI budget. It is too low, because retrieval, evals, and provisioned capacity sit on other bills, and it is flat where the real number is not, because it does not know that requests per active user is climbing. The symptom: a budget that is fine for two months and breached in the third with nobody surprised except finance. On a Fortune-500 financial data platform with eight figures of annual AI and cloud spend, the first feature-level breakdown found one of the top three AI products running several times the portfolio average; the invoice-based budget had no line for it.

The rule

Budget per active user, not per API bill, and re-forecast when requests per active user move by a quarter. Cost per active user is the one number that moves with adoption in a way the invoice cannot explain, and it is the one that turns into cost per subscriber when pricing comes up.

Model and cloud prices change monthly; the illustrative figures above are for arithmetic, not for planning, and any per-token price should be checked against the vendor’s live pricing page.

FAQ

How much does AI cost per month for a small business?

For a team building on an API rather than buying seats, monthly cost is active users times requests per user times cost per request, plus any retrieval or eval infrastructure. A chat feature for a few hundred users often runs in the hundreds of dollars a month at the token estimate; the ride-alongs add 30 to 80 percent, and usage per user tends to double within two quarters.

How much does it cost to run an AI chatbot per month?

A chat assistant is an output-token problem: 45 to 60 percent of cost per request was output on production products, and session length drives it. Monthly cost is requests per active user times cost per request times users, with the p95 session carrying most of the bill. Budget on the heavy user, not the median.

What is the cost of AI per user?

Cost per active user is cost per request times requests per active user per month; cost per subscriber is that times the share of subscribers who are active. It is the number to budget and price on, because it moves with adoption. On production products, requests per active user doubled within two quarters, so the per-user number at launch understates the year.

Why does our AI cost go up every month when traffic is flat?

Requests per active user climbing, a prompt template that grew, output length creeping, a cache hit rate that fell, or a retry rate that rose after a model change. None of them shows on an infrastructure dashboard, which is why the monthly number needs cost per request per feature underneath it.

What to do next

Put your three inputs into the LLM cost calculator at launch usage, double the requests per user, and look at the second number; that is the budget. When it needs to become a forecast finance runs alone, that is FinOps consulting for AI products, and the rest of the FinOps for AI practice is on the home page.


I put a number on what an AI product costs per request, per active user, and per subscriber, early enough to change the model, the architecture, or the price. If you’re shipping something with a model behind it and nobody can tell you what it costs, email me.

Have a number nobody can explain?

Send a note and I will get back to you.