The AI Cost Stack

The AI cost stack: three layers, not three departments arguing past each other. What each layer answers, where its number lives, and which article owns it.

Ask three people at the same company what the AI feature costs and you get three numbers, and none of them is wrong. The engineer says cost per request, built from tokens. Finance says the total on the cloud bill, or the spreadsheet next to it, spend with no drivers behind it. Whoever owns pricing says a margin estimate, built on the vendor’s pricing page times a token guess. Three honest answers to three different questions, and the reason they never reconcile in a meeting is that nobody has named the thing all three sit on: the AI cost stack.

The AI cost stack is three layers built on each other, not three departments with three dashboards. Cost per request is the unit, measured from tokens and usage. Pricing and margin is what that unit is worth charging for. FinOps for AI is the operating discipline that keeps both current as usage, models, and pools change. Skip a layer and the other two stop agreeing with each other, which is exactly the meeting above.

That’s the direct version: the AI cost stack answers one question that a cloud bill, a pricing page, or a spreadsheet answers alone. Cost per request tells you what one call costs. Pricing and margin tells you what to charge for it. FinOps for AI tells you whether either number still holds once usage, the model, or the provisioned pool changes under it.

The AI cost stack, in one table

Each layer answers a different question, and each already has its own deep dive on this site. This is the map, not a fourth argument.

LayerQuestion it answersWhere the number livesThe deep dive
Cost per requestWhat does one call, one active user, one subscriber cost?Token metering at the model gatewayWhat one AI request actually costs
Pricing and marginWhat do we charge, and what survives as margin?The unit cost run through a pricing model’s mathAI pricing models compared
FinOps for AIDoes either number still hold as usage, models, or pools change?Allocation, forecasting, and governance built on the two layers aboveFinOps for AI

Layer one: cost per request

Cost per request is the floor the other two layers stand on. It’s built from nine drivers, not one: input tokens, output tokens, cached tokens, retries, tool calls, retrieval, and the evals that ride along with a call, priced at the model’s rate rather than guessed from the vendor’s pricing page times an estimate. On the production AI products I’ve measured, those ride-alongs added 30 to 80 percent on top of the tokens-times-price number most launch budgets start from.

Two things make this layer worth measuring on its own before it feeds the next one. First, the drivers move independently: a prompt edit changes output length, which is a pricing event every time, whether anyone treats it as one or not. Second, the unit compounds upward. Cost per active user is cost per request times requests per user; cost per subscriber divides that by the share of subscribers who are active. Get the base unit wrong and both numbers above it are wrong by the same multiple.

What one AI request actually costs has the full driver breakdown and the ranges each landed at in production. The cheapest AI API is rarely the cheapest AI product and how much AI costs to run per month are the two spokes off it; the LLM cost calculator turns the driver list into a number for your own product.

Layer two: pricing and margin

Pricing and margin takes the unit from layer one and asks what it’s worth charging for. Every AI pricing model, per seat, usage or metered, credits, hybrid, tiered, or outcome-based, is a bet on how usage per seat grows, and the bet only pays off if the margin curve under it was checked against the unit cost before the pricing page went live. On the products I’ve measured, requests per active user doubled within two quarters of launch, which moved a flat seat price’s margin curve more than any change in model pricing did.

This is the layer where cost per request stops being an engineering number and becomes a pricing decision, and where the CFO’s version of the question (does this feature make money) gets an answer built on the same unit the engineer is already tracking, instead of a separate estimate that quietly drifts from it.

AI pricing models compared has all twelve models against the margin curve; the AI business case a CFO will sign is the spoke that turns the margin number into the five inputs a finance review actually asks for. Fractional product leadership for AI economics is the engagement built around owning this layer before the pricing page is public.

Layer three: FinOps for AI

FinOps for AI is the layer that keeps the first two honest as the product scales, and it’s the newest of the three, which is why it still needs defining. Cloud FinOps assumes allocation by tag, spend that tracks infrastructure, and elastic capacity; inference breaks all three, because a shared model endpoint serves five features at once, spend moves with usage instead of provisioning, and a provisioned throughput pool is fixed until someone resizes it. The fix at every one of those breaks is the same move: allocate by feature and by token instead of by resource, so cost per request from layer one and the margin call from layer two are both still traceable once five teams share one endpoint.

On a Fortune-500 financial data company’s shared model layer, moving allocation from the account level to the request level was what made cost per request, cost per active user, and cost per subscriber exist for the first time, broken out by feature instead of buried in one line called inference. That’s the operating discipline this layer names: not a new set of numbers, but the reporting path that keeps layers one and two accurate once more than one feature and more than one team depend on them.

FinOps for AI covers where the cloud FinOps assumptions break and the fix for each. FinOps KPIs for AI products, cloud cost governance for AI spend, and what FinOps tools can’t see about AI spend are the three spokes off it. FinOps consulting for AI products is the engagement that builds all three layers, in order, for a team that doesn’t have them yet.

The mistake most teams make

The mistake isn’t picking the wrong layer. It’s mistaking the layer you own for the whole stack. Engineering ships cost per request and assumes pricing will pick it up. Whoever owns pricing sets a model and assumes finance’s forecast will hold under it. Finance builds a forecast off last quarter’s total and assumes the unit underneath hasn’t moved since. Each assumption is reasonable on its own, and none of the three people holding it is wrong about their own layer.

The symptom is a meeting where three people bring three numbers and none of them are talking about the same thing, and everyone leaves assuming someone else owns the gap. Nobody does, by default, because the stack itself isn’t anyone’s job until someone names it as one.

The rule

A cost number is only as useful as the layer it answers for, so say which of the three questions it answers before anyone acts on it. Cost per request answers what one call costs, nothing about what to charge. A pricing model answers what survives as margin, nothing about whether the unit under it has drifted. A FinOps allocation answers whether the number still holds at scale, nothing about whether the price was ever right. Handed across a table without that label attached, all three get treated as the same number, which is how the meeting above happens in the first place.

FAQ

What is the AI cost stack?

Three layers built on each other: cost per request (the unit, from token metering), pricing and margin (what that unit is worth charging for), and FinOps for AI (the operating discipline that keeps both current as usage and models change). Each layer already exists on this site as its own article; this piece is the map between them.

Is FinOps for AI the same thing as cloud FinOps?

No. Cloud FinOps assumes spend tracks provisioned infrastructure and can be allocated by resource tag. Inference breaks that assumption, because a shared model endpoint serves several features at once and spend moves with usage, not with what’s provisioned. FinOps for AI allocates by feature and by token instead, which is the fix cloud FinOps tooling doesn’t do by default.

What’s the difference between cost per request and cost per active user?

Cost per request is the base unit: what one call costs, priced from its tokens, retries, and tool calls. Cost per active user multiplies that by how many requests a typical active user generates in a period. One is a property of the model call; the other is a property of how the product gets used.

Which pricing model protects margin as usage grows?

Usage or metered pricing holds margin roughly flat because the customer absorbs the variable cost. Per-seat pricing’s margin falls as usage per seat grows, and goes negative past the crossover where a heavy user’s inference cost exceeds their seat price. The choice depends on whether the buyer will tolerate a variable bill, not on which model is generically “better.”

Do I need all three layers if the product hasn’t launched yet?

Layer one, cost per request, is buildable pre-launch from a token estimate and a planned architecture, which is what a pre-build cost model is for. Layer two needs at least a rough usage assumption. Layer three, the FinOps allocation and forecasting layer, matters once more than one feature shares infrastructure, which is usually after launch, not before it.

How does the AI cost stack relate to the FinOps Foundation’s FinOps Framework?

The FinOps Framework is the broader discipline the FinOps Foundation defines for cloud spend generally: allocation, forecasting, and governance capabilities that predate AI workloads. The AI cost stack is the version of that discipline specific to inference, where the unit is a request and a token, not a provisioned instance.

What to do next

Each layer of the stack has its own deep dive linked above, and the LLM cost calculator turns layer one into a number for your own product. When the gap is the reporting path that keeps all three layers current as the product scales, that’s FinOps consulting for AI products. Start on the home page for the rest of the practice this stack is built from.


I put a number on what an AI product costs per request, per active user, and per subscriber, early enough to change the model, the architecture, or the price. If you’re shipping something with a model behind it and nobody can tell you what it costs, email me.

Have a number nobody can explain?

Send a note and I will get back to you.