Cloud Cost Governance for AI Spend: The Controls Finance Can Run Alone

The governance program was fine until the AI line item showed up. Budgets by account, a monthly review, a cost console finance could read on its own. Then a product with a model behind it shipped, the inference bill doubled while the instance count stayed flat, and the monthly review had a number with nothing underneath it.

Finance’s spreadsheet has the total and none of the drivers. The provider’s cost console stops at the API call, so it cannot say which feature made the calls. Engineering can pull the usage export but nobody in finance can run that query, so the answer arrives a week late and by email. Cloud cost governance built for compute governs the account. Inference spends by the feature, and the controls have to move with it.

The direct version: cloud cost governance for AI spend is the set of controls that let finance see, forecast, and act on inference cost without an engineer in the room: budgets with a pacing line, a driver-based forecast, anomaly alerts on unit cost, and an allocation that gives every feature its share of a shared model pool. On a Fortune-500 financial data platform with eight figures of annual AI and cloud spend, that set replaced the provider console as finance’s reporting path.

Cloud cost governance versus cloud cost management

Management is the doing: rightsizing, reservations, cleanup. Governance is the deciding: who owns a number, what threshold triggers what action, which report both sides trust. The FinOps Foundation’s framework holds both, and the split matters for AI spend because the doing (caching, routing, prompt length) sits with product engineering while the deciding sits with whoever owns the P&L.

The FinOps framework phases (inform, optimize, operate) when the workload is inference

The Foundation’s three phases hold. What each one needs changes.

PhaseCloud compute versionInference version
InformSpend by account and tagCost per request, per active user, per subscriber, by feature
OptimizeRightsizing, reservations, commitmentsCache hit rate, model routing, prompt length, pool sizing
OperateMonthly review, budget variancePacing alerts on unit cost, re-measure on every prompt change

The row that trips teams is Inform. A program cannot optimize or operate a number it reports at the wrong grain, and the account grain is wrong for a shared model endpoint serving five features. The sibling piece on what changes when the workload is inference covers why; this one covers the controls.

The twelve controls, by what each one catches

Same profile for each: what it governs, what drives it, how it behaved in production, and the caveat.

Budgets and pacing (burn rate)

A budget is a ceiling; a pacing line is the daily spend that stays under it. Inference paces on requests per active user, which doubled within two quarters of launch on the products I have measured, so a budget set at launch volume is breached mid-year. Pace against the forecast curve, not a flat monthly twelfth.

Caveat: a budget without a pacing line is only discovered when it is gone.

Driver-based forecasting (forecast variance)

A forecast built on the drivers (requests, tokens per request, cache hit rate, price per token) rather than on last month’s total. Driver-based forecasting landed within 5 percent of actual on a nine-figure multi-cloud estate, and it is the only forecast shape that survives a prompt change, because the changed driver is visible in the model.

Caveat: the drivers have to be re-derived when the product changes shape. A forecast that held for three quarters missed by almost thirty when a feature ran ten times the inference per session.

Anomaly detection (cost spike alerts)

Alerts that fire when spend leaves its band. The band has to be per feature and per unit, because total spend rises with adoption and that is not an anomaly. Generic thresholds fire on every legitimate scale-up or stay silent while one feature’s cost per request climbs.

Caveat: anomaly detection bought before allocation is fixed has nothing to define “normal” against.

Unit cost reporting (cost per request, per active user, per subscriber)

The metric layer everything else reads. Cost per request explains a move; cost per subscriber lands next to revenue. Features on the same model differed ten to one in production, which no account-level report can show.

Caveat: the report has to come from the usage export joined to product analytics, not from the invoice divided by requests.

Feature-level allocation (request labeling)

Every model call carries the feature that triggered it, so cost attributes by token count against a feature instead of by which account hosts the endpoint. This is the control that made unit economics exist at all on a shared model layer.

Caveat: labeling is an engineering change, and it lands only when the P&L owner asks for it by name.

Shared pool allocation (provisioned throughput, PTU)

Provisioned capacity is a fixed cost five applications draw on and none of them sized. The allocator splits it by each application’s share of throughput. At low utilization, a request on a provisioned pool ran two to three times the pay-per-token equivalent, and without the allocator no application owner could see that.

Caveat: report the pool as pay-per-call and it looks flat for a month, then steps.

Commitment coverage (committed use, reserved capacity)

The share of steady-state spend covered by a commitment. On compute this is a savings lever. On inference it is a governance lever too: the commitment sets the pool, and the pool sets the allocation. Under-commit and unit cost is high; over-commit and the pool sits idle at two to three times the pay-per-token rate.

Caveat: commit to the forecast’s p50, not its peak.

Showback and chargeback (cost allocation to teams)

Showback presents each team its number; chargeback bills it. A team that has watched its own number for two quarters accepts it. A team billed cold fights the number, and the fight is rarely about the arithmetic.

Caveat: the first disputed number breaks the model unless there is a documented way to re-run it.

Tagging policy (labels, cost categories)

Tags still govern the cloud side: the retrieval index, the vector database, the eval pipeline. What tags cannot do is split a shared endpoint by feature; that is request labeling’s job.

Caveat: a tagging policy that reaches 95 percent coverage on compute can reach 0 percent on inference, because the endpoint is one resource.

Policy as code (spend caps, guardrails)

A spend cap that is an enforced rule instead of a slide. For inference: per-feature rate limits, a maximum output length, a model tier a feature may not exceed without sign-off. Policy in a document holds until someone forgets, and someone forgets.

Caveat: cap the unit, not the total. A total cap throttles the product’s best customers first.

The reporting path (BI layer versus provider console)

Where finance reads the number. The provider console reports what was billed at the API call. A BI layer on the usage export reports tokens, features, and units; that is what FP&A ran unassisted after the rebuild, pacing, budget, forecast, and anomaly views in a tool they already opened.

Caveat: the console is free and the BI layer is a project. The project pays for itself the first time the forecast holds.

Rate levers finance can see (caching, batch, model routing)

The optimize-phase levers as governance inputs. Cache hit rate, batch share, and model mix per feature belong on the finance dashboard, because each moves unit cost by a third or more and none shows on an infrastructure metric.

Caveat: finance should see them, not set them. The lever belongs to the product team; the number belongs to both.

Provider console versus BI layer: when to rebuild

Rebuild when a shared endpoint serves more than one feature, when finance asks a question only the usage export can answer, or when the forecast has missed twice. Before that, the console plus a monthly export is enough.

Showback versus chargeback: when to switch

Switch to chargeback only after two quarters of showback with no disputed numbers, and only when the team being charged holds the lever that changes the number. Charging a product team for a pool an infrastructure team sized produces a meeting where everyone agrees the number is high.

Pay-per-token versus provisioned: which to govern how

Pay-per-token spend is governed by unit cost and pacing. Provisioned spend is governed by utilization and allocation. Mixing the two on one dashboard is how the AI line item stops making sense; the per-token features look expensive on busy days and the provisioned features look free until the resize.

The mistake: governing the account when the spend is the feature

The mistake is running cloud cost governance for AI at the account level because that is where the console reports. The total is right and no decision can be made on it, because the account holds five features and one of them is losing money. The symptom: finance asks why the AI line item did not move when usage doubled, and engineering cannot answer without a query finance cannot run. The first feature-level breakdown on that financial data platform found one of the top three AI products running several times the portfolio average; the account view had hidden it for two quarters.

The rule

Every governance control runs on a unit finance can read without an engineer, or it is not governance yet. A budget on the account, an alert on total spend, a forecast on last month’s invoice: each one is a control that needs an engineer to explain, which means finance is not running it.

This piece describes controls, not accounting treatment. How AI spend is capitalized, expensed, or recognized is a question for your finance and legal team, not this page.

FAQ

What is cloud cost governance?

The policies, ownership, and controls that decide who is accountable for cloud spend and what happens when it moves: budgets, forecasts, allocation, alerts, and the reporting path finance uses. The FinOps Foundation places it inside the operate phase of its framework. For AI spend, the controls have to work at the feature and token level, not the account.

What are the FinOps framework phases?

The FinOps Foundation defines three: inform (visibility and allocation), optimize (rate and usage levers), and operate (continuous governance). They repeat as a loop rather than a sequence. Inference stresses inform hardest, because a shared model endpoint cannot be allocated by tag, and the other two phases inherit whatever grain inform reports at.

What is a cloud cost management framework?

A structured way to allocate, report, optimize, and govern cloud spend, usually built on the FinOps framework’s phases and capabilities. For a team running AI products, the framework needs one addition: unit cost per feature, so cost per request and per subscriber sit alongside spend by account.

What are cloud cost management best practices for AI workloads?

Label every model call with its feature, report cost per request and per subscriber rather than the invoice, forecast on drivers rather than history, allocate provisioned pools by throughput share, and put the whole thing in a tool finance already opens. Cache hit rate and model mix belong on the finance dashboard too.

What to do next

Pick one shared endpoint, label its calls by feature for a month, and put cost per request per feature next to the account total. When the gap between those two numbers needs to become a reporting path finance runs alone, that is the engagement on FinOps consulting for AI products. The rest of the FinOps for AI practice is on the home page.


I put a number on what an AI product costs per request, per active user, and per subscriber, early enough to change the model, the architecture, or the price. If you’re shipping something with a model behind it and nobody can tell you what it costs, email me.

Have a number nobody can explain?

Send a note and I will get back to you.