
Two numbers from the FinOps Foundation’s 2026 State of FinOps survey belong in the same sentence. Ninety-eight percent of FinOps practices now manage AI spend, up from thirty-one percent two years earlier. And the top request those practitioners make of FOCUS, the billing data spec their tools are built on, is coverage for AI workloads. Nearly every team owns the AI bill; the standard underneath their tooling does not yet describe it.
That is the gap FinOps tools leave open. A FinOps tool shows the cloud bill by account, service, and tag, a total with no unit behind it. It cannot show what one request, one active user, or one subscriber costs, because the usage that drives inference spend lives in product analytics and the model gateway, not the billing export.
The direct version: FinOps tools see AI spend as line items (model endpoints, provisioned throughput, GPU hours, vector databases) allocated by tag. They do not see tokens per request, cache hit rate, retries, or feature-level usage, so cost per request and cost per subscriber never appear. On the production AI products I have measured, that missing join was the entire unit-economics layer.
FinOps tools for AI spend: what each capability sees
The FinOps Foundation’s framework lists the capabilities a practice needs; FinOps platforms and the native consoles from AWS, Azure, and Google Cloud implement some version of the twelve below. The table is what each showed on a Fortune-500 financial data platform with eight figures of annual AI and cloud spend, abstracted to shape.
| Capability | Sees on cloud spend | Sees on AI spend | Missing |
|---|---|---|---|
| Cost allocation | Spend by tag, account, project | The model endpoint as one line | Which feature made the call |
| Showback and chargeback | Spend by team | Spend by the endpoint’s owner | Spend by the feature’s owner |
| Anomaly detection | Spikes on daily spend | Spikes on endpoint spend | A feature’s unit cost climbing under a flat total |
| Budgets and pacing | Burn against a monthly number | Same | Burn per active user |
| Forecasting | Trend on the bill | Trend on the bill | The drivers: users, requests per user, tokens |
| Commitment management | RI, savings plan, CUD coverage | Provisioned throughput coverage | Pool utilization per feature |
| Rightsizing | Idle and oversized compute | Idle GPU nodes | Prompt length, model tier, cache hit rate |
| Kubernetes allocation | Pod and namespace cost | GPU pod cost | Which requests ran on the pod |
| Unit economics | Cost per business metric, if fed one | Cost per request, if fed requests | The feed |
| FOCUS ingestion | Normalized multi-cloud billing | Normalized rows for AI services | A token, a request, a cache read |
| GPU and pool allocation | GPU hours by cluster | The provisioned pool as one line | The pool split by application |
| SaaS and vendor spend | Invoices by vendor | The model API invoice as one line | Usage behind the invoice |
Every row has the same shape: the tool sees the resource; the cost is decided at the request.
The twelve capabilities, profiled
Cost allocation (tagging, showback coverage)
Assigns each billing line to an owner by tag, account, or project. Tagging coverage reached 95 percent on compute and zero on the shared model endpoint, because the endpoint is one resource serving five features; the cloud around the model (retrieval index, vector database, eval pipeline) allocates fine. Caveat: resource-level allocation cannot split one endpoint across the features calling it. That takes request labels, a change in the application, not the tool.
Showback and chargeback (cost attribution)
Reports allocated spend to the team that owns the resource. Caveat: on AI products the endpoint belongs to a platform team and the roadmap that moves the cost belongs to a feature team, so chargeback to the endpoint owner sends the bill to the people without the lever. The governance controls that fixed this route it to feature owners instead.
Anomaly detection (cost alerts)
Flags a day or a service that departs from its baseline on total spend; good at catching a runaway training job or an orphaned cluster. Caveat: on inference, total spend rises with adoption, so the alert fires on every legitimate scale-up and stays silent while one feature’s cost per request climbs under a flat total. The first feature-level breakdown on that platform found one of the top three AI products running several times the portfolio average; no anomaly had fired.
Budgets and pacing (burn rate)
Tracks spend against a number set at launch, usually the vendor’s pricing page times a token estimate. Caveat: that budget was wrong by the end of the first quarter on the products I have measured, because requests per active user doubled within two quarters of launch. Pacing against a fixed number reports the miss; pacing against cost per active user times the user forecast explains it.
Forecasting (spend projection)
Projects the bill forward from its own trend. Caveat: a trend forecast held within 2 percent for three quarters on that platform and missed by almost thirty in the fourth, because one feature ran ten times the inference per session and the trend had no driver for it. Driver-based forecasting (users, requests per user, cost per request, by feature) landed within 5 percent on a nine-figure multi-cloud estate; the tool carries no drivers unless someone feeds them.
Commitment management (reserved instances, savings plans, committed use)
Tracks coverage and utilization of commitments and recommends more, on compute and on the provisioned throughput pool. Caveat: at low utilization a request on a provisioned pool ran two to three times the pay-per-token equivalent. Coverage percent reports the pool as covered and cannot show that it was half empty.
Rightsizing (workload optimization, waste reduction)
Lists oversized or idle compute and the savings if resized. Right for training clusters, GPU nodes, the retrieval stack. Caveat: the levers that move inference cost (prompt length, model tier, cache hit rate, retry rate) are not resources, so the tool has no recommendation for them, and its savings backlog grows while the largest lever on the AI product sits outside it. The survey has mature practitioners reporting diminishing returns from traditional optimization.
Kubernetes cost allocation (container cost)
Splits shared cluster spend across pods, namespaces, and labels. Right for self-hosted model serving on GPU nodes. Caveat: it allocates the pod, not the request. On the same deployed model with the same instance count, cost per request moved by more than half within a quarter after a prompt template changed; nothing infrastructural moved, so nothing in the allocation did.
Unit economics (cost per unit)
Divides allocated spend by a business metric; the survey lists defining unit economics as a rising priority. Caveat: the tool divides; it does not count. The numerator is the allocation problem above, and the denominator (requests, active users, subscribers, per feature) lives in product analytics. On that platform the join was built outside the tool, and features on the same model differed ten to one in cost per request once it existed.
FOCUS ingestion (multi-cloud billing normalization)
Reads billing exports from several providers into the FOCUS specification, which the FinOps Foundation maintains. Caveat: FOCUS normalizes billing rows, and a billing row for a model endpoint is spend by hour, not by request. Practitioners’ top expansion request for the spec is AI workloads: the same gap in one line.
GPU and provisioned-throughput allocation (PTU pools)
Reports GPU hours by cluster and provisioned throughput as a reserved line, billed flat whether used or not. Caveat: the pool serves several applications and the tool shows one line. Allocating it needs each application’s share of throughput, which comes from the gateway, not the bill. A planner that split a shared pool by throughput closed a total-cost-of-ownership gap the AI product reporting had carried for two quarters.
SaaS and vendor spend (AI vendor invoices)
Ingests model vendor and SaaS invoices alongside cloud billing; 90 percent of practices now manage SaaS spend or plan to, per the survey. Caveat: an invoice is a total. The usage export behind it (input tokens, output tokens, cached reads, by key) turns the invoice into cost per request, and the tool ingests the invoice.
The thirteenth capability: token-level metering
The capability none of the twelve carry is the one the AI product runs on. Tokens in, tokens out, cached reads, retries, tool calls, and the feature label on each request live in the model gateway or an LLM observability layer; that is where cost per request is born. A tool that ingests it can compute unit economics; one that does not shows the endpoint as a line. FinOps for AI covers why the cloud assumptions under the other twelve fail on inference.
Buy, extend, or build
The provider’s console (AWS Cost Explorer, Azure Cost Management, Google Cloud Billing) and a third-party FinOps platform see the same billing rows; the platform adds multi-cloud normalization, allocation rules, and a place to put a unit metric. Neither sees a token. So: buy a platform for the cloud estate around the model. Extend it with gateway usage if it can take a custom feed with a feature label. Build the join yourself when it cannot, or when the unit metrics belong in the tool finance already uses. On that financial data platform the provider console came off finance’s reporting path, replaced by billing data joined to product analytics in the tool FP&A already opened. They run it unassisted.
The mistake: buying the tool before the feed exists
Selecting a FinOps platform for its AI cost dashboard and finding, a quarter in, that it shows the model endpoint as one line because nothing in the application labels a request. The symptom: a savings backlog that grows every quarter while cost per request per feature is not on the dashboard at all. The fix is upstream: a feature label on every request, a usage export from the gateway, and one definition of an active user that finance and product share.
The rule
A FinOps tool sees AI spend only as far down as the data it is fed, and the bill stops at the endpoint. Feed it request-level usage with a feature label and unit economics works: cost per request from token metering, cost per subscriber from product analytics, the bill’s movement from drivers instead of trend. Feed it the bill and it reports a total, in a nicer chart than the console.
FAQ
What are FinOps tools?
Software that ingests cloud and vendor billing data and implements the FinOps Framework’s capabilities: allocation, showback and chargeback, anomaly detection, budgets, forecasting, commitment management, rightsizing, and unit economics. Native cloud consoles and third-party platforms both qualify. For AI spend, what matters is whether the tool can ingest request-level usage, not which capabilities it lists.
Do FinOps tools support AI and LLM cost?
At the resource level, yes: model endpoints, provisioned throughput, GPU nodes, and vendor invoices appear as billing lines and can be allocated by tag. Cost per request, cost per active user, and cache hit rate need token-level usage with a feature label from the model gateway. Whether a tool can take that feed is the question to ask before buying.
What is the best FinOps tool for AI spend?
The one that can ingest request-level usage from your model gateway with a feature label, divide it into cost per request and cost per subscriber, and put the result in the reporting path finance already runs. On the products I have measured that join was built outside every tool evaluated; the feed mattered more than the tool.
What does FOCUS change for AI spend?
FOCUS, the FinOps Foundation’s open billing specification, normalizes cost and usage rows across providers so tools stop parsing three formats. A FOCUS row for a model endpoint is still spend by hour, not by request, which is why practitioners’ top expansion request for the spec in the 2026 survey is AI workloads.
What to do next
Before evaluating a FinOps tool for the AI product, check whether every request carries a feature label and whether the gateway exports usage by it; without both, the tool shows the endpoint as a line. The LLM cost calculator gives the per-request number the tool should reproduce. When the join between billing and product analytics needs to become a reporting path finance runs alone, that is FinOps consulting for AI products. The rest of the FinOps for AI practice is on the home page.
I put a number on what an AI product costs per request, per active user, and per subscriber, early enough to change the model, the architecture, or the price. If you’re shipping something with a model behind it and nobody can tell you what it costs, email me.
Have a number nobody can explain?
Send a note and I will get back to you.