
The AI business case went to the CFO with a market-size slide, a productivity-uplift slide, and a cost line that was the vendor’s pricing page times an estimated token count. It came back with one question: what does this cost per subscriber when usage is ten times what the pilot showed. Nobody had the number, because the pilot was the only usage anyone had measured.
That is the shape of most AI business cases that stall. They argue value in units finance cannot check and state cost in a unit that stops being true the first time usage shifts. The cases that get signed are shorter. They carry five numbers finance can recompute, five assumptions written down so they can be wrong in public, and none of the two slides a CFO stops reading at.
The direct version: an AI business case a CFO will sign states cost per request at launch and at ten times usage, break-even usage, cache-adjusted gross margin per subscriber, and the forecast variance the case tolerates, with the assumptions behind each one named. I built the decision tooling behind that shape at the direct request of a CFO, CTO, and CPO at a Fortune-500 financial data company, so product teams could run it before a line of code was written.
The AI business case, by number
| Number | Where it comes from | Why it is in the case |
|---|---|---|
| Cost per request at launch | Token split, cache hit rate, retries, in the calculator | The unit everything else is built from |
| Cost per request at 10x usage | Same inputs with the cache and retry rates usage changes | Shows whether the unit holds as the product scales |
| Break-even usage | Cost per subscriber against price, by usage tier | The line the product has to stay under |
| Cache-adjusted margin per subscriber | Cost per active user times activation, at measured hit rate | The margin in the same unit as revenue |
| Forecast variance tolerance | How far actual can miss before the case fails | The honest part; what the CFO checks first |
The five numbers
Same profile for each: what it is, what moves it, how it behaved in production, and the caveat.
Cost per request at launch (unit cost at pilot usage)
Model, retrieval, and eval spend per call, built from the drivers rather than the invoice. On production products, the ride-alongs beyond tokens (retries, tool calls, retrieval, evals) added 30 to 80 percent on top of the tokens-times-price estimate. A case that carries only the token estimate is understated before it is presented.
Caveat: the prompt you tested with is shorter than the one you ship.
Cost per request at ten times usage (unit cost at scale)
The same number recomputed where usage changes the inputs: cache hit rate rises with traffic, retries rise if a cheaper model is swapped in under load, and provisioned capacity changes the marginal cost entirely. At low utilization a request on a provisioned pool ran two to three times the pay-per-token equivalent; at high utilization, below it.
Caveat: ten times usage is not ten times users. Requests per active user doubled within two quarters on its own.
Break-even usage (break-even point)
The usage per subscriber at which cost per subscriber equals the price. Under seat pricing the case has a crossover by construction; the number states where it is. If break-even sits below the heavy-user cohort’s usage, the case has a negative-margin cohort and needs a pricing change, not a caching plan.
Caveat: compute it at three usage tiers, not one. The median tier passes and the top tier is where the decision lives.
Cache-adjusted gross margin per subscriber
Price minus cost per subscriber, where cost per subscriber is cost per active user times activation, at the cache hit rate the product will see. Under 30 percent hit rate the cached price is a rounding error; above 60 percent, pricing on uncached rates overstates cost per request by a third or more. The case states which it assumed.
Caveat: plan on the uncached price until a month of production data says otherwise.
Forecast variance tolerance
How far actual cost can miss the forecast before the case no longer holds. Driver-based forecasting landed within 5 percent of actual on a nine-figure estate; a forecast that ignored a change in the unit of work missed by almost thirty. A case that tolerates 10 percent variance and is built on a method that misses by thirty is a case with no tolerance at all.
Caveat: the CFO reads this number first, because it is the one that says whether the author expects to be checked.
The five assumptions that have to be written down
Requests per active user
The input launch estimates skip and the one that moved most. State the pilot number, state what the case assumes at month twelve, and make the second at least double the first.
Activation rate (share of subscribers who are active)
Cost per subscriber depends on it, and it rises after launch. A costly feature at 20 percent activation can be cheaper per seat than a modest feature everyone uses; the case has to say which one this is.
Cache hit rate
The assumption most likely to be wrong pre-build. The hit rate in a test harness with one user is not the hit rate with a thousand.
Retry and re-prompt rate
Never in the pre-build estimate, so it arrives as unexplained variance at the first invoice. On agents and structured-output features, 10 to 20 percent of cost per request; on plain chat, 5 to 10.
Capacity model (pay-per-token or provisioned)
Whether the product runs on a metered rate or a share of a committed pool. The two produce different cost curves as usage grows, and a case that does not say which one it assumed has two answers to every question.
The two numbers a CFO stops reading at
Total addressable market (TAM)
A number the CFO cannot check and the product cannot invoice. It belongs in a fundraising deck. In a business case it signals that the author does not have the unit economics, because a team that had cost per subscriber would have led with it.
Productivity uplift (hours saved)
Hours saved times a loaded hourly rate is a revenue-like number with no revenue in it. Finance has seen it justify every software purchase for twenty years. If the product is sold, the case states margin per subscriber; if it is internal, it states cost per resolved task against the cost of the task it replaces.
Pre-build versus post-launch: two different cases
Before the build, the case runs on the LLM cost calculator with assumed inputs and is judged on whether it survives requests per active user doubling. After launch, the same five numbers come from the usage export joined to product analytics, and the case is re-signed against measured values. The second version is the AI ROI measurement finance asked for; the first is the permission to build.
When to kill the case before the build
Kill it when cost per subscriber at ten times usage exceeds the price under every pricing model on the table, or when break-even sits below the heavy-user cohort and the pricing model cannot change. Either one means the architecture or the price has to move first, and the comparison of AI pricing models is where that conversation goes next. A case that fails this test and gets built anyway becomes the feature that runs several times the portfolio average and hides in the blended number for two quarters.
The mistake: building the case on the median user
The mistake is running every number at the median user’s usage because the median is what the pilot produced. Margin looks healthy on average and negative on the cohort that uses the product most. The symptom: a case that passes at launch and a renewal conversation a year later where the biggest customer is the one the product loses money on.
The rule
A business case is signed on the five numbers it can be wrong about, not the two it cannot. Cost per request at launch and at scale, break-even, cache-adjusted margin, variance tolerance, each with its assumption beside it. If any of the five is missing, the CFO is being asked to sign a slide.
This piece covers the cost side of the case. How AI spend is capitalized or expensed, and how usage revenue is recognized, are questions for your finance and legal team.
FAQ
How do you build a business case for an AI product?
Start from cost per request built from the drivers (tokens, cache, retries, tool calls), scale it to ten times pilot usage, compute cost per subscriber against the price at three usage tiers, and state the variance the case tolerates. Write the five assumptions down. Value claims that finance cannot recompute, such as market size or hours saved, go last or not at all.
How do you measure AI ROI after launch?
Re-run the same five numbers on measured data: cost per request per feature from the usage export, cost per subscriber against price, break-even against the heavy-user cohort, and forecast variance against actual. AI ROI measurement is the pre-build case re-signed on production values, not a new set of metrics invented after the fact.
What is the ROI of generative AI for a product team?
For a product that is sold, gross margin per subscriber after inference cost, tracked as usage grows; for an internal tool, cost per resolved task against the task it replaced, retries included. Blended ROI across a portfolio hides the feature losing money; on production products, features on the same model differed ten to one in cost per request.
What should an AI ROI analysis include?
Unit cost at launch and at scale, break-even usage, margin per subscriber at the measured cache hit rate, the variance tolerance, and the five assumptions (requests per user, activation, cache hit rate, retry rate, capacity model). Plus a line for the infrastructure outside the model bill: retrieval, evals, and any provisioned pool share.
What to do next
Take the feature spec, put its inputs into the LLM cost calculator at pilot usage and at ten times, and see whether cost per subscriber clears the price at the top tier. When the case needs a product leader who has built it before to carry it into the finance review, that is fractional product leadership for AI economics. The rest of the FinOps for AI practice is on the home page.
I put a number on what an AI product costs per request, per active user, and per subscriber, early enough to change the model, the architecture, or the price. If you’re shipping something with a model behind it and nobody can tell you what it costs, email me.
Have a number nobody can explain?
Send a note and I will get back to you.