
RAND Corporation interviewed fifty experienced AI engineers and researchers in 2024 and found that leadership-driven failure, not a technical one, was the reason eighty-four percent of them named for a project falling apart. More than eighty percent of AI projects fail overall, twice the rate for non-AI IT projects, and one interviewee described the survivors as routinely shipping at “roughly half of what they could have been.” That number describes a project that never reaches users.
The harder failure to see is the one that ships, works, and keeps running anyway. Finance’s spreadsheet shows the total AI bill for the quarter, not which feature drives it. Engineering’s dashboard shows requests per day and latency, not what each request earns against what it costs. A feature can sit between those two views, technically successful and financially bleeding, for two or three quarters before either one flags it.
The direct version: an AI feature that shipped and works can still be losing money, and twelve leading indicators show up in usage and cost data well before a quarterly P&L review would catch them, from requests-per-user drift to a stale cost-per-resolved-task check-in. They are grouped below by where each one shows up first.
The twelve margin signals
Each signal below is named by what it looks like in the data, not by a stage or a sprint, because none of these show up on a calendar. They show up in one of three places: the usage data product teams already watch, the cost data finance already watches, or the process around the feature that nobody watches at all.
Usage signals (visible in product analytics before Finance ever sees a number)
- Requests-per-user drift. Usage per active user creeps past what the launch business case assumed, and nobody re-priced the feature against the new number.
- Activation without cost-per-user tracking. The feature has an activation metric on a dashboard somewhere and no corresponding cost-per-user metric next to it.
- A degraded cache hit rate. A prompt structure or product edit moved the cache breakpoint, the hit rate fell, and cost per request rose with it, invisible unless someone is watching that specific number.
- A climbing retry rate. Retries and re-prompts on quality grounds are a silent cost multiplier that rarely shows up as its own line anywhere.
Cost signals (visible in billing or cost data, before anyone connects it to a specific feature)
- Model-tier drift. Routing has quietly shifted a share of traffic to a costlier fallback tier, usually after a reliability incident, and never shifted back.
- Context bloat. Tokens per request are growing without a corresponding quality gain, a common byproduct of stacking more instructions and examples into a prompt over time.
- Cost per request rising while cost per active user flatlines. This combination usually means usage is concentrating in a smaller, heavier-using segment that the blended average is hiding.
- A shared-pool allocation that never gets rebalanced. A feature’s share of a shared inference pool or provisioned-throughput commitment changes as its traffic does, and the allocation formula rarely gets revisited to match.
Process signals (visible only if someone is looking for them, since no dashboard tracks them by default)
- No re-run since launch. Nobody has re-calculated the feature’s unit economics since the original business case, regardless of how much the underlying model, usage, or pricing has moved.
- No named owner for the margin number. Ownership got assigned for uptime and for the roadmap, not for cost per resolved task, so no one is accountable when it drifts.
- A roadmap review that cites growth with no margin caveat. Usage growth gets presented as the win with no adjoining cost-per-user trend, which is how a feature can be called a success and lose money in the same meeting.
- A stale cost-per-resolved-task check-in. More than a quarter has passed since anyone looked at the feature’s cost per resolved task specifically, as opposed to its total spend.
Why the aggregate view misses all twelve
| What Finance sees | What Engineering sees | What neither sees |
|---|---|---|
| Total AI and cloud spend by month | Requests per day, latency, error rate | Cost per resolved task, by feature |
| Spend by service or model endpoint | Active users and retention | Whether a specific feature’s margin moved since launch |
| A quarterly variance versus budget | Uptime and incident count | Which of the twelve signals above just tripped |
Finance’s spreadsheet and Engineering’s dashboard are both correct for what they measure. Neither one was built to hold a feature’s unit economics next to its usage trend, which is the join that reveals a margin problem before it becomes a budget conversation.
The mistake most teams make
The most common mistake is not missing one signal. It is treating usage growth and cost growth as two separate conversations, reviewed on two separate cadences, by two different people who rarely compare notes on the same feature in the same meeting. A product review covers adoption. A finance review covers the bill. Neither one is set up to say “this specific feature’s cost per active user rose forty percent while adoption rose twenty,” because that sentence needs both datasets in the same room.
The observable symptom is a feature that keeps appearing on the roadmap review as a success story, quarter after quarter, with no cost-per-user line anywhere on the slide. By the time it shows up in a budget review as a problem, the gap has usually been compounding for two or three cycles.
The rule
Re-check a feature’s unit economics the moment any one of the twelve signals above trips, not on the next scheduled review. A calendar-based cadence, quarterly or annual, means a signal that appears in month one goes unaddressed until month three at the earliest. A threshold-based rule catches it the month it happens, which is the difference between a five percent margin miss and a fifty percent one.
Kill, fix, or re-price
Not every tripped signal means the feature should die. Three responses fit three different situations:
- Fix it when the signal traces to something correctable in isolation: a cache regression from a recent prompt change, a retry loop from a model-quality issue, or a routing bug sending traffic to the wrong tier. These are engineering fixes, not strategy decisions, and usually resolve the margin problem without touching the feature’s design.
- Re-price it when usage patterns have genuinely shifted (a feature is being used more heavily, by more engaged users, than the original pricing model assumed) and the value delivered still justifies the cost, just not at the current price point.
- Kill it when neither a fix nor a re-price closes the gap: the unit economics were wrong at launch, the model class required to deliver the feature well has no affordable path, or the segment using it heavily is not one the business wants to subsidize. This is the hardest of the three to say out loud in a roadmap review, which is exactly why it is the one most often avoided past the point it should have been made.
The calculator turns any of the twelve signals into a number: model tier, token split, and cache hit rate in, cost per request and cost per active user out, so the fix, re-price, or kill decision has a figure behind it instead of a hunch.
FAQs
How do you know an AI feature is losing money if it looks successful? Usage growth and cost growth get reviewed separately by design, on different cadences, by different teams. A feature can show rising adoption on a product dashboard and rising cost per user on a billing report at the same time, and neither review is set up to put both numbers on the same slide.
What’s a reasonable cost per active user for an AI feature? There is no universal number. What matters is the trend against the feature’s own launch estimate and its pricing model, not a benchmark against other products. A cost per active user that has grown faster than the price or plan tier it sits under is the signal that matters, regardless of the absolute figure.
Should an underperforming AI feature be killed or re-priced? Re-price it if usage has genuinely grown and the value still holds at a higher price point. Kill it if the unit economics were wrong at launch or the model class required has no affordable path, since a fix or re-price only helps when the underlying gap is closeable.
How often should a team re-check AI feature margins? On a threshold, not a calendar. Re-checking every time usage per user, cache hit rate, or model tier moves meaningfully from the launch estimate catches a margin miss the month it happens, instead of waiting for the next scheduled review to notice a problem that started two cycles earlier.
What to do next
Run the feature’s current numbers through the LLM cost calculator to see where it sits against its launch estimate, then check that number against the FinOps KPIs that actually caught a margin problem in production versus the ones that didn’t. If the fix, re-price, or kill call needs a second set of eyes, that is the conversation a fractional product leader for AI economics exists to have, backed by the same FinOps for AI discipline behind every number on this site, and detailed further on the FinOps consulting page. It pairs directly with the business case and pricing models work that got the feature built in the first place.
I put a number on what an AI product costs per request, per active user, and per subscriber, early enough to change the model, the architecture, or the price. If you’re shipping something with a model behind it and nobody can tell you what it costs, email me.
Have a number nobody can explain?
Send a note and I will get back to you.