AI Pricing Models Compared: Seat, Usage, Credits, Hybrid, Outcome

The AI feature shipped inside the existing seat price because that was the fastest decision. Six months later finance asks what it costs per subscriber and the honest answer is a range, because the pricing call was made on the vendor’s pricing page times a token estimate, and usage has shifted twice since.

Every AI pricing model is a bet on how usage per seat grows. Per-seat pricing bets it stays flat. Usage pricing bets the customer will tolerate a variable bill. Credits bet the customer will not do the token math. The bet pays only if the team knows the margin curve under each model before the pricing page exists; most learn it from the first quarter’s invoice.

The direct version: AI pricing models fall into three families, seat, usage, and outcome, plus the hybrids between them. Under seat pricing, margin per subscriber falls as usage grows; under usage pricing, it holds; under outcome pricing, it depends on retries. On the production AI products I have measured, requests per active user doubled within two quarters of launch, which moved the seat-pricing margin curve more than any model price change did.

AI pricing models compared: the margin curve

What happens to gross margin per customer as usage per seat grows. Twelve models, four columns.

ModelMargin as usage per seat growsWho carries usage riskFits when
Per seat (flat)Falls, goes negative at the crossoverVendorUsage per seat is bounded by the workflow
Per active userFalls slower; inactive seats are freeVendorActivation is low or uneven
Usage (metered)HoldsCustomerUsage varies ten to one across customers
CreditsHolds if credits track cost; drifts if notSharedMultiple features on different models
Hybrid (seat + allowance + overage)Falls to the allowance, then holdsShared above the allowanceMost B2B products with a heavy-user tail
Tiered (good, better, best)Falls within a tier, resets at the nextVendor within a tierUsage clusters into recognizable cohorts
Outcome (per resolved task)Depends on retries and resolution rateVendorResolution is measurable and the eval exists
AI add-on (feature gated)Falls, but only for buyers who opted inVendorFeature is optional for most of the base
Free tier (freemium)Negative by designVendorCost per free user is a rounding error
Enterprise commit (platform fee)Holds if the commit is sized on the forecastSharedBuyer wants a predictable annual number
Fair-use capFalls to the cap, then stopsCustomer above the capSeat pricing with a runaway-user problem
Per agent or per workflowDepends on calls per runVendorAgents with a measured p90 run cost

Read the second column first. Wherever margin “falls” there is a crossover where a customer costs more than they pay, and the job is knowing which customers are past it.

The models, by family

Per seat (flat subscription, per user per month)

A fixed price per seat regardless of use. Margin per seat is price minus cost per subscriber, which is cost per active user times activation. Both rise after launch: activation climbs as the feature gets found, and requests per active user doubled within two quarters.

Caveat: the median user is fine. The heaviest users cost several times the median, and they are the ones who renew.

Per active user (cost per MAU pricing)

A seat price charged only for users active that month. Cuts the inactive-seat subsidy and prices closer to cost, at the price of a bill the buyer cannot predict.

Caveat: “active” needs one definition finance, product, and the contract all share.

Usage-based (metered, per token, per request)

Charged on what was consumed. Margin per unit is fixed, so revenue grows with cost. Where the cost spread between customers is widest (ten to one across features on the same model in production), seat pricing hides the most and metering hides nothing.

Caveat: procurement wants a number for the year. Usage pricing without a cap or a commit loses deals to a worse product with a flat price.

Credits (token abstraction)

Usage priced in a unit the vendor defines, so the customer sees credits instead of tokens. Works while one credit tracks one unit of cost. Drifts when different models sit behind the same credit, so a credit spent on a small-model feature and one spent on a large-model feature cost the vendor very different amounts.

Caveat: re-derive the credit’s cost every time model routing changes, or margin moves without a pricing decision.

Hybrid (seat plus allowance plus overage)

A seat price with an allowance and metered overage above it. Fits most B2B AI products because it prices the median user flat and the heavy user by use. The decision is where the allowance sits.

Caveat: set the allowance at the p80 user and a fifth of the base pays overage, which is the share that generates the support tickets.

Tiered (good, better, best)

Seat tiers with different caps, models, or features. Margin resets at each boundary. Works when usage clusters into cohorts a tier can name; fails when the heavy users sit in the cheapest tier because nothing stops them.

Caveat: a tier defined by model quality rather than usage puts the most expensive inference in the tier with the most seats.

Outcome-based (per resolved task, per completed workflow)

Charged when the product does the job: a ticket resolved, a document produced. The cost side is cost per resolved task, retries included. On agent products, retries ran 10 to 20 percent of cost per request, and a failed run costs the same as a successful one.

Caveat: outcome pricing needs the eval that defines “resolved,” and the eval has its own inference cost.

AI add-on (feature-gated pricing)

The AI feature sold as a separate line on the existing seat. Isolates the margin question to the buyers who opted in and gives the P&L owner a clean cost per subscriber for the feature alone.

Caveat: an add-on nobody buys tells you the price was wrong, not the feature.

Free tier (freemium)

Zero price, inference cost carried by the vendor. Sustainable only when cost per free user is a rounding error next to conversion. The number to hold: cost per free active user times the share who never convert.

Caveat: a free tier on an agent product is a free tier on the most expensive request shape you ship.

Enterprise commit (platform fee plus committed usage)

An annual commitment sized on forecast usage, with true-up or overage. Margin holds when the commit is sized on a driver-based forecast, which landed within 5 percent of actual on a nine-figure estate.

Caveat: the buyer’s forecast is not your forecast. Size on your usage data for that cohort.

Fair-use cap

Seat pricing with a ceiling on monthly use. Margin falls to the cap and stops. The cheapest fix for a runaway-user problem, and the one buyers resent least when the cap is above the p95 user.

Caveat: a cap below p95 turns the best customers into the churn cohort.

Per agent or per workflow run

A price per autonomous run. Cost per run is a distribution, because tool calls plus retrieval ran 15 to 30 percent of cost per request with a long tail. Price off the p90 run, not the mean.

Caveat: a run that fans out into more calls next quarter, because the workflow got more capable, moves cost per run without moving the price.

Seat versus usage: when usage pricing is required, not preferred

Usage pricing is required when the heaviest cohort’s monthly inference cost exceeds the seat price. At that point seat pricing has a negative-margin cohort by construction, and no amount of caching fixes the arithmetic. Below that line, a hybrid with an allowance does the same job with a predictable bill.

Pre-build versus in production: when to set the price

Before launch, the number comes from the LLM cost calculator: token split, requests per user, cache hit rate, at three usage tiers. Good for choosing between seat and hybrid. After launch, cost per subscriber comes from the usage export joined to product analytics, and that is the only number to re-price on. Set the price on the pre-build estimate and never revisit it, and the launch estimate is half the real number by the time the pricing page is public.

When to switch from seat to hybrid

Switch when three things are true at once: cost per subscriber is above a third of the seat price, the top decile of users drives more than half of inference cost, and requests per active user is still climbing. Any one alone is a caching or context-trimming problem. All three together is a pricing problem.

The mistake: pricing the median user

The mistake is setting the AI price on the median user’s cost, because the median is what the estimate produced. Margin looks healthy on average and negative on the cohort that uses the product most. The symptom: a stable blended margin and a renewal where the biggest customer is the one the product loses money on. The first feature-level breakdown at a Fortune-500 financial data company showed one of its top three AI products running several times the portfolio average; the blended number had hidden it for two quarters.

The rule

Price off cost per subscriber for the heavy-user cohort, and re-price whenever requests per active user move by a quarter. The cohort clause is what keeps the best customers profitable. The re-price clause is there because the usage curve moves faster than the pricing page does.

Revenue recognition under usage and credit pricing is an accounting question for your finance team; this piece covers the margin math, not the treatment.

FAQ

What are the main AI pricing models?

Three families: seat (a flat price per user), usage (a metered price per token, request, or credit), and outcome (a price per resolved task), plus hybrids that combine a seat price with a usage allowance. Most B2B AI products land on a hybrid because it prices the median user flat and the heavy user by consumption.

What is the best pricing model for a generative AI product?

The one whose margin holds at the usage level your heaviest cohort reaches, not the one that is easiest to sell. Seat pricing works when usage per seat is bounded by the workflow. Hybrid works when there is a heavy-user tail. Pure usage is required when the heaviest cohort’s inference cost exceeds any flat price you could charge.

How do enterprise generative AI pricing models differ from self-serve?

Enterprise deals add a committed annual number, usually a platform fee plus a usage commitment with true-up, because procurement needs a predictable line item. The margin question is the same; the commit has to be sized on your usage forecast for that cohort, not the buyer’s estimate.

What to do next

Pull one month of usage by user, sort by cost, and put the top decile’s cost per subscriber next to the seat price. If it is above the line, the pricing decision is already made and nobody has written it down. When that call needs a product leader who has built the cost model, that is fractional product leadership for AI economics. The per-request math is in what one AI request actually costs; the rest of the FinOps for AI practice is on the home page.


I put a number on what an AI product costs per request, per active user, and per subscriber, early enough to change the model, the architecture, or the price. If you’re shipping something with a model behind it and nobody can tell you what it costs, email me.

Have a number nobody can explain?

Send a note and I will get back to you.