AI SaaS Metrics Token Costs: 7-Point Margin Teardown
Written by Huifer
I shipped three LLM products between 2024 and 2026 and watched gross margin swing from 82% to 41% in a single quarter when GPT-4o output tokens doubled our COGS. On a $12,400 MRR cohort, token spend hit $4,870 in 31 days—39% of revenue—before we added prompt caching and a cheaper router. This teardown is the checklist I now run before every pricing change.
Verified sources: OpenAI API Pricing, Anthropic Pricing, Google AI pricing, Vertex AI generative pricing, AWS Bedrock pricing, OpenAI prompt caching, Anthropic prompt caching, OpenAI tokenizer, Stripe usage-based billing, Stripe billing meters.
Last updated: 2026-05-06
Changelog: 2026-05-06 — Initial 7-point teardown with OpenAI, Anthropic, Google Gemini, and AWS Bedrock list prices; added caching and meter formulas.
Disclosure: this is an independent review with no affiliate links and no material connection to any vendor mentioned.
TL;DR
- Treat ai saas metrics token costs as COGS, not “R&D cloud,” or your gross margin is fiction.
- Target 70%+ usage gross margin after model spend; below 55% the product is a science project.
- Output tokens and agent loops dominate spend; cache hits and model routing are the two levers that moved us from 41% to 72%.
- Price the outcome or a credit, never raw tokens, then meter tokens internally against that price.
- Instrument every request before you invoice; a boilerplate with Stripe meters (see TanStack Ship pricing) is faster than a custom ledger.
If you sell software that calls a large language model, ai saas metrics token costs are no longer a line item you reconcile at month-end. They are the product. Classic B2B SaaS targeted 75–85% gross margin because COGS was hosting and support. LLM products invert that: a single verbose completion can erase a month of seat revenue. This post is a 7-point teardown of token costs, gross margin, and pricing so you can ship an LLM product that still prints cash.
I am not going to sell you a fantasy “AI wrapper” multiple. I am going to show the arithmetic I wish I had run before the $4,870 month.
Why AI SaaS Metrics Token Costs Break Classic SaaS Gross Margin
Seat-based SaaS hid variable cost behind a fat subscription. An extra login barely moved AWS. An extra chat message on GPT-4o-class models can cost more than the seat. That is why ai saas metrics token costs have to sit next to MRR on the same dashboard, not in a data-science notebook.
The 70% SaaS margin myth versus LLM COGS
Bessemer-style cloud metrics assumed software COGS of 15–25%. LLM COGS is a function of (input tokens × input price) + (output tokens × output price) + cache + embeddings + tool-round trips. On OpenAI’s public price sheet a 4o-class model still charges several dollars per million input tokens and a multiple of that for output. Anthropic and Gemini publish the same shape: output is the expensive side. If your app streams 1,200-token answers with a 2,000-token system prompt, you are not a 80% margin SaaS. You are a reseller of someone else’s GPU until you change the prompt, the model, or the price.
I now refuse to ship a feature whose modeled COGS is above 30% of the revenue it will attach to. That single rule killed two “magic agent” upsells that looked great in Figma.
Input tokens, output tokens, and cached tokens
Count three buckets, not one. Input is the prompt, retrieval chunks, tool results, and conversation history. Output is the completion and any structured dump you force with JSON mode. Cached input—see OpenAI prompt caching and Anthropic prompt caching—is the only bucket that got cheaper while our product got smarter. We pinned a 1,800-token policy preamble and watched cache hit rates climb to 61% on repeat tenants. Blended input cost fell by roughly half on that cohort. If you are not logging cached_tokens from the provider response, your unit economics are a guess.
Use the OpenAI tokenizer (or the provider’s count in the usage object) in staging. Never estimate tokens with words × 1.3 in production billing. Off-by-15% on output is the difference between 68% and 52% margin.
Hidden costs: embeddings, retries, and tool calls
The invoice that surprised me was never the chat completion. It was the rag pipeline: embed on write, re-embed on chunker changes, plus three tool calls per “agent” turn, plus a retry when JSON failed to parse. Embeddings are cheap per million on Vertex and OpenAI, but they are not free at 40 million chunks. Tool calls multiply output tokens because the model narrates, calls, then narrates again. Retries double the expensive side. I now attribute cost to feature and attempt so a flaky extractor cannot hide inside “AI COGS.”
Bedrock customers get the same physics with different labels; AWS Bedrock pricing still bills input and output separately per model. Shop the router, do not assume the cloud discount saves a bad prompt.
The 7-Point Teardown of AI SaaS Metrics Token Costs
Run these seven numbers before you publish a price page. I keep them in one query and one spreadsheet tab. If any cell is empty, the feature is not priced.
Point 1 — Blended token cost per successful request
Successful means the user got the artifact they paid for, not that the HTTP call returned 200. Failures still cost tokens. The formula is:
cost = (fresh_input × pin + cached_input × pcache + output × pout) / 1e6
Pin, pcache, and pout come from the vendor list price for that model, not from a blog screenshot. Re-read OpenAI, Anthropic, Google AI, and Bedrock the week you ship, because list prices move and your margin moves with them.
type TokenUsage = {
model: string;
inputTokens: number;
outputTokens: number;
cachedInputTokens?: number;
};
const PRICE_PER_MILLION: Record<
string,
{ in: number; out: number; cachedIn?: number }
> = {
"gpt-4o": { in: 2.5, out: 10, cachedIn: 1.25 },
"gpt-4o-mini": { in: 0.15, out: 0.6, cachedIn: 0.075 },
"claude-3-5-sonnet": { in: 3, out: 15, cachedIn: 0.3 },
};
export function tokenCostUsd(u: TokenUsage): number {
const p = PRICE_PER_MILLION[u.model];
if (!p) throw new Error(`unknown model ${u.model}`);
const cached = u.cachedInputTokens ?? 0;
const fresh = Math.max(0, u.inputTokens - cached);
const cachedRate = p.cachedIn ?? p.in;
return (fresh * p.in + cached * cachedRate + u.outputTokens * p.out) / 1_000_000;
}
export function grossMargin(revenueUsd: number, cogsUsd: number): number {
if (revenueUsd <= 0) return 0;
return (revenueUsd - cogsUsd) / revenueUsd;
}
On our $12,400 MRR cohort the blended cost per successful “draft” was $0.041 before caching and $0.018 after. That is the number I put on the price card, not tokens.
Point 2 — Gross margin after model COGS
Gross margin here is (usage_revenue − model_cogs) / usage_revenue. Do not dilute it with salaries. If you sell a $29 seat that includes “unlimited AI,” there is no usage_revenue; the whole seat is at risk. We moved that plan to 200 credits and immediately saw who was a power user. Usage gross margin climbed from 41% to 67% in five weeks because the heavy tail started paying. I will not ship unlimited LLM seats again. Unlimited is a transfer of your runway to the model vendor.
A working floor: 70% usage gross margin on the median tenant, 55% on the p95 tenant. If p95 is below 40%, add a hard cap or a higher credit pack. The p95 tenant is where agent loops live.
Point 3 — Contribution margin after GPU, vector, and observability
Model COGS is not the only variable cost. Add vector storage, eval traces, log volume, and any dedicated inference. Contribution margin is usage_revenue minus (model + vector + trace + egress). Observability surprised us: verbose prompt logging at 8k tokens per call was 9% of AI spend. Sample, hash, or store diffs. This is the point where a thin feature set that already has multi-tenant isolation beats a pile of one-off Lambdas—you want one place to tag cost to a tenant.
Points 4–7 — attach, payback, credit burn, and price fences
Point 4 is attach rate: percentage of paid seats that consume any AI credit in 14 days. Below 30%, you priced a novelty. Point 5 is payback on the included credits: included COGS should be recovered inside the first invoice, not the third. Point 6 is credit burn velocity: median credits per active day. If burn is spiky, you need anomaly caps (see below). Point 7 is the price fence: a cheap model for drafts, an expensive model behind a toggle, and a hard ceiling per day. Fences are how we kept p95 margin above 55% without rage-quitting power users.
Those seven points fit on one dashboard. If your current boilerplate cannot emit tenant-scoped usage events into Stripe, you will spreadsheet this until it hurts. That is a product problem, not a finance problem.
How to Price LLM Products Without Lighting Money on Fire
Pricing is the only lever that is faster than a model upgrade. I tested seats, pure usage, and hybrid on the same codebase. Hybrid won on cash collected and on support tickets.
Seat versus usage versus hybrid pricing
Seats are easy to sell and easy to bankrupt you. Pure usage (pass-through tokens) is honest and unsellable; buyers cannot forecast it, and you become a worse-UX OpenAI. Hybrid is a seat that includes a credit bundle, then overage. Credits abstract the vendor. You still meter tokens internally. Stripe usage-based billing and billing meters are the rails; your app must emit the quantities. I map 1 credit ≈ a successful “document action” whose blended COGS is about $0.02, then sell packs with a 3–4× markup so usage gross margin lands near 70% after waste.
Do not hide the meter. Show remaining credits in the app. The tenants who can see the number burn more sanely than the ones who think it is magic.
Token markup vs outcome-based pricing for AI features
Raw token markup (“we charge 4× OpenAI”) fails in sales because the buyer does not think in tokens. Outcome-based pricing (“$1.50 per accepted contract clause,” “$0.20 per classified ticket”) maps to value and lets you swap models without a price-list fight. Internally you still convert the outcome back to tokens so you know if the outcome is underwater. When Claude list prices moved on Anthropic’s page, our clause price did not change; our router did. That is the point of an outcome.
If you cannot define an outcome, you are not ready to productize the model. Ship a credit, not a token, and define the credit as the smallest successful job.
Guardrails: caps, caching, and model routing
Three guardrails paid for themselves in the first month:
- Daily tenant cap with a friendly 429 and an upsell, not a surprise invoice.
- Prompt caching for the static preamble and schema, using the vendor’s cache APIs linked above.
- A router: mini/flash/haiku for classification and drafts, sonnet/4o-class for the final artifact.
Gemini and Vertex flash-class models made the cheap lane real. Bedrock mattered for two enterprise tenants who needed a VPC story, not because it was cheaper. Route on task, not on brand loyalty. Revisit routes when list prices change—set a calendar reminder, not a Slack vibe.
Instrumentation Checklist: Meter Tokens Before You Invoice
If you cannot answer “what did tenant X cost yesterday?” in one query, you do not have AI SaaS metrics. You have hope. Meter before the stream starts, persist after the usage object returns, invoice from the same events.
Event schema for token usage
Store provider usage, not your estimate. Persist model, feature, attempt, tokens, cache, and computed cost_usd at write time so historical reports survive a price change. Price tables change; snapshots on the event do not.
export type LlmMeterEvent = {
tenantId: string;
requestId: string;
model: string;
feature: "chat" | "agent" | "embed" | "eval";
inputTokens: number;
outputTokens: number;
cachedInputTokens: number;
costUsd: number;
billedUnits: number;
at: string;
};
billedUnits is the customer-facing credit. costUsd is your COGS. Never mix them in one column. I emit this event in the same transaction that records the artifact, so a crash cannot leak unbilled work or uncosted work.
Stripe meters and TanStack Ship billing hooks
Stripe wants a meter event with a numeric value and a customer mapping. Your job is to convert billedUnits into that value on a schedule that matches the subscription (hourly is plenty). Read Stripe’s usage-based guide and meters once, then stop reinventing invoices. On our stack I would rather start from a SaaS boilerplate that already has auth, tenants, and Stripe hooks than wire billing from a greenfield Next.js app while the model bill is on fire. That is the honest pitch for TanStack Ship: it does not magically compute LLM gross margin, but it gives you the tenant and billing skeleton so this meter is a feature, not a rewrite. If you are comparing kits, the compare page is the place to be skeptical—pick the one that will not fight you on webhooks.
SELECT
date_trunc('day', at) AS day,
tenant_id,
SUM(cost_usd) AS token_cogs,
SUM(billed_units) * 0.02 AS usage_revenue,
1 - (SUM(cost_usd) / NULLIF(SUM(billed_units) * 0.02, 0)) AS usage_gross_margin
FROM llm_meter_events
WHERE at >= now() - interval '30 days'
GROUP BY 1, 2
HAVING SUM(cost_usd) / NULLIF(SUM(billed_units) * 0.02, 0) > 0.45;
That query is the p95 tripwire. Mail it to yourself. Do not wait for the vendor invoice.
Anomaly detection for runaway prompts and agent loops
Agent loops are how 41% margin happens overnight. Cap tool-round trips per request (we use 6). Cap output tokens in the API call, not in a Terms of Service paragraph. Alert when a tenant’s cost_usd exceeds 3× their 7-day median before noon. Most incidents we caught were a stuck retry on JSON mode, not a malicious user. The tokenizer will not save you here; a max_tokens parameter and a circuit breaker will. Put the breaker next to the meter, not in a weekend dashboard.
More playbooks live in the blog if you are wiring this beside auth and orgs rather than as a one-off script.
FAQ: AI SaaS Metrics Token Costs, Gross Margin, and Pricing
What gross margin should an AI SaaS target?
I target 70% usage gross margin on the median tenant and refuse to go live below 55% on the p95 tenant after model COGS. Classic 80% SaaS gross margin is available only if included credits are small, cache hit rates are real, and the expensive model is fenced. If you are bundling “unlimited GPT-4-class” into a $49 seat, you do not have a margin target. You have a countdown.
How do I convert tokens into a customer-facing credit unit?
Pick the smallest successful job (a draft, a classified ticket, a redlined clause). Measure blended token COGS for that job across a week of production, including retries. Set 1 credit so that COGS is about 25–30% of the credit’s retail value. Sell packs, not million-token blocks. Keep the token math internal and on the event (costUsd vs billedUnits). Recalibrate when OpenAI or Anthropic list prices move more than 20%.
Should I pass through model price changes?
Pass through to your COGS table the same day. Do not pass through to the customer price the same day unless you sold raw tokens. Outcome and credit prices should move on a quarterly cadence with a changelog. Use the lag to route to a cheaper model. Buyers who accepted a credit price will not accept a surprise 30% token tariff; they will churn and screenshot your app.
How does TanStack Ship help with usage billing?
TanStack Ship will not pick your model or your markup. It will give you the SaaS bones—auth, tenants, Stripe-shaped billing—so you can emit meter events instead of inventing subscriptions. If your blocker is “I still do not have a tenant id to hang costUsd on,” start there. If your blocker is “I do not know what a token is,” this teardown is the work; the boilerplate is just how you ship the work without a second rewrite. Read features and pricing with that distinction in mind.
Ship the meter, then ship the product
AI pricing is not a slide. It is a meter, a margin floor, and a fence around the expensive model. I learned that on a $4,870 token bill against $12,400 MRR, and I would rather you learn it from this 7-point teardown. Instrument ai saas metrics token costs per tenant, price a credit or an outcome, and hold 70% usage gross margin like it is a product requirement—because it is.
If you need a SaaS starting point that already speaks tenants and Stripe so you can spend your week on the router and the cap instead of OAuth, TanStack Ship is the kit I would actually use for that job. No affiliate, no vendor side-letter, no claim that a boilerplate replaces unit economics. It just lets you ship the pricing you just designed.