AI pricing is understood now but the token is still just a cost AI pricing is now understood, but the token remains a cost, not a price, argues Arnon Shimoni in an expert opinion piece. Shimoni identifies three working strategies: bundling tokens into plans or seats to drive adoption, pricing on the value tokens produce (e.g., Fin's $0.99 per resolution), and gating usage to prevent runaway costs. He notes that AI add-on launches surged in 2025, with 1/5th of companies adding them in Q2 2025, often as an additional shape. AI pricing is understood now but the token is still just a cost Insights Read time: 8 min Arnon Shimoni ✓ Expert opinion Open pretty much any AI pricing guide published since 2025 including some of my old ones and it starts the same way: pricing AI is the hardest, most unsolved problem in software and no one has figured it out . Well, I disagree by this point. The reason it still feels unsolved is one substitution people keep making, over and over and over again. Because you look at a token and that's considered the price. A token isn't a rpice. Yes, a token is what the model costs you to run an input. The output costs you more output tokens run 3-8x input and the outcome, the thing your customer wants from you, isn't measured in tokens at all. Once you stop pricing the cost of your own input and start pricing the value on top of it, the "playbook" if you can call it that is becoming a bit clearer. The standard advice we see all over makes the substitution for you, as every explainer on token pricing lands on the same instruction: work out what a request costs to run, then set the token price to cover it, plus a margin. Cost-plus. It's clean, it's defensible, and it quietly assumes the token is the thing you sell. For a foundation-model API, fine... the token is the product. For an application built on top, cost-plus on the token caps your price at your own COGS curve, right when the customer is buying an outcome worth many times that. So what does work now? What's working right now I see three main moves which you can make, but a combo of all three is also possible. Bundle the token into a plan or seat to buy adoption Notion folded AI into its higher-priced plans instead of selling it as a separate add-on. Cursor sells seats with included usage Pro at $20 a month, Ultra at $200 with far more headroom . Perplexity does the same across its $20 to $325 tiers. You eat the token cost and bet that adoption pays it back through retention and expansion. It's a bet on your own product. Not my favourite structure as a consumer, but for land-and-expand it works. Price on the value the tokens produce Obviously we all know Fin by now with their $0.99 per resolution, on top of a $49 base that includes the first 50 - we also have Zendesk, AgentForce, Justt, Chargeflow and so many more doing this. The tokens are the vendors' cost to manage, invisible to the buyer. When you can define the outcome and measure it, this tracks value so much better than anything else on the list. Gate it so it can't run away from you This is the move that makes the other two safe, and it's one that's easy to skip over. When you create a bundled plan with no ceiling and no entitlement gating, you're inviting users to consume much more than what they pay for. Cursor's tiers are gating in practice with each seat having a usage threshold and the next tier up is priced for the accounts that outgrow it. When I had a look through our partner PricingSaaS's data, https://pulse.pricingsaas.com/ I can really see where the need to add AI features separately shows up. AI add-on launches were the hot thing in 2025 - much more intense than in 2024 with 1/5th of companies adding it in Q2 2025 - often as an additional shape. Add-ons, when done correctly at least, with some form of gating can protect both the vendor and the buyer. You hold your margin tight by billing for an add-on, and the customer gets a voluntary credit/balance they can watch and budget against. Unfortunately, when done badly it can become some invisible thing that trips you up. A lot of times customers experience as a trap and they eventually churn. It's not a ceiling per se, it's a live balance that the customer shouldn't have to get stuck on. Pricing KPIs you should know The normal SaaS metrics/KPIs are less important with AI… While I said AI pricing is understood, this is how you'd track it to actually prove it. Luckily, there are just five of them, and MRR isn't one of them: KPI | What it tells you | Rough benchmark | |---|---|---| | Your AI gross margin but opposite | AI gross margin rose from 41% 2024 to 52% 2026 , floor now around 60-65% | | The accounts that make money | Watch the spread, not the average | | Credit breakage on one side, overrun on the other side | Aim for a band, not a ceiling | | How much of the usage drives expansion | Rising is usually the healthy sign | | Does the model work longer-term | Usage/hybrid 115-130%+ vs 95-105% for flat | That second one is the one I like the most, and it's one that gets neglected a bunch. Like many other companies you may find more comfort in an average - but your top 5% of users are either your best case studies or your margin losers, and the average will never tell you which... How to measure events One big habit we noticed at Solvimon is some companies meter every billable event in real time and attribute the cost to the customer and the task, instead of reconciling it after the invoice gets finalized - and that's a winning activity. We believe you can't effectively bundle what you find difficult to track and it gets harder as you want to align value. It's damn near impossible to charge for an outcome you haven't defined and measured. We try to do this underneath your app - so Solvimon's metering ../usage-metering turns raw events into billable, attributable quantities per customer per outcome or task. FAQ Is AI pricing actually a solved problem? I'd say yes because the models are understood: bundle, meter, value/ outcome-based ../glossary/outcome-based-pricing , hybrid ./hybrid-pricing-is-the-default-now-heres-the-data . What's unsolved for most teams is the plumbing underneath, i.e., metering ../usage-metering and gating consumption per customer accurately enough to trust the margin. The strategy is known, doesn't mean it looks identical for everyone. Should I charge my customers per token? Only if the token is genuinely your product e.g., a foundation-model API or raw inference . For 99% of applications built on top, the token is your cost of input, and billing it straight to the customer passes your supplier's meter through at a markup which feels bad for the customer. Bundle it or price the outcome instead. What's a healthy token-cost-to-revenue ratio? It maps inversely to gross margin. AI-native gross margins climbed from about 41% in 2024 to 52% in 2026, with the durable floor landing around 60-65%, and hybrid SaaS-AI products sitting higher. If token cost eats more than 40% of revenue, look at the pricing model before you blame the model provider. What is outcome-based pricing? You charge when the AI delivers a defined result, e.g., a resolved ticket or a completed task. Intercom's Fin $0.99 per resolution is the clearest live example. It tracks value closely, and it asks you to define the outcome precisely and handle the cases where the AI half-succeeds. How do I measure gross margin per customer? Attribute token and infra cost to each account, not just aggregate COGS, then set it against what the account pays. This needs per-customer metering. The blended average hides your loss-making whales, which are usually the accounts you're proudest of. Ready for billing v2? Solvimon is monetization infrastructure for companies that have outgrown billing v1. One system, entire lifecycle, built by the team that did this at Adyen.