How to Structure a Bring-Your-Own-Model AI Pricing Contract Before You Sign Enterprise buyers are increasingly asking AI vendors to let them bring their own large language models (BYOM) to cut per-token spend, and most startup SaaS teams lack contract language for it. Salesforce's Agentforce BYOM program through Amazon Bedrock demonstrates a two-tier structure where the platform fee stays fixed and inference cost passes through. Contracts must separate platform fees from model inference costs, write usage caps in compute or request units, and name liability for customer-chosen model outputs. Enterprise buyers are now asking AI vendors to let them swap in their own LLM to cut per-token spend, and most startup SaaS teams are walking into that negotiation with no contract language ready for it. - A bring-your-own-model AI pricing contract has to separate your platform fee from model inference cost, or you lose your margin the day a customer switches models - Usage caps must be written in compute or request units, not dollars, because BYOM removes your ability to price by token markup - Model liability clauses need to name who is responsible when a customer's chosen model produces a bad output, a hallucinated figure, or a compliance breach - Salesforce's Agentforce BYOM program through Amazon Bedrock shows the two-tier structure that works: platform fee stays fixed, inference cost passes through Here's what's happening. A procurement lead at a mid-size insurer or bank sits down with your sales team and says they already have an enterprise agreement with OpenAI or Anthropic, they already pay for Azure OpenAI Service credits, and they don't want to pay you a second markup on top of tokens they're already buying elsewhere. They want to plug their own model into your product. You say yes, because saying no loses the deal. Then legal and finance realize nobody wrote the contract for this. That's the gap right now. Vendor contracts for AI SaaS were built around a simple assumption: you buy the model, you mark it up, you bill the customer a blended rate per seat or per token. BYOM breaks that assumption completely. The customer brings the compute. Your revenue model, and your legal exposure, both change shape. Most founders have never negotiated this because almost nobody has published what the clauses should actually say. In a standard SaaS-plus-AI deal, you control the whole stack. You pick the model, you eat the API cost, you price the product with enough margin to cover it, and if OpenAı raises prices you quietly adjust your own pricing at renewal. The customer never sees the seams. BYOM removes that control. The customer's model choice, their rate limits, their region, their fine-tuning, all sit outside your contract with your inference provider. You're now selling orchestration, workflow, and interface on top of infrastructure you don't own and can't fully audit. That changes three things at once: how you price, what you cap, and who's liable when the model gets something wrong. Get any one of them wrong and you either bleed margin or inherit risk that isn't yours to carry. Split the platform fee from the model cost, always The single most important structural decision is separating what you charge for your software from what the customer pays for inference. Bundle those two together and BYOM customers will negotiate your platform fee down to nothing, because they'll compare your all-in price against a competitor's bare API rate and assume you're the expensive part. Salesforce's approach with Agentforce is worth looking at directly, because it's a real, live example of a large vendor solving this exact problem. Salesforce lets enterprise customers bring their own model through Amazon Bedrock rather than defaulting to Salesforce's own model layer, and the commercial structure keeps the Agentforce platform fee billed per conversation or per action separate from the underlying model inference cost, which the customer pays directly to their cloud provider. Salesforce doesn't try to mark up tokens it never touches. It prices the orchestration layer, and only the orchestration layer. Copy that shape. Your contract should have two line items, not one: a platform or orchestration fee that's yours to set and defend, and a pass-through or bring-your-own inference cost that the customer controls. Do not let a customer's BYOM request collapse both into a single negotiated number. The moment you do, you've lost the ability to explain what you're actually charging for. Usage caps have to move from dollars to compute units Most startup pricing tiers are written in dollar thresholds, like Also read: How Do Crypto Options Vaults Work, and Who Really Keeps the Yield https://startupfortune.com/how-do-crypto-options-vaults-work-and-who-really-keeps-the-yield/ • How to Structure a Token Vesting Schedule Before TGE Without Tanking Your Price https://startupfortune.com/how-to-structure-a-token-vesting-schedule-before-tge-without-tanking-your-price/ • How Restaking Slashing Insurance Actually Works, and Who Gets Paid Last https://startupfortune.com/how-restaking-slashing-insurance-actually-works-and-who-gets-paid-last/