{"slug": "what-can-10-buy-across-different-llm-apis", "title": "What Can $10 Buy Across Different LLM APIs?", "summary": "A developer compared how many requests a $10 budget buys across three LLM APIs using a fixed 4,000-input/1,000-output token workload, finding GPT-6 Luna yields roughly 11,111 requests, Gemini 3.8 Flash about 1,481, and Claude Sonnet 5.5 around 556. The analysis notes that output tokens account for over half the per-request cost despite being only 20% of the tokens, and that at 100,000 requests the same workload costs about $90, $675, and $1,800 respectively. The developer argues token price answers only cost per request, not cost per completed task, which requires separate evaluation.", "body_md": "LLM API pricing pages usually quote prices per million tokens.\n\nThat is useful for billing, but it is not always the easiest way to think about cost when building an application.\n\nA question I find more intuitive is:\n\nIf I have a $10 API budget, how many real requests can I make?\n\nThe answer depends heavily on the ratio between input and output tokens.\n\nLet's use one simple workload and compare a few models.\n\nAssume each API request contains:\n\n```\n4,000 input tokens\n1,000 output tokens\n```\n\nThat gives us a 5,000-token request with an 80/20 input-output split.\n\nFor this comparison, I am using the pricing currently recorded in my dataset:\n\n| Model | Input / 1M tokens | Output / 1M tokens | \n|---|---|---|\n| GPT-6 Luna | $0.10 | $0.50 | \n| Gemini 3.8 Flash | $0.75 | $3.75 | \n| Claude Sonnet 5.5 | $2.00 | $10.00 | \n\nThese are standard token rates and do not include special long-context tiers or other processing modes.\n\nNow let's turn those prices into something more concrete.\n\nFor one request:\n\n```\nInput:\n4,000 / 1,000,000 × $0.10\n= $0.0004\n\nOutput:\n1,000 / 1,000,000 × $0.50\n= $0.0005\n\nTotal:\n$0.0009 per request\n```\n\nWith a $10 budget:\n\n```\n$10 / $0.0009\n≈ 11,111 requests\n```\n\nThat is roughly:\n\n**44.4 million input tokens**\n\nand\n\n**11.1 million output tokens**\n\nfor the $10 budget.\n\nNow run exactly the same workload through Gemini 3.8 Flash.\n\n```\nInput:\n4,000 / 1,000,000 × $0.75\n= $0.003\n\nOutput:\n1,000 / 1,000,000 × $3.75\n= $0.00375\n\nTotal:\n$0.00675 per request\n```\n\nWith $10:\n\n```\n$10 / $0.00675\n≈ 1,481 requests\n```\n\nSame workload.\n\nSame $10.\n\nVery different request capacity.\n\nFor Claude Sonnet 5.5:\n\n``` php\nInput:\n4,000 / 1,000,000 × $2\n= $0.008\n\nOutput:\n1,000 / 1,000,000 × $10\n= $0.01\n\nTotal:\n$0.018 per request\n```\n\nA $10 budget gives approximately:\n\n```\n$10 / $0.018\n≈ 556 requests\n```\n\nSo for this particular workload:\n\n| Model | Approx. requests for $10 | \n|---|---|\n| GPT-6 Luna | 11,111 | \n| Gemini 3.8 Flash | 1,481 | \n| Claude Sonnet 5.5 | 556 | \n\nThat difference is large enough to matter when an application moves from experimentation to production traffic.\n\nBut this table does **not** mean GPT-6 Luna is automatically the best choice.\n\nImagine Model A needs one request to solve a coding task.\n\nModel B is cheaper per token but needs:\n\nModel B may still produce a higher total cost per completed task.\n\nThis is why I think there are two separate questions developers should ask.\n\nThat is mostly arithmetic.\n\nThat is an evaluation problem.\n\nToken price answers the first question.\n\nIt does not answer the second.\n\nThere is another interesting detail in the example.\n\nThe request contains four times as many input tokens as output tokens:\n\n```\n4,000 input\n1,000 output\n```\n\nYet for all three models in this example, the output portion costs more per token.\n\nFor Claude Sonnet 5.5, the single request costs:\n\n```\nInput:  $0.008\nOutput: $0.010\n```\n\nSo only 20% of the tokens account for more than half of the cost.\n\nThis makes output control surprisingly important.\n\nIf an agent produces verbose intermediate reasoning, oversized summaries, unnecessary code explanations, or repeated generated content, reducing output may save more money than aggressively shortening the prompt.\n\nA $10 experiment may not sound important.\n\nNow imagine the application processes 100,000 of these requests.\n\nUsing the same 4,000-input / 1,000-output workload:\n\n```\nGPT-6 Luna:\n100,000 × $0.0009\n≈ $90\n\nGemini 3.8 Flash:\n100,000 × $0.00675\n≈ $675\n\nClaude Sonnet 5.5:\n100,000 × $0.018\n≈ $1,800\n```\n\nAt that point model choice is no longer a minor implementation detail.\n\nIt becomes part of the product economics.\n\nThe examples above assume all input is charged at the standard input rate.\n\nReal applications can be different.\n\nMany agent workloads repeatedly send information such as:\n\nWhen an API supports discounted cached input, part of that input may become significantly cheaper.\n\nSo a more complete calculation looks like:\n\n```\nTotal cost =\n  normal input tokens × input price\n+ cached input tokens × cached price\n+ output tokens × output price\n```\n\nFor applications with large reusable prompts, caching can materially change the economics.\n\nAnother thing I try not to assume is:\n\nThis model supports a one-million-token context window, therefore one million tokens always cost the normal input rate.\n\nProviders can have special pricing rules for large contexts, different processing tiers, batch workloads, caching, or other modes.\n\nA model's context window tells you what is technically possible.\n\nIt does not necessarily tell you what that workload will cost.\n\nThat is why pricing comparisons should preserve the provider source and the date when the pricing was verified.\n\nI was repeatedly doing calculations like these manually, so I built a simple [LLM API Cost Calculator](https://iilib.com/workbench/llm-api-cost).\n\nInstead of asking only for a model price, it lets you enter a workload:\n\n```\ninput tokens / request\ncached input tokens / request\noutput tokens / request\nnumber of requests\n```\n\nThere is also a budget-oriented way to think about the problem: start with a fixed amount of money and estimate the workload capacity.\n\nI deliberately keep the pricing source and verification date next to the results because a perfectly accurate calculation based on outdated pricing is still a wrong answer.\n\nThe underlying rates are also available separately in the [LLM API Pricing reference](https://iilib.com/research/llm-api-pricing).\n\nAfter looking at enough pricing tables, I have become less interested in:\n\nWhich model has the cheapest token?\n\nThe metric I would rather know is:\n\nHow much does it cost to successfully complete one useful task?\n\nFor some applications, a very inexpensive model will win.\n\nFor others, paying more for a model that completes the task reliably in fewer steps may be cheaper overall.\n\nA useful evaluation therefore needs both:\n\n```\ncost per request\n×\nrequests per successful task\n```\n\nThat gets much closer to the economics of a real AI application than a simple price-per-million-token leaderboard.\n\nPricing changes frequently. The figures above reflect the pricing data I was using in October 2026. Always verify current provider pricing and pricing conditions before making production cost decisions.", "url": "https://wpnews.pro/news/what-can-10-buy-across-different-llm-apis", "canonical_source": "https://dev.to/adam_z/what-can-10-buy-across-different-llm-apis-4i1n", "published_at": "2026-10-07 02:37:32+00:00", "updated_at": "2026-10-07 02:47:43.981484+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-agents", "ai-tools"], "entities": ["GPT-6 Luna", "Gemini 3.8 Flash", "Claude Sonnet 5.5"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-can-10-buy-across-different-llm-apis", "markdown": "https://wpnews.pro/news/what-can-10-buy-across-different-llm-apis.md", "text": "https://wpnews.pro/news/what-can-10-buy-across-different-llm-apis.txt", "jsonld": "https://wpnews.pro/news/what-can-10-buy-across-different-llm-apis.jsonld"}}