What Can $10 Buy Across Different LLM APIs? A developer compared how many requests a $10 budget buys across three LLM APIs using a fixed 4,000-input/1,000-output token workload, finding GPT-6 Luna yields roughly 11,111 requests, Gemini 3.8 Flash about 1,481, and Claude Sonnet 5.5 around 556. The analysis notes that output tokens account for over half the per-request cost despite being only 20% of the tokens, and that at 100,000 requests the same workload costs about $90, $675, and $1,800 respectively. The developer argues token price answers only cost per request, not cost per completed task, which requires separate evaluation. LLM API pricing pages usually quote prices per million tokens. That is useful for billing, but it is not always the easiest way to think about cost when building an application. A question I find more intuitive is: If I have a $10 API budget, how many real requests can I make? The answer depends heavily on the ratio between input and output tokens. Let's use one simple workload and compare a few models. Assume each API request contains: 4,000 input tokens 1,000 output tokens That gives us a 5,000-token request with an 80/20 input-output split. For this comparison, I am using the pricing currently recorded in my dataset: | Model | Input / 1M tokens | Output / 1M tokens | |---|---|---| | GPT-6 Luna | $0.10 | $0.50 | | Gemini 3.8 Flash | $0.75 | $3.75 | | Claude Sonnet 5.5 | $2.00 | $10.00 | These are standard token rates and do not include special long-context tiers or other processing modes. Now let's turn those prices into something more concrete. For one request: Input: 4,000 / 1,000,000 × $0.10 = $0.0004 Output: 1,000 / 1,000,000 × $0.50 = $0.0005 Total: $0.0009 per request With a $10 budget: $10 / $0.0009 ≈ 11,111 requests That is roughly: 44.4 million input tokens and 11.1 million output tokens for the $10 budget. Now run exactly the same workload through Gemini 3.8 Flash. Input: 4,000 / 1,000,000 × $0.75 = $0.003 Output: 1,000 / 1,000,000 × $3.75 = $0.00375 Total: $0.00675 per request With $10: $10 / $0.00675 ≈ 1,481 requests Same workload. Same $10. Very different request capacity. For Claude Sonnet 5.5: php Input: 4,000 / 1,000,000 × $2 = $0.008 Output: 1,000 / 1,000,000 × $10 = $0.01 Total: $0.018 per request A $10 budget gives approximately: $10 / $0.018 ≈ 556 requests So for this particular workload: | Model | Approx. requests for $10 | |---|---| | GPT-6 Luna | 11,111 | | Gemini 3.8 Flash | 1,481 | | Claude Sonnet 5.5 | 556 | That difference is large enough to matter when an application moves from experimentation to production traffic. But this table does not mean GPT-6 Luna is automatically the best choice. Imagine Model A needs one request to solve a coding task. Model B is cheaper per token but needs: Model B may still produce a higher total cost per completed task. This is why I think there are two separate questions developers should ask. That is mostly arithmetic. That is an evaluation problem. Token price answers the first question. It does not answer the second. There is another interesting detail in the example. The request contains four times as many input tokens as output tokens: 4,000 input 1,000 output Yet for all three models in this example, the output portion costs more per token. For Claude Sonnet 5.5, the single request costs: Input: $0.008 Output: $0.010 So only 20% of the tokens account for more than half of the cost. This makes output control surprisingly important. If an agent produces verbose intermediate reasoning, oversized summaries, unnecessary code explanations, or repeated generated content, reducing output may save more money than aggressively shortening the prompt. A $10 experiment may not sound important. Now imagine the application processes 100,000 of these requests. Using the same 4,000-input / 1,000-output workload: GPT-6 Luna: 100,000 × $0.0009 ≈ $90 Gemini 3.8 Flash: 100,000 × $0.00675 ≈ $675 Claude Sonnet 5.5: 100,000 × $0.018 ≈ $1,800 At that point model choice is no longer a minor implementation detail. It becomes part of the product economics. The examples above assume all input is charged at the standard input rate. Real applications can be different. Many agent workloads repeatedly send information such as: When an API supports discounted cached input, part of that input may become significantly cheaper. So a more complete calculation looks like: Total cost = normal input tokens × input price + cached input tokens × cached price + output tokens × output price For applications with large reusable prompts, caching can materially change the economics. Another thing I try not to assume is: This model supports a one-million-token context window, therefore one million tokens always cost the normal input rate. Providers can have special pricing rules for large contexts, different processing tiers, batch workloads, caching, or other modes. A model's context window tells you what is technically possible. It does not necessarily tell you what that workload will cost. That is why pricing comparisons should preserve the provider source and the date when the pricing was verified. I was repeatedly doing calculations like these manually, so I built a simple LLM API Cost Calculator https://iilib.com/workbench/llm-api-cost . Instead of asking only for a model price, it lets you enter a workload: input tokens / request cached input tokens / request output tokens / request number of requests There is also a budget-oriented way to think about the problem: start with a fixed amount of money and estimate the workload capacity. I deliberately keep the pricing source and verification date next to the results because a perfectly accurate calculation based on outdated pricing is still a wrong answer. The underlying rates are also available separately in the LLM API Pricing reference https://iilib.com/research/llm-api-pricing . After looking at enough pricing tables, I have become less interested in: Which model has the cheapest token? The metric I would rather know is: How much does it cost to successfully complete one useful task? For some applications, a very inexpensive model will win. For others, paying more for a model that completes the task reliably in fewer steps may be cheaper overall. A useful evaluation therefore needs both: cost per request × requests per successful task That gets much closer to the economics of a real AI application than a simple price-per-million-token leaderboard. Pricing changes frequently. The figures above reflect the pricing data I was using in October 2026. Always verify current provider pricing and pricing conditions before making production cost decisions.