cd /news/large-language-models/what-can-10-buy-across-different-llm… · home › topics › large-language-models › article
[ARTICLE · art-146491] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

What Can $10 Buy Across Different LLM APIs?

A developer compared how many requests a $10 budget buys across three LLM APIs using a fixed 4,000-input/1,000-output token workload, finding GPT-6 Luna yields roughly 11,111 requests, Gemini 3.8 Flash about 1,481, and Claude Sonnet 5.5 around 556. The analysis notes that output tokens account for over half the per-request cost despite being only 20% of the tokens, and that at 100,000 requests the same workload costs about $90, $675, and $1,800 respectively. The developer argues token price answers only cost per request, not cost per completed task, which requires separate evaluation.

by read5 min views1 publishedOct 7, 2026

LLM API pricing pages usually quote prices per million tokens.

That is useful for billing, but it is not always the easiest way to think about cost when building an application.

A question I find more intuitive is:

If I have a $10 API budget, how many real requests can I make?

The answer depends heavily on the ratio between input and output tokens.

Let's use one simple workload and compare a few models.

Assume each API request contains:

4,000 input tokens
1,000 output tokens

That gives us a 5,000-token request with an 80/20 input-output split.

For this comparison, I am using the pricing currently recorded in my dataset:

Model Input / 1M tokens Output / 1M tokens
GPT-6 Luna $0.10 $0.50
Gemini 3.8 Flash $0.75 $3.75
Claude Sonnet 5.5 $2.00 $10.00

These are standard token rates and do not include special long-context tiers or other processing modes.

Now let's turn those prices into something more concrete.

For one request:

Input:
4,000 / 1,000,000 × $0.10
= $0.0004

Output:
1,000 / 1,000,000 × $0.50
= $0.0005

Total:
$0.0009 per request

With a $10 budget:

$10 / $0.0009
≈ 11,111 requests

That is roughly:

44.4 million input tokens

and

11.1 million output tokens

for the $10 budget.

Now run exactly the same workload through Gemini 3.8 Flash.

Input:
4,000 / 1,000,000 × $0.75
= $0.003

Output:
1,000 / 1,000,000 × $3.75
= $0.00375

Total:
$0.00675 per request

With $10:

$10 / $0.00675
≈ 1,481 requests

Same workload.

Same $10.

Very different request capacity.

For Claude Sonnet 5.5:

Input:
4,000 / 1,000,000 × $2
= $0.008

Output:
1,000 / 1,000,000 × $10
= $0.01

Total:
$0.018 per request

A $10 budget gives approximately:

$10 / $0.018
≈ 556 requests

So for this particular workload:

Model Approx. requests for $10
GPT-6 Luna 11,111
Gemini 3.8 Flash 1,481
Claude Sonnet 5.5 556

That difference is large enough to matter when an application moves from experimentation to production traffic.

But this table does not mean GPT-6 Luna is automatically the best choice.

Imagine Model A needs one request to solve a coding task.

Model B is cheaper per token but needs:

Model B may still produce a higher total cost per completed task.

This is why I think there are two separate questions developers should ask.

That is mostly arithmetic.

That is an evaluation problem.

Token price answers the first question.

It does not answer the second.

There is another interesting detail in the example.

The request contains four times as many input tokens as output tokens:

4,000 input
1,000 output

Yet for all three models in this example, the output portion costs more per token.

For Claude Sonnet 5.5, the single request costs:

Input:  $0.008
Output: $0.010

So only 20% of the tokens account for more than half of the cost.

This makes output control surprisingly important.

If an agent produces verbose intermediate reasoning, oversized summaries, unnecessary code explanations, or repeated generated content, reducing output may save more money than aggressively shortening the prompt.

A $10 experiment may not sound important.

Now imagine the application processes 100,000 of these requests.

Using the same 4,000-input / 1,000-output workload:

GPT-6 Luna:
100,000 × $0.0009
≈ $90

Gemini 3.8 Flash:
100,000 × $0.00675
≈ $675

Claude Sonnet 5.5:
100,000 × $0.018
≈ $1,800

At that point model choice is no longer a minor implementation detail.

It becomes part of the product economics.

The examples above assume all input is charged at the standard input rate.

Real applications can be different.

Many agent workloads repeatedly send information such as:

When an API supports discounted cached input, part of that input may become significantly cheaper.

So a more complete calculation looks like:

Total cost =
  normal input tokens × input price
+ cached input tokens × cached price
+ output tokens × output price

For applications with large reusable prompts, caching can materially change the economics.

Another thing I try not to assume is:

This model supports a one-million-token context window, therefore one million tokens always cost the normal input rate.

Providers can have special pricing rules for large contexts, different processing tiers, batch workloads, caching, or other modes.

A model's context window tells you what is technically possible.

It does not necessarily tell you what that workload will cost.

That is why pricing comparisons should preserve the provider source and the date when the pricing was verified.

I was repeatedly doing calculations like these manually, so I built a simple LLM API Cost Calculator.

Instead of asking only for a model price, it lets you enter a workload:

input tokens / request
cached input tokens / request
output tokens / request
number of requests

There is also a budget-oriented way to think about the problem: start with a fixed amount of money and estimate the workload capacity.

I deliberately keep the pricing source and verification date next to the results because a perfectly accurate calculation based on outdated pricing is still a wrong answer.

The underlying rates are also available separately in the LLM API Pricing reference.

After looking at enough pricing tables, I have become less interested in:

Which model has the cheapest token?

The metric I would rather know is:

How much does it cost to successfully complete one useful task?

For some applications, a very inexpensive model will win.

For others, paying more for a model that completes the task reliably in fewer steps may be cheaper overall.

A useful evaluation therefore needs both:

cost per request
×
requests per successful task

That gets much closer to the economics of a real AI application than a simple price-per-million-token leaderboard.

Pricing changes frequently. The figures above reflect the pricing data I was using in October 2026. Always verify current provider pricing and pricing conditions before making production cost decisions.

── more in #large-language-models 4 stories · sorted by recency
── more on @gpt-6 luna 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-can-10-buy-acro…] indexed:0 read:5min 2026-10-07 · —