cd /news/ai-infrastructure/how-many-gpus-is-1m-b-t-tokens · home › topics › ai-infrastructure › article
[ARTICLE · art-146904] src=cedana.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

How many GPUs is 1M/B/T tokens?

Cedana published a tokens-to-GPUs calculator showing that serving 1 trillion tokens in one month (30 days) on Llama 3.3 70B requires approximately 367 H100 GPUs at a base case, with a range of 211 to 853. The calculator derives the figure by converting the monthly volume to 385,802 tokens per second, applying a 0.475 token-mix factor for 183,256 output-equivalent tokens per second, and dividing by 50% utilization and 1,000 tokens per second per H100. The same 1 trillion tokens per month works out to 815 A100 GPUs, 2,547 V100 GPUs, or 92 B200 GPUs at base case, and the largest single sensitivity is per-GPU throughput, which moves the H100 count from 244 at 1,500 tok/s to 814 at 450 tok/s.

read5 min views1 publishedOct 6, 2026
How many GPUs is 1M/B/T tokens?
Image: source

» Cedana / Capacity planning

Tokens have become the unit that many practitioners use to report AI growth. The trouble is that the numbers are hard to picture. Is a trillion tokens a month a lot? It depends. The same volume needs a different amount of hardware depending on the GPU, the model, and the workload (agentic coding, chat). So we built a calculator. Enter a token volume and pick a model. It tells you how many GPUs that works out to, with a range that shows how firm the estimate is. Every default links to its source, and you can swap in your own measurements.

1 trillion tokens in one month (30 days) on Llama 3.3 70B needs approximately 367 H100 GPUs. Confidence is moderate.

How the calculator gets the H100 number #

Each step changes one quantity. Read from left to right.

  1. 01Start1 trilliontokens in one month (30 days)
  2. 02÷ 2,592,000 seconds385,802tokens per second, average
  3. 03× 0.475 for the token mix183,256output-equivalent tokens per second
  4. 04÷ 50% utilization366,512tokens per second of installed capacity
  5. 05÷ 1,000 tok/s for each H100367H100 GPUs

Which assumption moves the H100 result most #

Each bar shows the GPU count when one assumption moves from its best case to its worst case. The other assumptions stay at the base value. The longest bar is the assumption to measure first.

244 at 1,500 tok/s

814 at 450 tok/s

229 at 80%

611 at 30%

251 at 0.1 of output

482 at 0.5 of output

GPU tok/s per GPU Tokens per GPU per day GPUs per copy Low Base High 8-GPU nodes
A100 450 40.93M 1 436 815 1,966 102
V100 144 13.10M 3 1,338 2,547 7,122 319
H100 1,000 90.95M 1 211 367 853 46
B200 4,000 363.79M 1 47 92 252 12

Reference speed is tokens per second for one GPU, as low, base, and high. "Measured" values come from public benchmarks across loose and strict latency targets. "Estimated" values have no direct benchmark.

Model Total / active Weights Reference speed B200 ÷ H100 Basis Source
Llama 3.1 8B 8B / 8B 8 GB H100: 3,000, 6,000, 12,500 2.3, 4, 6 Estimated. One measured point for the high value. morph »
Llama 3.3 70B 70B / 70B 70 GB H100: 450, 1,000, 1,500 2.3, 4, 6 Measured on H100 and B200. inferencex »cerebrium »
gpt-oss 120B 117B / 5.1B 65 GB H100: 740, 1,400, 2,600 6, 9, 12 Measured on H100 and B200. Weight size is not verified. inferencex »
DeepSeek V4 Flash 284B / 13B 160 GB B200: 1,500, 4,000, 9,000 5, 9, 14 Estimated. No benchmark found. Scaled from V4 Pro and gpt-oss by active parameters. datacamp »
MiniMax M3 428B 428B / n/a 428 GB H100: 157, 240, 537 3.5, 4.5, 5.3 Measured on H100 and B200. Weight size assumes 8-bit. inferencex »
DeepSeek R1 / V3 671B / 37B 671 GB H100: 23, 75, 266 12, 16, 21 Measured on H100 and B200. Weight size assumes 8-bit. inferencex »
GLM-5.3 753B / 40B 755 GB B200: 1,400, 3,000, 7,000 4, 8, 16 Measured on B200 only. Sources disagree: GLM-5 at FP4 gives 1,417 to 2,935, GLM-5.3 gives 4,329 to 10,243. H100 ratio is estimated. inferencex 5.3 »inferencex 5 fp4 »size »
DeepSeek V4 Pro 1600B / 49B 865 GB B200: 468, 1,000, 2,800 6, 14, 21 Measured on B200 only. H100 ratio is estimated from DeepSeek R1. inferencex »size »
Assumption Value Basis Source
A100 ÷ H100 0.35, 0.45, 0.60 H100 measured at 1.8x to 2.9x the A100 on 70B models hyperstack »perplexity »
V100 ÷ H100 0.10, 0.18, 0.25 Specifications only: 125 against 312 TFLOPS, 900 against 2,039 GB/s. No LLM serving benchmark. spheron »
Input cost 0.1, 0.3, 0.5 Derived from 128:128 and 2024:128 runs on 4x H100 (result 0.33) e2e networks »
GPU memory 32, 80, 80, 180 GB V100, A100, H100, B200. The fit check adds 10% to the weight size. inferencex »

Values read on 2 October 2026.

  • The result is a planning estimate. Expect an error of 2x in each direction for measured presets, and more for estimated presets. Measure your model on your hardware before you buy.
  • Speed changes with the latency target. The low and high values of each preset come from strict and loose latency targets.
  • For large mixture-of-experts models, the B200 lead over the H100 is 4x to 21x. FP4 support and memory size cause this. The ratio is specific to each model.
  • A100 and V100 results for the large models are theoretical. One model copy needs more GPUs than one node contains.
  • The V100 value has no measured LLM serving data. It comes from hardware specifications.
  • The range combines independent errors as a root sum of squares in log space. It is not a statistical confidence interval.
  • The calculator does not include prompt caching, speculative decoding, reasoning-token overhead, failures, or regional duplication.
── more in #ai-infrastructure 4 stories · sorted by recency
── more on @cedana 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-many-gpus-is-1m-…] indexed:0 read:5min 2026-10-06 · —