How many GPUs is 1M/B/T tokens? Cedana published a tokens-to-GPUs calculator showing that serving 1 trillion tokens in one month (30 days) on Llama 3.3 70B requires approximately 367 H100 GPUs at a base case, with a range of 211 to 853. The calculator derives the figure by converting the monthly volume to 385,802 tokens per second, applying a 0.475 token-mix factor for 183,256 output-equivalent tokens per second, and dividing by 50% utilization and 1,000 tokens per second per H100. The same 1 trillion tokens per month works out to 815 A100 GPUs, 2,547 V100 GPUs, or 92 B200 GPUs at base case, and the largest single sensitivity is per-GPU throughput, which moves the H100 count from 244 at 1,500 tok/s to 814 at 450 tok/s. » Cedana / Capacity planning Tokens to GPUs Calculator Tokens have become the unit that many practitioners use to report AI growth. The trouble is that the numbers are hard to picture. Is a trillion tokens a month a lot? It depends. The same volume needs a different amount of hardware depending on the GPU, the model, and the workload agentic coding, chat . So we built a calculator. Enter a token volume and pick a model. It tells you how many GPUs that works out to, with a range that shows how firm the estimate is. Every default links to its source, and you can swap in your own measurements. 1 trillion tokens in one month 30 days on Llama 3.3 70B needs approximately 367 H100 GPUs. Confidence is moderate. How the calculator gets the H100 number Each step changes one quantity. Read from left to right. 1. 01Start1 trilliontokens in one month 30 days 2. 02÷ 2,592,000 seconds385,802tokens per second, average 3. 03× 0.475 for the token mix183,256output-equivalent tokens per second 4. 04÷ 50% utilization366,512tokens per second of installed capacity 5. 05÷ 1,000 tok/s for each H100367H100 GPUs Which assumption moves the H100 result most Each bar shows the GPU count when one assumption moves from its best case to its worst case. The other assumptions stay at the base value. The longest bar is the assumption to measure first. 244 at 1,500 tok/s 814 at 450 tok/s 229 at 80% 611 at 30% 251 at 0.1 of output 482 at 0.5 of output | GPU | tok/s per GPU | Tokens per GPU per day | GPUs per copy | Low | Base | High | 8-GPU nodes | |---|---|---|---|---|---|---|---| | A100 | 450 | 40.93M | 1 | 436 | 815 | 1,966 | 102 | | V100 | 144 | 13.10M | 3 | 1,338 | 2,547 | 7,122 | 319 | | H100 | 1,000 | 90.95M | 1 | 211 | 367 | 853 | 46 | | B200 | 4,000 | 363.79M | 1 | 47 | 92 | 252 | 12 | Reference speed is tokens per second for one GPU, as low, base, and high. "Measured" values come from public benchmarks across loose and strict latency targets. "Estimated" values have no direct benchmark. | Model | Total / active | Weights | Reference speed | B200 ÷ H100 | Basis | Source | |---|---|---|---|---|---|---| | Llama 3.1 8B | 8B / 8B | 8 GB | H100: 3,000, 6,000, 12,500 | 2.3, 4, 6 | Estimated. One measured point for the high value. | morph » https://www.morphllm.com/vllm-benchmarks | | Llama 3.3 70B | 70B / 70B | 70 GB | H100: 450, 1,000, 1,500 | 2.3, 4, 6 | Measured on H100 and B200. | inferencex » https://inferencex.semianalysis.com/compare/llama-3-3-70b-b200-vs-h100 cerebrium » https://cerebrium.ai/blog/benchmarking-vllm-sglang-tensorrt-for-llama-3-1-api | | gpt-oss 120B | 117B / 5.1B | 65 GB | H100: 740, 1,400, 2,600 | 6, 9, 12 | Measured on H100 and B200. Weight size is not verified. | inferencex » https://inferencex.semianalysis.com/compare/gptoss-120b-b200-vs-h100 | | DeepSeek V4 Flash | 284B / 13B | 160 GB | B200: 1,500, 4,000, 9,000 | 5, 9, 14 | Estimated. No benchmark found. Scaled from V4 Pro and gpt-oss by active parameters. | datacamp » https://www.datacamp.com/blog/deepseek-v4 | | MiniMax M3 428B | 428B / n/a | 428 GB | H100: 157, 240, 537 | 3.5, 4.5, 5.3 | Measured on H100 and B200. Weight size assumes 8-bit. | inferencex » https://inferencex.semianalysis.com/compare/minimax-m3-b200-vs-h100 | | DeepSeek R1 / V3 | 671B / 37B | 671 GB | H100: 23, 75, 266 | 12, 16, 21 | Measured on H100 and B200. Weight size assumes 8-bit. | inferencex » https://inferencex.semianalysis.com/compare/deepseek-r1-b200-vs-h100 | | GLM-5.3 | 753B / 40B | 755 GB | B200: 1,400, 3,000, 7,000 | 4, 8, 16 | Measured on B200 only. Sources disagree: GLM-5 at FP4 gives 1,417 to 2,935, GLM-5.3 gives 4,329 to 10,243. H100 ratio is estimated. | inferencex 5.3 » https://inferencex.semianalysis.com/compare/glm-5-3-b200-vs-h200 inferencex 5 fp4 » https://inferencex.semianalysis.com/compare-precision/glm-5-1-b200-fp4-vs-fp8 size » https://artificialanalysis.ai/models/glm-5-3 | | DeepSeek V4 Pro | 1600B / 49B | 865 GB | B200: 468, 1,000, 2,800 | 6, 14, 21 | Measured on B200 only. H100 ratio is estimated from DeepSeek R1. | inferencex » https://inferencex.semianalysis.com/compare/deepseek-v4-b200-vs-h200 size » https://www.morphllm.com/deepseek-v4 | | Assumption | Value | Basis | Source | |---|---|---|---| | A100 ÷ H100 | 0.35, 0.45, 0.60 | H100 measured at 1.8x to 2.9x the A100 on 70B models | hyperstack » https://www.hyperstack.cloud/technical-resources/performance-benchmarks/llm-inference-benchmark-comparing-nvidia-a100-nvlink-vs-nvidia-h100-sxm perplexity » https://hub-prod.perplexity.ai/hub/blog/turbocharging-llama-2-70b-with-nvidia-h100 | | V100 ÷ H100 | 0.10, 0.18, 0.25 | Specifications only: 125 against 312 TFLOPS, 900 against 2,039 GB/s. No LLM serving benchmark. | spheron » https://www.spheron.network/blog/nvidia-a100-vs-v100/ | | Input cost | 0.1, 0.3, 0.5 | Derived from 128:128 and 2024:128 runs on 4x H100 result 0.33 | e2e networks » https://docs.e2enetworks.com/docs/tir/benchmarks/inference benchmarks/ | | GPU memory | 32, 80, 80, 180 GB | V100, A100, H100, B200. The fit check adds 10% to the weight size. | inferencex » https://inferencex.semianalysis.com/chips/b200 | Values read on 2 October 2026. - The result is a planning estimate. Expect an error of 2x in each direction for measured presets, and more for estimated presets. Measure your model on your hardware before you buy. - Speed changes with the latency target. The low and high values of each preset come from strict and loose latency targets. - For large mixture-of-experts models, the B200 lead over the H100 is 4x to 21x. FP4 support and memory size cause this. The ratio is specific to each model. - A100 and V100 results for the large models are theoretical. One model copy needs more GPUs than one node contains. - The V100 value has no measured LLM serving data. It comes from hardware specifications. - The range combines independent errors as a root sum of squares in log space. It is not a statistical confidence interval. - The calculator does not include prompt caching, speculative decoding, reasoning-token overhead, failures, or regional duplication.