Benchmarking Qwen 3.8 27B Across Inference Providers: Together, Fireworks, Doubleword, and g factor A benchmark of Qwen 3.8 27B across inference providers found Together AI FP8 TP2 delivered the highest aggregate output throughput at every concurrency level tested, reaching 2,701.47 tokens per second at concurrency 64 versus 2,611.88 for g factor MTP4, 2,649.71 for Doubleword Public FP8, 2,462.44 for vanilla vLLM FP8 DP2, and 1,145.08 for Fireworks FP8 DP2. The tests, run by the gft-studio research platform on a fixed 2× NVIDIA H100 80GB SXM baseline with prompts averaging ~564 input tokens and 128 generated output tokens at temperature 0 and seed 42, covered concurrency levels 1, 2, 4, 8, 16, 32, and 64 using AIPerf 0.12.0. Fireworks FP8 DP2 plateaued near 1,145 tokens per second from concurrency 32 onward, while Doubleword and vanilla vLLM scaled through concurrency 64. If you look at vendor landing pages or benchmarks on social media, every inference provider claims to be “the fastest engine on Earth.” You see sleek bar charts showing thousands of tokens per second, single-digit Time-To-First-Token, and promises of dramatic cost savings. Then you deploy a 27-billion parameter reasoning model like Qwen 3.8 27B into production, route real multi-turn traffic through those endpoints, and practical systems trade-offs immediately emerge. Single-stream latency behaves very differently from high-concurrency batch throughput. Hardware interconnects dictate whether Tensor Parallelism flies or grinds to a halt. And subtle gateway interpretations of reasoning tokens can quietly balloon your generation budgets. Over the past several weeks, we ran an exhaustive series of empirical benchmarks across our research platform gft-studio to evaluate Qwen 3.8 27B across dedicated infrastructure and leading inference providers: Together AI, Fireworks AI, Nebius, Doubleword, and g factor . Below, we share the verified engineering telemetry: how parallel topology Tensor Parallelism vs. Data Parallelism shapes decode latency, how next-generation B200 hardware scales over H100 baselines, the real-world sweet spot of multi-token speculative decoding MTP4 vs. MTP8 , how prefix caching and prompt speculation behave on structured workflows, and what happens to latency when concurrency pushes to 64 parallel streams. 1. The Experimental Setup: Isolating Real Performance To make an inference benchmark meaningful, you have to eliminate confounding variables. Comparing a 7B model on FP8 with a 70B model on BF16 tells you nothing. Comparing an API called from a laptop in London with a server hosted in Oregon tells you about transit latency, not engine throughput. Here is how we standardized our test harness: - The Model: Qwen/Qwen3.8-27B and its official FP8 quantized variant , pinned to tokenizer revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 . - The Workload: Standardized AIPerf 0.12.0 test suites streaming over public HTTPS. Prompts averaged ~564 input tokens , requesting exactly 128 generated output tokens at temperature 0 and random seed 42. - Concurrencies Tested: Standard evaluation across concurrency levels 1, 2, 4, 8, 16, 32, and 64. Low-concurrency cells ran 16 warmup requests followed by 60 strictly measured requests across three separate repetitions 180 measured requests per cell . High-concurrency cells c16, c32, c64 ran 64 warmup requests and 256 measured requests per repeat. - The Hardware Baseline: A fixed hardware budget of 2× NVIDIA H100 80GB SXM GPUs per target except where explicitly comparing next-generation B200 accelerators . 2. Qwen 3.8 27B FP8 Throughput: Fireworks vs. Together vs. vLLM c1–c64 In real-world serving, systems rarely operate at a single concurrency point. Early morning traffic might see solitary interactive queries, while daytime peaks bombard your cluster with dozens of simultaneous streams. Below is the full empirical comparison across Together AI, Fireworks AI FP8 , Doubleword, vanilla standalone vLLM, and g factor with MTP4 speculative decoding: | Qwen 3.8 27B · Aggregate Output Throughput tok/s Across Concurrency 1–64 | | | | | | |---|---|---|---|---|---| | Concurrency | g factor MTP4 | Together FP8 TP2 | Fireworks FP8 DP2 | Doubleword Public FP8 | Vanilla vLLM FP8 DP2 | |---|---|---|---|---|---| | c = 1 | 133.00 | 189.61 | 114.49 | 65.86 | 69.29 | | c = 2 | 250.65 | 350.57 | 227.15 | 138.48 | 136.81 | | c = 4 | 437.41 | 612.35 | 401.46 | 291.29 | 253.51 | | c = 8 | 769.13 | 1033.44 | 640.02 | 534.71 | 459.16 | | c = 16 | 1252.34 | 1590.84 | 1056.65 | 935.52 | 836.18 | | c = 32 | 1904.97 | 2579.22 | 1146.75 | 1716.53 | 1601.44 | | c = 64 | 2611.88 | 2701.47 | 1145.08 | 2649.71 | 2462.44 | Looking across this spectrum reveals three distinct operational regimes: 1. Low Concurrency c = 1 to 2 : Together AI leads single-stream speed 189.62 tok/s at c1, 350.57 tok/s at c2 . Because Together serves the model via single-node Tensor Parallelism TP2 , both H100s collaborate on every token decode step. g factor with MTP4 delivers 133.00 tok/s at c1 and 250.65 tok/s at c2, comfortably outperforming Fireworks 114.49 / 227.15 tok/s and Vanilla vLLM 69.29 / 136.81 tok/s . 2. Medium Concurrency c = 4 to 8 : As requests multiply, multi-replica data parallelism DP2 hits its stride. g factor reaches 437.41 tok/s at c4 and 769.13 tok/s at c8 , pulling ahead of Fireworks DP2 640.02 tok/s and Vanilla vLLM 459.16 tok/s . 3. High Concurrency Saturation c = 16 to 64 : At c32 and c64, all modern engines push past 1,500 to 2,600 tokens per second. Together reaches 2,674 tok/s, Doubleword reaches 2,650 tok/s, and g factor reaches 2,612 tok/s. But raw throughput at high concurrency is only half the story—you also have to examine the latency bill. The Latency Bill Under Load It is easy to generate thousands of tokens per second if you let requests sit in a queue. What matters for interactive user experience is Time-To-First-Token TTFT : how long the user stares at a blank screen before text begins to stream. | Time-To-First-Token Under Load · Mean p95 TTFT Milliseconds | | | | | |---|---|---|---|---| | Concurrency | Together FP8 | Fireworks FP8 | Doubleword FP8 | Vanilla vLLM FP8 | |---|---|---|---|---| | c = 1 | 141.6 ms | 1088.4 ms | 1093.1 ms | 278.0 ms | | c = 2 | 174.3 ms | 390.4 ms | 1059.8 ms | 468.2 ms | | c = 4 | 217.1 ms | 435.3 ms | 826.8 ms | 538.0 ms | | c = 8 | 221.8 ms | 1323.3 ms | 786.7 ms | 656.9 ms | | c = 16 | 269.1 ms | 1034.9 ms | 919.5 ms | 705.6 ms | | c = 32 | 245.8 ms | 2687.1 ms | 975.5 ms | 736.0 ms | | c = 64 | 1701.7 ms | 7714.6 ms | 1598.6 ms | 1195.3 ms | Notice what happens between c32 and c64. On Fireworks, throughput stays essentially flat 1,146 tok/s → 1,145 tok/s , while p95 TTFT surges from 2.69 seconds to 7.71 seconds . The GPUs are fully saturated; adding more requests in flight simply queues them up at the door without producing more tokens per second. 3. Architectural Topology: Tensor Parallelism vs. Independent Replicas The most important architectural lesson from our benchmark is that two GPUs do not make a system; how those two GPUs are connected makes the system . Why Together Won Single-Stream Latency: High-Speed NVLink Look at the single-stream results: Together achieved 189.6 tok/s and an Inter-Token Latency ITL of just 3.9 milliseconds , compared to 69–133 tok/s on independent single-GPU replicas. Why? Because Together deployed Tensor Parallelism TP2 inside a single physical server. In TP2, every linear layer in Qwen 3.8 27B is sliced across both GPUs. For every single token decode step, GPU 0 and GPU 1 compute their respective matrix slices and exchange intermediate activations via an all-reduce collective. At concurrency 1, both H100s collaborate on that single user’s request simultaneously. In Data Parallelism DP2 , by contrast, GPU 0 handles User A while GPU 1 handles User B. At concurrency 1, User A only uses one GPU, with decode speed physically bounded by that single chip’s memory bandwidth 3.35 TB/s on H100 SXM . Evaluating Cross-Node Tensor Parallelism Seeing Together’s single-stream TP2 speed, we tested an experiment: what happens if you run TP2 across two separate 1× H100 cloud instances connected via standard datacenter virtual networking VPC rather than intra-node NVLink? The result demonstrated the severe cost of network barrier latency: | Concurrency | DP2 2 Separate Nodes | Cross-Node TP2 Over Network | Single-Node TP2 NVLink | Network Interconnect Penalty | |---|---|---|---|---| | c = 1 | 73.71 tok/s | 24.98 tok/s | 189.61 tok/s | 2.95x slower than DP2 | | c = 4 | 274.64 tok/s | 56.47 tok/s | 612.35 tok/s | 4.86x slower than DP2 | | c = 8 | 483.19 tok/s | 75.75 tok/s | 1,033.44 tok/s | 6.38x slower than DP2 | At concurrency 8, cross-node TP2 throughput collapsed from 483 tok/s down to 75.75 tok/s . The physical explanation is straightforward: During single-token decoding, GEMV computations for a 27B model take only 5 to 10 microseconds . Within a single chassis over NVLink, exchanging activations takes ~2 microseconds over 900 GB/s channels. Across separate physical servers over standard cloud networking, that same transfer requires 800 to 2,000 microseconds . The GPUs spent over 95% of their execution time stalled at network barriers waiting for TCP packets. The Golden Rule of Topology: Never run Tensor Parallelism across physical machines unless you have dedicated multi-rail InfiniBand. If your GPUs reside on separate nodes, always deploy independent replicas with data parallelism DP . 4. The Hardware Leap: What Happens on NVIDIA B200? While H100 remains the workhorse of enterprise inference, next-generation NVIDIA Blackwell B200 accelerators are entering production. We benchmarked Fireworks AI running Qwen 3.8 27B on a dedicated 2× B200 deployment: | Hardware Generation Leap · Fireworks Dedicated · Qwen 3.8 27B | | | | |---|---|---|---| | Concurrency | 2× H100 BF16 Fireworks | 2× B200 BF16 Fireworks | Observed Hardware Speedup | |---|---|---|---| | c = 1 | 97.09 tok/s | 153.38 – 161.32 tok/s | 1.58x – 1.66x | | c = 4 | 347.12 tok/s | 571.82 – 597.11 tok/s | 1.65x – 1.72x | | c = 8 | 622.18 tok/s | 979.49 – 1,001.32 tok/s | 1.57x – 1.61x | Moving from H100 to B200 delivered a clean 1.6x to 1.7x throughput increase with zero code changes. This speedup is directly explained by memory hardware specifications: - NVIDIA H100 SXM features 3.35 TB/s of HBM3 memory bandwidth. - NVIDIA B200 features 8.00 TB/s of ultra-dense HBM3e bandwidth a 2.38x hardware leap . Because memory-bound autoregressive decoding scales near-linearly with memory bandwidth, B200 allows a 2-GPU cluster to cross the 1,000 tok/s barrier even in unquantized 16-bit precision. 5. Case Study: Reasoning Token Budgets Across API Gateways During our evaluation of Qwen3.8-27B-FP8 on Doubleword’s public API, we encountered an instructive systems interaction that illustrates how reasoning-capable models interface with standard throughput benchmarks. We configured our AIPerf test harness with standard parameters to measure decode throughput at a fixed length: {"max completion tokens": 128, "temperature": 0} While standard non-reasoning requests complete in 1 to 2 seconds for 128 tokens, these initial requests ran for approximately 247 seconds . An inspection of the returned payload counters explained the extended duration: | Metric | Configured Parameter | Returned API Telemetry | |---|---|---| | Reasoning Tokens