In the agentic era, inference latency becomes business latency. A security agent investigating an intrusion or a coding agent implementing and testing a fix can make dozens of sequential model calls. Every wait compounds. Faster inference shortens the loop between reasoning, action, and recovery.
Twenty engines in two weeks #
That is why INT21 chose ultra-fast inference as its first target. Two engineers directed the generation of 20 engines across seven model categories in two weeks: autoregressive language, diffusion language, speech recognition, speech generation, OCR, image generation, and video generation.
The engines use Rust with C++ and CUDA components, keeping model execution out of Python and PyTorch. That is one design choice, not an explanation for every gain. The measurements capture the combined effect of the runtime, kernels, execution strategy, and serving path.
We benchmarked a 15-model subset against SGLang and vLLM, including their Omni implementations where applicable. The interactive explorer above keeps each result attached to its workload, metric, hardware allocation, and limitations. There is no single score for this suite: decoding, usable audio, and complete image generation measure different things.
Faster decoding is only part of the story #
The text-generation workload uses approximately 8K input tokens and 1K output tokens, one request at a time, across three rounds after warmup. Rates measure client-observed decoding after the first token—or the first completed block for DiffusionGemma—and include subsequent serving overhead.
INT21 leads on the listed decode rates for all six text models. MiMo reaches 1,308 tokens/s, versus 540 for tuned SGLang and 1,011 for vLLM. GLM-5.3 reaches 753.2 tokens/s, versus 528.3 and 465.0. The advantage is narrower on GLM-5.3-Flash: 748.0 tokens/s, versus 627.1 and 713.2.
GLM-5.3-Flash also shows why decoding speed cannot stand in for completion time. INT21 takes 1,527 ms to produce its first token, compared with 242 ms for SGLang and 163 ms for vLLM. On this workload, that initial delay outweighs the decode lead. Prefill and serving latency remain clear optimization targets.
DiffusionGemma needs a different interpretation. It emits 256-token blocks; its rate is calculated as the remaining 768 tokens divided by the mean time from the first block to completion, checked against all 96 measured requests per engine. This excludes prefill and first-block generation, but it is not an isolated denoiser measurement. INT21 uses FP8 experts while the displayed baselines use BF16. Its reported advantage over a separately tested vLLM FP8 configuration is 1.81×.
MiMo uses the tuned SGLang configuration, and Qwen uses the faster tested vLLM nightly build. For DeepSeek, the supplied rates of 1,844 / 830.4 / 791.0 tokens/s imply 2.22× / 2.33×; the draft’s 2.66× / 2.62× ratios require reconciliation with the underlying runs. The explorer consistently calculates ratios from the displayed measurements, which may differ slightly from ratios computed using unrounded values.
Speech needs a clock that matches the experience #
For recognition, we measure request latency. For speech generation, we measure when usable audio arrives. For OCR, we measure when the complete response arrives. Treating all three as “inference speed” would hide what a user actually experiences. Fish Audio S2 Pro returns its first PCM audio in 48.55 ms on average, versus 254.88 ms for SGLang-Omni and 91.53 ms for vLLM-Omni. But the official serving paths differ in prompting and sampling, vLLM’s optional attention fast path fell back, and all three engines experienced early simulated playback gaps. Fast first audio does not guarantee smooth playback.
Higgs Audio uses a common 800 ms audio buffer because the engines return different initial chunk sizes. INT21 reaches that buffer in a median 103 ms, versus 129 ms and 394 ms. Its first-audio-only advantages are smaller: 1.25× / 1.60×.
The recognition and OCR tests have important limits. Both ASR tests returned empty cleaned transcripts from synthetic noise, so they measure audio processing and short-output latency, not speech accuracy or sustained transcription. MiMo ASR uses core SGLang and vLLM-Omni; Qwen ASR uses SGLang-Omni and core vLLM. Both TTS tests use Omni baselines, while OCR uses the core frameworks. On multi-page OCR, vLLM’s output quality diverged substantially: the latency ratios do not establish equivalent extraction quality.
Complete audio generation, on the same GPU allocation #
Hunyuan AuK provides a comparison of complete audio-generation requests: AuK base, 32 steps, one B200 per engine, concurrency one. Each result is the median of five warmed runs.
For requested audio durations of 3, 6, and 12 seconds, INT21 completes requests in 127, 139, and 188 ms, respectively. Across those durations, that is 1.73–2.30× faster than SGLang-Omni and 5.39–7.54× faster than vLLM-Omni. Switch durations in the explorer to see how each engine’s latency changes.
Images: lower latency, with more GPUs #
The image results show substantial deployment-latency reductions, but the hardware allocations differ. Qwen-Image-2.1 uses eight B200s for INT21 and four for each baseline. SenseNova U1.5 uses eight for INT21, one for SGLang, and two for vLLM-Omni.
At 1K resolution, Qwen-Image-2.1 completes in a mean 0.473 seconds, versus 2.077 and 1.347 seconds. SenseNova completes in 0.459 seconds, versus 6.361 and 3.142 seconds. These results describe the tested deployments; they do not establish equal-resource efficiency or cost advantages. The selected baseline recipes do not establish universal GPU limits either.
SenseNova’s vLLM 2K result is preliminary, with three measured requests. INT21 used its documented NCCL fallback. Image results are means; the video result is a median.
For MiniMax-H3 VDN, INT21 and SGLang each use eight B200s to generate a 14.375-second, 1344×768 video with stereo audio in eight steps. Median completion times are 9.21 and 10.20 seconds, a 1.11× advantage. No official vLLM VDN support was found for this test.
What would make this an AlphaGo moment? #
The strongest finding is the combination of development speed and measured performance across architectures. Engineers can direct AI-generated implementations across a broad set of models, then use benchmark evidence to guide the next optimization cycle.
The limits matter. These are primarily warm, concurrency-one benchmarks, many with synthetic inputs; some audio and OCR tests ran on shared nodes. They do not establish production-load performance or equivalent output quality across the suite. The benchmark summary supplied for this article does not include the underlying run reports, so the individual runs have not been independently verified here.
The next gains have concrete targets: prefill, serving overhead, streaming cadence, and hardware utilization. Broader validation must also test realistic inputs, concurrent load, and output quality.
An “AlphaGo moment” for inference would mean AI changing how we build the systems that run AI—and how quickly we can improve them. These results are an encouraging step toward that possibility. They also make the remaining work visible.