High-performance single-GPU inference for selected model checkpoints and GPUs
NInfer, a from-scratch C++/CUDA inference engine, achieves up to 1,313.8 aggregate decode tokens per second on a single NVIDIA GeForce RTX 5090 for the Qwen3.6-35B-A3B model at concurrency 8, with the…