High-performance single-GPU inference for selected model checkpoints and GPUs NInfer, a from-scratch C++/CUDA inference engine, achieves up to 1,313.8 aggregate decode tokens per second on a single NVIDIA GeForce RTX 5090 for the Qwen3.6-35B-A3B model at concurrency 8, with the Qwen3.6-27B NVFP4 profile reaching 1,146.9 tok/s. The engine supports explicitly registered Qwen checkpoints via local CLI or OpenAI-/Anthropic-compatible HTTP APIs. Selected checkpoints. Maximum single-GPU inference performance. NInfer is a from-scratch C++/CUDA inference engine for explicitly registered Qwen checkpoints on a single NVIDIA GeForce RTX 5090. It runs text, image, and video prompts through a local CLI or OpenAI-/Anthropic-compatible HTTP APIs. NInfer deliberately supports a closed set of model artifacts instead of acting as a general model runtime: | Model | Weights | NInfer artifact | Size | SHA-256 | |---|---|---|---|---| | groupwise-int qwen3 6 27b.ninfer 7b51600ffd10632b9660f56085efdd9b751d79733ad32036a652234b64bebe7b Qwen3.6-27B NVFP4 https://huggingface.co/neroued/Qwen3.6-27B-nvfp4-NInfer nvfp4 qwen3 6 27b nvfp4.ninfer bce5f00d066c0f20f1317bf1fdcb458264cf95837c3b1f3fbec163694627893a Qwen3.8-27B https://huggingface.co/neroued/Qwen3.8-27B-NInfer groupwise-int qwen3 8 27b.ninfer eec39564993d6e9c7d5e383382a760f093465c9d163ec9a1bd6b80199514bf3e Qwen3.6-35B-A3B https://huggingface.co/neroued/Qwen3.6-35B-A3B-NInfer groupwise-int qwen3 6 35b a3b.ninfer 1fb9ea0b5b8561e49d9604115ec89e5d9f2b6f6434e32c37c57fffd480a325d2 The two Qwen3.6-27B weight profiles bind to the registered qwen3 6 27b target; the version-2 artifact identity selects the profile without a separate runtime flag. Qwen3.8-27B is separately registered as qwen3 8 27b and shares the 27B execution package while using W8 token-embedding and full-output-head weights. The nvfp4 profile uses W4A4 Tensor Core MMA for prefill and A16 NVFP4 kernels for decode. All three 27B artifacts retain the same Text, Vision, MTP, prefix-reuse, CLI, and serving routes. The published measurements currently cover the three Qwen3.6 artifact profiles. Qwen3.8-27B is supported by current NInfer builds but is not yet included in the benchmark campaign. Saturated decode was measured on an RTX 5090 with INT8 group-64 KV cache, CUDA Graphs, MTP3, and one 8,192-token generation per active request. The values below are aggregate committed decode throughput from complete one-second intervals in which the actual decode batch remained equal to the configured concurrency. Each profile should be read independently. | Model profile | C=1 | C=2 | C=4 | C=8 | C8 / C1 | |---|---|---|---|---|---| Qwen3.6-27B groupwise-int | 185.8 tok/s | 247.0 tok/s | 309.5 tok/s | 535.0 tok/s | 2.88× | Qwen3.6-27B nvfp4 | 202.4 tok/s | 399.7 tok/s | 699.7 tok/s | 1,146.9 tok/s | 5.67× | Qwen3.6-35B-A3B groupwise-int | 593.0 tok/s | 877.7 tok/s | 1,166.0 tok/s | 1,313.8 tok/s | 2.22× | At C=8, Qwen3.6-35B-A3B reaches 1,313.8 aggregate decode tok/s . The 27B NVFP4 profile reaches 1,146.9 tok/s and 5.67× its C=1 throughput. The single-request corpus was measured on the same GPU with INT8 group-64 KV cache, CUDA Graphs, and a 1,024-token prefill chunk. Each reported fixture uses five fixed seeds after server warm-up. The two measured targets are reported independently and are not cross-target comparisons. The two 27B weight profiles are reported separately. Requests were submitted serially to a persistent server. Qwen3.6-35B-A3B - MTP0 at a 7,680-token prompt: 15,544.3 prefill tok/s and 271.1 decode tok/s . - MTP0 at a 260,096-token prompt: 5,157.1 prefill tok/s and 188.2 decode tok/s . - MTP3 long reasoning: 620.3–726.2 decode tok/s with 72.7–82.8% acceptance . - MTP3 structured output: 770.9 decode tok/s , 89.1% acceptance , and 3.67 tokens/round . Qwen3.6-27B groupwise-int - MTP0 at a 7,680-token prompt: 3,218.1 prefill tok/s and 77.6 decode tok/s . - MTP0 at a 260,096-token prompt: 1,614.8 prefill tok/s and 54.8 decode tok/s . - MTP3 long reasoning: 161.9–175.4 decode tok/s with 73.4–78.8% acceptance . - MTP3 structured output: 193.0 decode tok/s , 88.7% acceptance , and 3.66 tokens/round . Qwen3.6-27B nvfp4 - MTP0 at a 7,680-token prompt: 11,191.5 prefill tok/s and 86.4 decode tok/s . - MTP0 at a 260,096-token prompt: 2,510.6 prefill tok/s and 59.9 decode tok/s . - MTP3 long reasoning: 213.1–231.0 decode tok/s with 76.3–81.1% acceptance . - MTP3 structured output: 252.2 decode tok/s , 89.8% acceptance , and 3.69 tokens/round . - Against groupwise-int on the same corpus and runtime options: 3.48× the 7,680-token prefill throughput , 1.55× the 260,096-token prefill throughput , and 30–32% higher MTP3 decode throughput . See Performance /Neroued/ninfer/blob/master/docs/performance.md for the full methodology, variability, reproduction command, and per-fixture results. Capability scores were measured through NInfer's OpenAI-compatible serving route with thinking enabled, MTP=3, and EvalScope 1.9.0 0-shot, rule scoring, one sample per problem : | Model profile | AIME 2025 | AIME 2026 | GPQA-Diamond | |---|---|---|---| | Qwen3.6-27B NVFP4 /Neroued/ninfer/blob/master/model-cards/Qwen3.6-27B-nvfp4-NInfer/README.md Qwen3.6-35B-A3B groupwise-int /Neroued/ninfer/blob/master/model-cards/Qwen3.6-35B-A3B-NInfer/README.md Qwen3.8-27B is supported but has not yet been added to this published evaluation campaign. These are single-sample results under that NInfer evaluation profile, not pass@k. See the model cards and full performance document /Neroued/ninfer/blob/master/docs/performance.md for correct/total counts and evaluation notes. NInfer currently requires: - 64-bit Linux; - NVIDIA GeForce RTX 5090 sm 120a ; - NVIDIA driver support for CUDA 13.1 and the CUDA Toolkit 13.1 or newer; - CMake 3.28 or newer and a C++20-capable host compiler; pkg-config ;- FFmpeg development libraries: libavformat = 60 , libavcodec = 60 , libavutil = 58 , and libswscale = 7 ; libcurl = 7.85 ;- Ninja, when using the commands below. The build rejects CUDA architectures other than 120a . There is no install target or packaged binary distribution; NInfer is run from its source build tree. git clone https://github.com/Neroued/ninfer.git cd ninfer cmake -S . -B build -G Ninja -DCMAKE BUILD TYPE=Release cmake --build build --parallel The default configuration builds: build/apps/ninfer build/apps/ninfer-serve Tests, benchmarks, and maintainer tools are excluded from the default build. Build the runtime image on a 64-bit Linux host with an RTX 5090, a CUDA 13.1-compatible NVIDIA driver, Docker, and the NVIDIA Container Toolkit https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html . docker build --tag ninfer:local . Download a model into models/ as described below, then run the HTTP server: docker run --rm \ --gpus '"device=0"' \ --publish 8080:8080 \ --volume "$PWD/models:/models:ro" \ ninfer:local \ ninfer-serve /models/qwen3 6 27b.ninfer \ --host 0.0.0.0 Run the CLI from the same image: docker run --rm \ --gpus '"device=0"' \ --volume "$PWD/models:/models:ro" \ ninfer:local \ ninfer /models/qwen3 6 27b.ninfer \ --prompt "Explain prefill and decode in three sentences." \ --max-new 256 Use the Hugging Face CLI to download one of the registered artifacts: hf download neroued/Qwen3.6-27B-NInfer \ qwen3 6 27b.ninfer \ --local-dir models Or the 27B NVFP4 weight variant: hf download neroued/Qwen3.6-27B-nvfp4-NInfer \ qwen3 6 27b nvfp4.ninfer \ --local-dir models Or Qwen3.8-27B: hf download neroued/Qwen3.8-27B-NInfer \ qwen3 8 27b.ninfer \ --local-dir models Or: hf download neroued/Qwen3.6-35B-A3B-NInfer \ qwen3 6 35b a3b.ninfer \ --local-dir models Current NInfer builds accept only the version-2 artifact container, and all four downloads above are version 2. Migration applies only to Qwen3.6 artifacts downloaded before their version-2 publication; Qwen3.8-27B was published directly as version 2. Migrate an older exact local file in place: python3 -m tools.artifact.migrate v1 to v2 models/qwen3 6 27b.ninfer Use the same command with qwen3 6 27b nvfp4.ninfer or qwen3 6 35b a3b.ninfer for those artifacts. The migration updates only container metadata; it does not rewrite the weight payload. Alternatively, download the current version-2 file again from its Hugging Face repository. Each .ninfer file contains the weights and frontend resources needed by NInfer. It is not a Transformers checkpoint, Safetensors distribution, or GGUF file. Each artifact is complete, while GPU residency is fixed at process startup. Speculative decoding is disabled by default, so MTP/DFlash state and the optimized proposal head are not uploaded. Vision is also disabled by default, so its weights, Vision scratch phase, and frozen request-transient allocation are omitted. Add --vision to the CLI or server process that must accept image or video input. Disabled capabilities cannot be enabled by a later request. DFlash is available only for the 35B-A3B target and is text-only. ./build/apps/ninfer models/qwen3 6 27b.ninfer \ --prompt "Explain prefill and decode in three sentences." \ --max-context 16384 \ --max-new 256 \ --spec mtp --draft-tokens 3 \ --lm-head-draft Use --messages FILE instead of --prompt for chat history, images, or videos: ./build/apps/ninfer models/qwen3 6 27b.ninfer \ --messages examples/cli/messages/image chart.json \ --max-context 8192 \ --max-new 128 \ --vision Answer content is written to stdout. Loading progress, reasoning, timing, throughput, memory, and speculative-decoding statistics are written to stderr. See the CLI guide /Neroued/ninfer/blob/master/docs/cli.md and committed examples /Neroued/ninfer/blob/master/examples/cli for structured input and runtime options. ./build/apps/ninfer-serve models/qwen3 6 27b.ninfer \ --max-context 16384 \ --kv-capacity auto \ --max-concurrency 2 \ --spec mtp --draft-tokens 3 \ --lm-head-draft The public model ID defaults to the artifact's identity.model id ; use --model-id only to publish a deployment-specific alias. Then send an OpenAI-style request: curl http://127.0.0.1:8080/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "qwen3.6-27b", "messages": {"role": "user", "content": "Reply with one short sentence."} , "max tokens": 64 }' The server also implements OpenAI Responses Core typed Items, semantic SSE, local continuation state, and function calls plus Anthropic Messages, token counting, and multimodal input. See HTTP serving /Neroued/ninfer/blob/master/docs/serving.md . All three registered model IDs support: - text generation with thinking and non-thinking prompt modes; - image, multi-image, video, and mixed multimodal messages; - chunked prefill and CUDA Graph decode; - startup-bounded small-scale concurrent serving with true batched decode; - MTP speculative decoding with draft windows from one to five; - BF16 and INT8 group-64 KV cache; - model- and thinking-mode-aware official sampling defaults, with explicit greedy, temperature, top-k, top-p, min-p, and presence/frequency-penalty overrides; - compatible-prefix reuse; - OpenAI Responses Core, OpenAI Chat Completions, and Anthropic Messages, including streaming and usage accounting; - prompt-rendered function tools and parsed tool calls. The 35B-A3B target additionally supports text-only DFlash speculative decoding with draft windows from one to fifteen. - Only the four model id, weights id artifact identities listed above are accepted product identities. - Execution is specialized for one RTX 5090 and one CUDA device. - One Engine owns one resident model and supports a startup-fixed capacity of 1–8 active requests. Decode-ready requests are compacted at round boundaries and executed in one batched model traversal. - NInfer does not provide large-scale or preemptive continuous batching, priority/QoS scheduling, multi-GPU execution, CPU/GPU offload, or distributed serving. --max-context is the logical ceiling of each sequence and is configurable up to the registered models' native 262,144-token limit. --kv-capacity N explicitly sizes the shared Main Text KV pool for all active and retained sequences, while --kv-capacity auto selects the largest usable capacity from the memory remaining after weights are loaded while preserving 1 GiB of sizing headroom. Omission defaults to one --max-context worth of pages. The resolved pool is fixed at startup and is not divided statically among request lanes.- Tool calls are parsed and returned to the client; NInfer does not execute tools. - The C++ headers are used by the in-tree applications and are not distributed as an installed SDK. NInfer is licensed under the Apache License 2.0 /Neroued/ninfer/blob/master/LICENSE . The published artifacts are derived from Qwen/Qwen3.6-27B https://huggingface.co/Qwen/Qwen3.6-27B , Qwen/Qwen3.8-27B https://huggingface.co/Qwen/Qwen3.8-27B , and Qwen/Qwen3.6-35B-A3B https://huggingface.co/Qwen/Qwen3.6-35B-A3B . The 27B NVFP4 artifact also uses the fixed packed weights from rdtand/Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm https://huggingface.co/rdtand/Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm . These source repositories are distributed under Apache-2.0. Vendored dependencies retain their own license files under third party/ .