Neural Nova – GPU optimization benchmarks for LLM workloads Neural Nova published GPU optimization benchmarks for LLM workloads, reporting throughput and cost gains across four models running on vLLM. Qwen3-235B-A22B on 8× NVIDIA H100-80GB posted +138.7% token/s and +58% cost savings, while GLM-5.2 on 8× AMD Instinct MI325X posted +26.8% token/s and +21.1% cost savings. Gemma-4-31B-it on 1× NVIDIA H100-80GB gained +66.0% token/s with +40% cost savings, and GPT-OSS-120B on 1× NVIDIA H100-80GB gained +24.5% token/s with +20% cost savings. Performance Benchmarks Explore Optimized AI models recipes across GPUs, frameworks, and deployment configurations. BENCHMARKED MODELS REASONING vLLM · 8× NVIDIA H100-80GB Qwen3-235B-A22B +138.7% token/s +58% Cost Savings Qwen3-235B-A22B · benchmark REASONING vLLM · 8× AMD Instinct MI325X GLM-5.2 +26.8% token/s +21.1% Cost Savings GLM-5.2 · benchmark MULTI-MODAL vLLM · 1× NVIDIA H100-80GB Gemma-4-31B-it +66.0% token/s +40% Cost Savings Gemma-4-31B-it · benchmark OPEN-WEIGHT vLLM · 1× NVIDIA H100-80GB GPT-OSS-120B +24.5% token/s +20% Cost Savings GPT-OSS-120B · benchmark PRODUCTION-READY PERFORMANCE Benchmark with confidence. Deploy with clarity. Use validated performance data to select the right model for your production workload. Talk to Our Team