{"slug": "neural-nova-gpu-optimization-benchmarks-for-llm-workloads", "title": "Neural Nova – GPU optimization benchmarks for LLM workloads", "summary": "Neural Nova published GPU optimization benchmarks for LLM workloads, reporting throughput and cost gains across four models running on vLLM. Qwen3-235B-A22B on 8× NVIDIA H100-80GB posted +138.7% token/s and +58% cost savings, while GLM-5.2 on 8× AMD Instinct MI325X posted +26.8% token/s and +21.1% cost savings. Gemma-4-31B-it on 1× NVIDIA H100-80GB gained +66.0% token/s with +40% cost savings, and GPT-OSS-120B on 1× NVIDIA H100-80GB gained +24.5% token/s with +20% cost savings.", "body_md": "# Performance Benchmarks\n\nExplore Optimized AI models recipes across GPUs, frameworks,\n\nand deployment configurations.\n\nBENCHMARKED MODELS\n\nREASONING\n\nvLLM · 8× NVIDIA H100-80GB\n\n### Qwen3-235B-A22B\n\n+138.7% token/s\n\n+58% Cost Savings\n\nQwen3-235B-A22B · benchmark\n\nREASONING\n\nvLLM · 8× AMD Instinct MI325X\n\n### GLM-5.2\n\n+26.8% token/s\n\n+21.1% Cost Savings\n\nGLM-5.2 · benchmark\n\nMULTI-MODAL\n\nvLLM · 1× NVIDIA H100-80GB\n\n### Gemma-4-31B-it\n\n+66.0% token/s\n\n+40% Cost Savings\n\nGemma-4-31B-it · benchmark\n\nOPEN-WEIGHT\n\nvLLM · 1× NVIDIA H100-80GB\n\n### GPT-OSS-120B\n\n+24.5% token/s\n\n+20% Cost Savings\n\nGPT-OSS-120B · benchmark\n\nPRODUCTION-READY PERFORMANCE\n\n## Benchmark with confidence. Deploy with clarity.\n\nUse validated performance data to select the right model for your production workload.\n\nTalk to Our Team", "url": "https://wpnews.pro/news/neural-nova-gpu-optimization-benchmarks-for-llm-workloads", "canonical_source": "https://www.neural-nova.com/benchmark", "published_at": "2026-09-10 20:30:50+00:00", "updated_at": "2026-09-10 20:43:12.985048+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-chips", "mlops"], "entities": ["Neural Nova", "vLLM", "NVIDIA H100-80GB", "AMD Instinct MI325X", "Qwen3-235B-A22B", "GLM-5.2", "Gemma-4-31B-it", "GPT-OSS-120B"], "alternates": {"html": "https://wpnews.pro/news/neural-nova-gpu-optimization-benchmarks-for-llm-workloads", "markdown": "https://wpnews.pro/news/neural-nova-gpu-optimization-benchmarks-for-llm-workloads.md", "text": "https://wpnews.pro/news/neural-nova-gpu-optimization-benchmarks-for-llm-workloads.txt", "jsonld": "https://wpnews.pro/news/neural-nova-gpu-optimization-benchmarks-for-llm-workloads.jsonld"}}