19:31
2026-07-09
marktechpost.com
artificial-intelligence
Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput
NVIDIA AI released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed hybrid MoE LLM that achieves up to 2.14x server throughput on 8xB200 nodes and enables 8 concurrent 1M-token requests on a single H100, …