Quad R9700's on AM4 Shenanigans
A user on the Level1Techs forum reported running four AMD Radeon R9700 GPUs on an AM4 platform to requantize a DeepSeek model to MXFP4 and run it in the vLLM engine, with the first layer of a 48-layer…
A user on the Level1Techs forum reported running four AMD Radeon R9700 GPUs on an AM4 platform to requantize a DeepSeek model to MXFP4 and run it in the vLLM engine, with the first layer of a 48-layer…
Baseten's technical breakdown of LLM inference efficiency identifies two categories of engineering choices: those that trade latency for throughput, such as batch sizing, tensor parallelism, expert pa…
Multiverse Computing's new Quantization-Aware Healing (QAH) method, applied to a GPT-OSS 120B model compressed to 60B parameters and quantized to MXFP4, produces a 4-bit model that outperforms its ful…
A team has run the full Kimi K3 model, a 2.8-trillion-parameter mixture-of-experts (MoE) model, on 80 RTX 5090 GPUs achieving 20 tokens per second single-stream inference on day one without tuning, ma…
Researchers propose MXAttention, a data-free post-training quantization framework for MXFP4 attention that introduces Universal Optimal Scaling (UOS) with a distribution-independent optimal scaling bo…
Vulkan 1.4.356 introduces the VK_EXT_shader_ocp_microscaling_types extension, adding support for Open Compute Project Microscaling MX data types (MXFP4, MXFP6, MXFP8, MXINT8) to improve machine learni…