11:39
2026-08-25
huggingface.co
machine-learning
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
Multiverse Computing's new Quantization-Aware Healing (QAH) method, applied to a GPT-OSS 120B model compressed to 60B parameters and quantized to MXFP4, produces a 4-bit model that outperforms its fulโฆ