These Researchers Just Shrunk an AI Model and Somehow Made It Smarter
Multiverse Computing researchers published a method called Quantization-Aware Healing on August 25 that shrank OpenAI's open GPT-OSS model from 120 billion parameters to 60 billion with 4-bit memory, …