Local Quantization and Multi-Backend Deployment with AMD Quark on Strix Halo
AMD detailed a local quantization workflow using AMD Quark 0.12.post1 that compressed the Qwen3.6-35B-A3B Mixture-of-Experts model from BF16 to W4A16 on a single Strix Halo system, cutting model weigh…