Running a 35B MoE model on a 2017 AMD RX 580 8GB via Vulkan (no ROCm/CUDA)
A developer successfully ran a 35-billion-parameter Mixture-of-Experts model on a 2017 AMD RX 580 8GB GPU using Vulkan, bypassing CUDA and ROCm. The project, called Polaris Revival, achieved 17-18 tokens per second for L…