04:13
2026-09-03
gist.github.com
machine-learning
README-omlx.md
A developer detailed their local AI setup using oMLX and Qwen models on an M5 Max Mac with 128GB unified memory, achieving 82-123 tok/s decode with a 35B-A3B MoE worker versus 15.5-17.8 tok/s for a 27…