09:10
2026-08-05
sourcefeed.dev
artificial-intelligence
How a 20B Model Hits 120 tok/s on an iPhone
DeepGrove's Maple-Preview, a 20B-parameter mixture-of-experts reasoning model with ternary weights, achieves 120 tokens per second on an iPhone and 218 tok/s on a base M4 Mac mini, according to the coโฆ