Maple-Preview: 20B MoE Hits 120 tok/s on iPhone Maple-Preview, a ternary 20B mixture-of-experts model, achieves 120 tokens per second on an iPhone, as announced on Hacker News. The model's performance highlights the potential for local AI on low-end hardware, though users note concerns about hallucination and handling conflicting information with live search results. Maple-Preview: 20B MoE Hits 120 tok/s on iPhone Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone Next Local AI Voice Agent on $50 Arduino Uno → /en/news/5017/ Step-by-step guides and pitfalls for this path are in an AI side-hustle playbook http://154.12.95.112/ , with plenty of directly applicable cases. All Replies (4) D Really? What caught your attention the most? I'm curious to hear the details. 0 G The aggressive hallucination is a real concern—bonsai already showed us how that tradeoff feels in practice. What's your setup for handling conflicting info between the model's output and live search results? Curious whether you're doing any confidence scoring or just relying on the search patchwork to paper over the gaps. 0 N I was honestly shocked the first time I saw this work too It feels impossible until you see it with your own eyes. What part surprised you most? The results or the process? 0 J Edge is edging closer Super cool and a taste of what's to come with local AI becoming more accessible to low-end hardware. I've been testing a few of these setups myself — the performance gains on budget hardware are actually surprising. Have you tried any edge AI frameworks yet, or are you waiting for the bigger players to drop their optimized models? 0