{"slug": "deepseek-v4-1-flash-runs-23-seconds-token-on-a-2020-16gb-m1-mac-mini", "title": "DeepSeek v4.1 flash runs 23 seconds/token on a 2020 16gb M1 Mac Mini", "summary": "A developer posting as FP4 Brain on X reported running DeepSeek V4.1 Flash locally on a 16GB M1 Mac Mini using original FP4/FP8 weights, SSD streaming, and a custom MLX runner built on the pipenetwork MLX port, achieving 108 seconds time-to-first-token and about 23 seconds per token. The developer said the setup improved from roughly 31 to 23 seconds per token by reusing allocation buffers, compiling weight decoding, and caching 4 GiB of dense weights, keeping weights on the SSD with selected experts and lookup rows loaded on demand; a 6 GiB cache was tried but was barely faster with more swapping. Code, tests, and logs are published in the GitHub repository atbender/deeps.", "body_md": "FP4 Brain on X: \"got deepseek V4.1 flash running locally on a 16GB m1 mac mini\noriginal FP4/FP8 weights, ssd streaming + custom mlx runner\n108s ttft and about 23s/token (not to be confused with tok/s)\"\n\ngot deepseek V4.1 flash running locally on a 16GB m1 mac mini\noriginal FP4/FP8 weights, ssd streaming + custom mlx runner\n108s ttft and about 23s/token (not to be confused with tok/s)\n\ngot deepseek V4.1 flash running locally on a 16GB m1 mac mini\noriginal FP4/FP8 weights, ssd streaming + custom mlx runner\n108s ttft and about 23s/token (not to be confused with tok/s)\n\ncode + recipe if you're not in a hurry\ngithub.com/atbender/deeps…\nweights stay on SSD. the runner loads selected experts + lookup rows on demand, with a 4 GiB cache for dense weights. built on @pipenetwork MLX port\n\ngot it from about 31 to 23 seconds/token by reusing allocation buffers, compiling weight decoding and caching 4 GiB of dense weights\ntried 6 GiB too but it was meh, barely faster and more swapping. that's why I stuck with 4 if anyone's wondering\ntests + logs in the repo", "url": "https://wpnews.pro/news/deepseek-v4-1-flash-runs-23-seconds-token-on-a-2020-16gb-m1-mac-mini", "canonical_source": "https://twitter.com/thefp4brain/status/2098424202168586367", "published_at": "2026-09-12 02:46:31+00:00", "updated_at": "2026-09-12 02:57:05.288298+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "mlops", "developer-tools"], "entities": ["DeepSeek V4.1 Flash", "FP4 Brain", "M1 Mac Mini", "MLX", "pipenetwork", "GitHub", "atbender/deeps"], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-1-flash-runs-23-seconds-token-on-a-2020-16gb-m1-mac-mini", "markdown": "https://wpnews.pro/news/deepseek-v4-1-flash-runs-23-seconds-token-on-a-2020-16gb-m1-mac-mini.md", "text": "https://wpnews.pro/news/deepseek-v4-1-flash-runs-23-seconds-token-on-a-2020-16gb-m1-mac-mini.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-1-flash-runs-23-seconds-token-on-a-2020-16gb-m1-mac-mini.jsonld"}}