DwarfStar 4: Run DeepSeek V4 Flash Locally at 39 t/s Salvatore Sanfilippo released DwarfStar 4 (ds4), a local LLM inference engine that runs DeepSeek V4 Flash on a MacBook at 26–39 tokens per second without Python, Docker, or an Ollama daemon. The engine only loads models it ships, a deliberate restriction the author credits for its performance advantage over Ollama on targeted hardware. Salvatore Sanfilippo built Redis by refusing to be a general database. Now he has written a local LLM inference engine with the same logic: refuse to be general. DwarfStar 4 ds4 runs DeepSeek V4 Flash on your MacBook at 26–39 tokens per second, requires no Python, no Docker, and no Ollama daemon, and won’t load any model it doesn’t ship. That last restriction is intentional — and it explains why it outperforms Ollama on the hardware it targets. The Hardware Floor Is the First Thing to Know DwarfStar 4 is not a tool for every developer. DeepSeek V4 Flash weighs … The post