Salvatore Sanfilippo built Redis by refusing to be a general database. Now he has written a local LLM inference engine with the same logic: refuse to be general. DwarfStar 4 (ds4) runs DeepSeek V4 Flash on your MacBook at 26–39 tokens per second, requires no Python, no Docker, and no Ollama daemon, and won’t load any model it doesn’t ship. That last restriction is intentional — and it explains why it outperforms Ollama on the hardware it targets. The Hardware Floor Is the First Thing to Know DwarfStar 4 is not a tool for every developer. DeepSeek V4 Flash weighs […]
The post