{"slug": "dwarfstar-compresses-frontier-models-to-run-them-on-local-machines", "title": "DwarfStar compresses frontier models to run them on local machines", "summary": "Redis creator Salvatore Sanfilippo published ds4, an MIT-licensed C inference engine, on May 7th, 2026, compressing routed experts in mixture-of-experts models to roughly two-bit precision so supported builds fit in roughly 96GB to 128GB of memory. DwarfStar's engine supports only selected DeepSeek V4 and V4.1, GLM 5.x and Qwen3.8 Flash Next models on high-memory Macs, CUDA and ROCm systems, exposing the same model state through a CLI, an HTTP server with OpenAI- and Anthropic-compatible endpoints, and a native agent. Sanfilippo wrote that \"AI is too critical to be just a provided service,\" arguing coding agents can let users adapt open-source projects to configurations their creators cannot support directly.", "body_md": "# DwarfStar compresses frontier models to run them on local machines\n\n**DwarfStar's ds4 is an MIT-licensed C engine for selected models on high-memory Macs, CUDA and ROCm systems, built by Redis creator Salvatore Sanfilippo.**\n\n        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)\n        · Published \n\nPrimary source: [DwarfStar](https://dwarfstar.sh/)\n\n## Why it matters\n\nDwarfStar shows how model-specific engineering can run selected open-weight models locally, while its memory requirements limit who can use it. Sanfilippo argues that coding agents could help users adapt open-source projects to configurations their creators cannot support directly.\n\n[Salvatore Sanfilippo](https://antirez.com/news/165?ref=runtimewire), who created Redis, built [DwarfStar](https://dwarfstar.sh/?ref=runtimewire) around a constraint most local AI software tries to avoid: support fewer models, then tune the whole stack for them. The resulting open-source engine, ds4, compresses selected large models to run on high-memory Macs and GPU systems, with a command-line interface, local APIs and a coding agent.\n\nSanfilippo published the [ds4 repository](https://github.com/antirez/ds4?ref=runtimewire) on May 7th, 2026, according to independent coverage by [RunLocal](https://runlocal.blog/blog/dwarfstar-ds4-one-model-inference-engine?ref=runtimewire). The DwarfStar site provides current documentation and a benchmark hub.\n\nSanfilippo's thesis is direct: \"AI is too critical to be just a provided service,\" [he wrote about ds4](https://antirez.com/news/165?ref=runtimewire). The engine is designed to run open-weight models on hardware users own, keeping model execution and code on the local machine. The approach requires hardware well above what most people already have.\n\n### Specialization is the product\n\nDwarfStar supports a limited set of models rather than serving as a general-purpose GGUF runner. Its current documentation lists selected DeepSeek V4 and V4.1 models, [GLM 5](https://runtimewire.com/models/z-ai/glm-5).x and [Qwen3.8 Flash Next](https://runtimewire.com/models/huggingface/qwen-qwen3.8-flash-next-c6e4ec527c4bd931), with supported model layouts and capabilities varying by backend. The project says it tests the model, prompt handling, tool calls, cache and interfaces together. Users cannot assume any arbitrary model file will work.\n\nThat narrow support surface lets Sanfilippo make model-specific choices. ds4 compresses the routed experts in mixture-of-experts models to roughly two-bit precision while preserving higher precision for shared components and other paths the project identifies as critical. In a large model, routed experts account for much of the parameter count, so lowering their precision can cut memory use substantially. DwarfStar's [documentation](https://dwarfstar.sh/docs/architecture/?ref=runtimewire) says this approach lets supported builds fit in roughly 96GB to 128GB of memory.\n\nThe engine also treats a model's key-value cache as persistent data. It can save prompt-prefix state to SSD and restore a matching prefix after a restart, avoiding the need to recompute that work from scratch. The same model state is exposed through a CLI, an HTTP server with OpenAI- and Anthropic-compatible endpoints, and a native agent. These tools let users keep long coding sessions local without rebuilding the workflow around a DwarfStar-only client.\n\nSanfilippo has described a second bet alongside local inference: that coding agents change what open-source software needs to be. In a recent [essay on software distribution](https://antirez.com/news/170?ref=runtimewire), he argued that a repository can serve as a template for users and their agents to adapt to different hardware and needs, rather than as a finished product that covers every configuration. He points to DwarfStar itself as an example: a strong implementation for a few models and backends can give coding agents a pattern to extend.\n\nThe project makes that trade-off explicit. A specialist engine may be easier to tune and validate for its chosen targets; users who need other models or devices may need a broader runtime or to adapt code themselves. Sanfilippo's argument is that coding agents lower the cost of that adaptation. It is a bet on how developers will work with open-source software, not a promise that ds4 already supports every machine or model.\n\n### The hardware bill remains real\n\nDwarfStar's own [benchmark table](https://dwarfstar.sh/benchmarks/?ref=runtimewire) reports 39.4 tokens per second of generation on an M5 Max with 128GB of memory at a 2,048-token context, and 27.6 tokens per second at 65,536 tokens. For a 128GB DGX Spark, it reports 18.1 and 13.8 tokens per second at those contexts. These are project-published measurements, not independent or apples-to-apples comparisons against other inference engines. The table also separates generation from prompt prefill, which measures how quickly the system ingests input; the two rates describe different parts of a workload.\n\nThe memory requirement limits who can use it. DwarfStar's website lists Apple Silicon machines with at least 64GB for some supported configurations, while its baseline [DeepSeek V4 Flash](https://runtimewire.com/models/azure/deepseek-v4-flash) Q2 setup is aimed at higher-memory systems. SSD streaming can extend what fits, but it does not make the capacity trade-off disappear. Running a large model locally remains a task for expensive hardware, even when the engine reduces the amount of memory it needs.\n\nSanfilippo began building [Redis](https://redis.io/blog/welcome-back-to-redis-antirez/?ref=runtimewire) in 2009, and the focused open-source database project later brought him back in 2024 as an evangelist after he had stepped away from day-to-day maintenance. DwarfStar applies a similar preference for a tightly defined system, this time to inference and the tools around it. The [MIT-licensed code](https://github.com/antirez/ds4?ref=runtimewire) is available for developers to inspect and adapt, while the project's own benchmarks and validation remain the primary evidence for its performance claims.\n\nFor now, ds4 is a local AI stack designed around a small set of large models: weights, inference engine, persistent cache, API and agent in one project. Its usefulness will depend on whether the codebase can remain useful as open models and hardware change, and whether agent-assisted adaptation can help a deliberately narrow tool serve more people without turning it into the generic runtime it set out not to be.", "url": "https://wpnews.pro/news/dwarfstar-compresses-frontier-models-to-run-them-on-local-machines", "canonical_source": "https://runtimewire.com/article/dwarfstar-local-inference-salvatore-sanfilippo", "published_at": "2026-10-02 21:03:49+00:00", "updated_at": "2026-10-02 21:08:58.304904+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "ai-agents", "ai-infrastructure"], "entities": ["DwarfStar", "ds4", "Salvatore Sanfilippo", "Redis", "DeepSeek V4", "GLM 5.x", "Qwen3.8 Flash Next", "RunLocal"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/dwarfstar-compresses-frontier-models-to-run-them-on-local-machines", "markdown": "https://wpnews.pro/news/dwarfstar-compresses-frontier-models-to-run-them-on-local-machines.md", "text": "https://wpnews.pro/news/dwarfstar-compresses-frontier-models-to-run-them-on-local-machines.txt", "jsonld": "https://wpnews.pro/news/dwarfstar-compresses-frontier-models-to-run-them-on-local-machines.jsonld"}}