DwarfStar's ds4 is an MIT-licensed C engine for selected models on high-memory Macs, CUDA and ROCm systems, built by Redis creator Salvatore Sanfilippo.
By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
· Published
Primary source: [DwarfStar](https://dwarfstar.sh/)
Why it matters #
DwarfStar shows how model-specific engineering can run selected open-weight models locally, while its memory requirements limit who can use it. Sanfilippo argues that coding agents could help users adapt open-source projects to configurations their creators cannot support directly.
Salvatore Sanfilippo, who created Redis, built DwarfStar around a constraint most local AI software tries to avoid: support fewer models, then tune the whole stack for them. The resulting open-source engine, ds4, compresses selected large models to run on high-memory Macs and GPU systems, with a command-line interface, local APIs and a coding agent.
Sanfilippo published the ds4 repository on May 7th, 2026, according to independent coverage by RunLocal. The DwarfStar site provides current documentation and a benchmark hub.
Sanfilippo's thesis is direct: "AI is too critical to be just a provided service," he wrote about ds4. The engine is designed to run open-weight models on hardware users own, keeping model execution and code on the local machine. The approach requires hardware well above what most people already have.
Specialization is the product
DwarfStar supports a limited set of models rather than serving as a general-purpose GGUF runner. Its current documentation lists selected DeepSeek V4 and V4.1 models, GLM 5.x and Qwen3.8 Flash Next, with supported model layouts and capabilities varying by backend. The project says it tests the model, prompt handling, tool calls, cache and interfaces together. Users cannot assume any arbitrary model file will work.
That narrow support surface lets Sanfilippo make model-specific choices. ds4 compresses the routed experts in mixture-of-experts models to roughly two-bit precision while preserving higher precision for shared components and other paths the project identifies as critical. In a large model, routed experts account for much of the parameter count, so lowering their precision can cut memory use substantially. DwarfStar's documentation says this approach lets supported builds fit in roughly 96GB to 128GB of memory.
The engine also treats a model's key-value cache as persistent data. It can save prompt-prefix state to SSD and restore a matching prefix after a restart, avoiding the need to recompute that work from scratch. The same model state is exposed through a CLI, an HTTP server with OpenAI- and Anthropic-compatible endpoints, and a native agent. These tools let users keep long coding sessions local without rebuilding the workflow around a DwarfStar-only client.
Sanfilippo has described a second bet alongside local inference: that coding agents change what open-source software needs to be. In a recent essay on software distribution, he argued that a repository can serve as a template for users and their agents to adapt to different hardware and needs, rather than as a finished product that covers every configuration. He points to DwarfStar itself as an example: a strong implementation for a few models and backends can give coding agents a pattern to extend.
The project makes that trade-off explicit. A specialist engine may be easier to tune and validate for its chosen targets; users who need other models or devices may need a broader runtime or to adapt code themselves. Sanfilippo's argument is that coding agents lower the cost of that adaptation. It is a bet on how developers will work with open-source software, not a promise that ds4 already supports every machine or model.
The hardware bill remains real
DwarfStar's own benchmark table reports 39.4 tokens per second of generation on an M5 Max with 128GB of memory at a 2,048-token context, and 27.6 tokens per second at 65,536 tokens. For a 128GB DGX Spark, it reports 18.1 and 13.8 tokens per second at those contexts. These are project-published measurements, not independent or apples-to-apples comparisons against other inference engines. The table also separates generation from prompt prefill, which measures how quickly the system ingests input; the two rates describe different parts of a workload.
The memory requirement limits who can use it. DwarfStar's website lists Apple Silicon machines with at least 64GB for some supported configurations, while its baseline DeepSeek V4 Flash Q2 setup is aimed at higher-memory systems. SSD streaming can extend what fits, but it does not make the capacity trade-off disappear. Running a large model locally remains a task for expensive hardware, even when the engine reduces the amount of memory it needs.
Sanfilippo began building Redis in 2009, and the focused open-source database project later brought him back in 2024 as an evangelist after he had stepped away from day-to-day maintenance. DwarfStar applies a similar preference for a tightly defined system, this time to inference and the tools around it. The MIT-licensed code is available for developers to inspect and adapt, while the project's own benchmarks and validation remain the primary evidence for its performance claims.
For now, ds4 is a local AI stack designed around a small set of large models: weights, inference engine, persistent cache, API and agent in one project. Its usefulness will depend on whether the codebase can remain useful as open models and hardware change, and whether agent-assisted adaptation can help a deliberately narrow tool serve more people without turning it into the generic runtime it set out not to be.