# Run DeepSeek V4 on Your Own Hardware With DwarfStar, the Redis Creator's New Inference Engine

> Source: <https://dev.to/jamilxt/run-deepseek-v4-on-your-own-hardware-with-dwarfstar-the-redis-creators-new-inference-engine-5hj>
> Published: 2026-10-04 03:04:41+00:00

The creator of Redis thinks your old GPU is not obsolete. Salvatore Sanfilippo, better known as antirez, published a project called DwarfStar (the repo is `antirez/ds4`) that runs DeepSeek V4 Flash, GLM 5.x, and DeepSeek V4 PRO entirely on consumer hardware: Macs, DGX Spark boxes, Strix Halo desktops, and older NVIDIA cards like Ada Lovelace and L40S. It hit the front page of Hacker News this week with hundreds of points, and the repo has already passed 21,000 stars.

The interesting part is not just the speed numbers. It is the design decision behind them. Instead of being a general GGUF runner like llama.cpp, DwarfStar is deliberately narrow: it supports a short list of models, ships its own quantized weights, and tests the whole stack together, from tensor loading to tool calls to the HTTP server. One model, made to feel finished.

Full disclosure before we go further: I have not run DwarfStar on my own machine. Everything below comes from the project's README and its official documentation on GitHub, all linked inline. This is a walkthrough of what it does, what hardware it actually needs, and how you would set it up, not a personal benchmark report. I will flag where the docs are honest about the limits.

A few facts from the README that frame everything else:

`ds4.c`), MIT licensed, with no dependency on GGML even though the docs openly credit llama.cpp's kernels and quantization work as the foundation.
That last point matters if you are comparing it to llama.cpp or Ollama. Those tools try to run everything. DwarfStar tries to run a few excellent models perfectly on machines people actually own. You cannot throw an arbitrary GGUF at it; the docs state plainly that other files will fail because the tensor layout, quantization mix, and metadata are specific to the weights the project publishes.

DeepSeek V4 Flash and the GLM Flash models are mixture-of-experts models. Most of their parameters are routed experts, and only a fraction activate per token. Two consequences make local inference practical:

This is why the same quantized model that would be hopeless in a dense architecture feels "quasi-frontier" here, in antirez's own words from the README.

This is the decision most people get wrong, so here is the matrix, straight from the documentation.

`ds4f-q2`, run `./ds4`. This is the setup the project treats as the baseline.` make cuda-spark`. This is the hardware antirez calls the main CUDA goal.` make cuda-generic`. This is the part that other backends often do not support. The docs report an eight-L40S setup running Flash Q2 at about 126 tokens per second aggregate across 16 concurrent sessions, essentially a small multi-user LLM server built from cards that are a GPU generation or two old.
The SSD streaming numbers are worth seeing because they change what "my laptop cannot fit that model" means. On a 128 GB M5 Max with automatic cache sizing, the docs' September measurements show GLM 5.3 Flash Q4 (a 177.77 GiB file) doing 121 tokens per second initial prefill and around 12 to 15 tokens per second generation, with a three-run median. DeepSeek Flash Vision MXFP4 (145.26 GiB) hit 300 tokens per second prefill and about 12 to 19 tokens per second generation. Those are reading speeds from a fast internal SSD, not a marketing claim, and the docs repeat that your workload will vary.

The happy path from the README takes four commands:

```
git clone https://github.com/antirez/ds4.git
cd ds4
make                      # Metal on Apple Silicon
./download_model.sh ds4f-q2
```

Then either the interactive CLI, the built-in agent, or the server:

```
./ds4 -p "Explain Redis streams in one paragraph."
./ds4-agent
./ds4-server --ctx 32768
```

A few details worth knowing before you start:

`./download_model.sh ds4f-q2` after an interrupted download and it continues with `curl -C -`. That matters when the file is 81 GiB.
`ds4-server` exposes an OpenAI-compatible endpoint on port 8000, and the included `ds4-agent` is a native coding agent that keeps token history and live model state together in local KV snapshots, so resumed sessions do not rebuild the whole prompt.

The docs include ready-made configs for wiring local DeepSeek into real agent tools:

`ANTHROPIC_BASE_URL` to your local server and points every model role, including subagents, at `deepseek-v4-flash`.` 127.0.0.1:8000`.
That last mile is what separates this from a benchmark toy. With cost fields literally set to zero in the Pi config, the pitch is simple: unlimited local tokens for your coding agent, with the privacy of a machine that never phones home. The tradeoff is capability. DeepSeek V4 Flash is strong, but it is not the same as frontier hosted models on the hardest tasks, and the docs' own evaluation section is careful to call `ds4-eval` "integration checks, not official leaderboard scores."

Two honest caveats from the same docs:

`--power`.
The README has a section called "AI full disclosure" where antirez states the software is developed with strong assistance from AI coding agents, with humans leading ideas, testing, and debugging, and says openly that if you are unhappy with AI-developed code, the project is not for you. He also argues that software should now ship as a working template for the biggest use cases, with users asking coding agents to adapt it to their specific hardware.

Whether or not you agree with that philosophy, it is a notable stance from someone who has written foundational infrastructure by hand for two decades, and the acknowledgements section draws the line clearly: this project leans on AI, while llama.cpp and GGML, which made it possible, were largely written by hand.

Here is my honest read of who should try it, as a checklist you can save:

The bigger story is what it says about the next year of local AI. Specialized engines per model family, quantization recipes built around MoE sparsity, and SSDs acting as an extension of RAM are all already here in one project. The gap between "consumer hardware" and "runs a frontier-class model" keeps shrinking, and this time the person shrinking it built Redis.

I write about AI tools, developer infrastructure, and practical engineering every week. Subscribe, it is free, and it helps me keep doing the hands-on breakdowns.

Have you tried DwarfStar or any local inference setup for coding agents? What hardware are you running, and what tokens per second are you seeing? I am collecting real-world numbers for a follow-up.
