DS4 (DwarfStar 4) in 2026: The Redis Creator's Engine Runs DeepSeek V4 Flash on a 128GB Mac, DGX Spark, or Strix Halo — Which Box Earns It?
TL;DR: DS4 (DwarfStar 4), the MIT-licensed inference engine from Redis creator Salvatore Sanfilippo (antirez), runs DeepSeek V4 Flash — 284B parameters — in about 96.5GB via an asymmetric 2/8-bit quant, which means a single 128GB machine now hosts a quasi-frontier model. Measured decode: 39.4 tok/s on a MacBook Pro M5 Max, 32 tok/s on Strix Halo (with the ds4fa speculative fork), 13.8 tok/s on a DGX Spark. The Mac is the fastest single box under $6,000; the Spark is the slowest of the three but the only one that clusters to 90 tok/s.
| Mac Studio M5 Max 128GB | GMKtec EVO-X2 128GB (Strix Halo) | NVIDIA DGX Spark | |
|---|---|---|---|
| Best for | Fastest single-box decode | Cheapest 128GB seat | CUDA + multi-node scaling |
| Price (Oct 2026) | ~$5,099 | ~$3,499–$3,649 | $4,699 |
| V4 Flash decode | ~35–39 tok/s (M5 Max class) | 15.6 tok/s stock, 32 with ds4fa | 13.8 tok/s (90 on 4 nodes) |
| The catch | No CUDA, no cheap upgrade path | ROCm path needs a fork for full speed | Slowest single unit of the three |
Honest take: If you’re buying one machine specifically to live with DeepSeek V4 Flash as a daily coding model, buy the Mac Studio M5 Max 128GB — it decodes 2.5× faster than a single Spark and costs $400 more. Buy Strix Halo only if $3,500 is the hard ceiling, and buy Sparks only if you already know you’ll run two or more.
When we ran the DeepSeek V4 hardware math in July, the verdict was blunt: V4 Flash’s smallest usable quant was ~103GB of weights, no consumer box could hold it, and the API at $0.14 per million input tokens was the correct answer for almost everyone. That article is now out of date in the most interesting way possible, and the reason is a side project from the guy who wrote Redis.
What DS4 actually is
DS4 — “DwarfStar 4” — is a native inference engine written in C by Salvatore Sanfilippo (antirez), released May 7, 2026 under the MIT license and sitting at roughly 23,000 GitHub stars as of early October. It hit the Hacker News front page in May with his “A few words on DS4” post, and it’s back on the front page this month because the last few weeks brought substantial speed and correctness improvements across all three backends: Metal, CUDA (including DGX Spark), and ROCm (Strix Halo).
DS4 refuses to be a general-purpose runtime, and that refusal is the entire design. It targets DeepSeek V4 Flash first — with additional support for DeepSeek V4 PRO on high-memory machines and GLM 5.2 — and optimizes everything around that one model family:
- An extremely asymmetric 2/8-bit quant recipe. The routed MoE experts — the bulk of a 284B model’s mass — get aggressive 2-bit imatrix quantization. The parts that break models when you squeeze them (shared experts, attention projections, routing) stay at 8-bit. The result is a ~96.5GB file that holds up in practice, where a naive uniform 2-bit quant of the same model degrades visibly.
- A disk-backed KV cache. Sessions persist to disk keyed by a SHA-1 hash and reload instantly. Restart your machine mid-project and you skip re-prefilling tens of thousands of context tokens — for agentic coding workflows, that’s the difference between a 10-minute warmup and none.
- A native coding agent (
ds4-agent), with inference driven in-process rather than over a socket, plus integrations for Claude Code, opencode, and Pi. - DSpark distributed inference: pipeline parallelism that pools RAM across networked hosts, and tensor parallelism across two Macs via RDMA over Thunderbolt 5 or across multiple CUDA GPUs.
The model it serves is the same one we profiled in July: DeepSeek V4 Flash, a 284B-total / 13B-active MoE with a 1M-token context window, MIT-licensed weights. Nothing about the model changed. What changed is that someone built a quant recipe and a runtime that make 128GB — not 160GB+ — the entry ticket.
The measured numbers, machine by machine
These are reported benchmarks from the DS4 README, the official benchmark page, and independent third-party runs — not our own lab numbers, and we’ve flagged the source for each. All decode figures are for the ~96.5GB 2/8-bit quant of V4 Flash unless noted. We sanity-checked every row against the memory-bandwidth ceiling (decode speed can’t exceed bandwidth ÷ bytes read per token); all of them pass, clustering around an effective 15–17GB read per token — consistent with 13B active parameters plus the 8-bit shared layers.
| Machine | Memory bandwidth | Decode (tok/s) | Prefill (tok/s) | Source |
|---|---|---|---|---|
| MacBook Pro M5 Max 128GB | 614 GB/s | 39.4 | 790 | dwarfstar.sh benchmarks |
| MacBook Pro M4 Max 128GB | 546 GB/s | ~32 (holds at 1M ctx) | ~205 @ 64K | dwarfstar.sh benchmarks |
| MacBook Pro M3 Max 128GB | 400 GB/s | 26.7 | 250 (11.7K prompt) | dwarfstar.sh benchmarks |
| DGX Spark (GB10, 128GB) | 273 GB/s | 13.8 | 344 | DS4 README / classmethod |
| Strix Halo 128GB (Ryzen AI Max+ 395) | 256 GB/s | 15.6 stock / 32 via ds4fa | ~250 (ds4fa) | slb350 tracker / Framework Community |
| RTX PRO 6000 Blackwell 96GB | 1,800 GB/s | 43 (31 @ 50K ctx) | — | loftllc.dev first look |
Three things in that table deserve explanation.
The Macs win on decode, and it scales with bandwidth. 400 → 546 → 614 GB/s maps almost linearly onto 26.7 → 32 → 39.4 tok/s. Apple’s unified memory is doing exactly what it’s priced for here. The M4 Max number has a detail worth repeating: decode holds ~32 tok/s even at a 1M-token context, because DeepSeek V4’s compressed attention keeps the KV cache tiny (more below).
The DGX Spark is the slowest single box — again. 13.8 tok/s stock tracks the Spark’s 273 GB/s bandwidth, the same constraint we documented in our three-way 128GB comparison. An independent test by Classmethod measured the decay curve: 13.2 tok/s at 2K context, 11.4 at 65K, 9.95 at 131K, and 7.94 at 262K. The compressed KV cache is remarkably small the whole way — 52MB at 2K context, just 3.63GB at 262K. A Blackwell-tuned community fork (Entrpi/ds4-on-spark) claims ~3× upstream prefill and ~1.5× decode, which would put a single Spark near 20 tok/s, but treat fork numbers as provisional.
Strix Halo’s headline number needs an asterisk. Upstream DS4 does 15.6 tok/s on a Ryzen AI Max+ 395. The 32 tok/s figure comes from ds4fa, a ROCm-focused effort that combines a 102GB mixed quant (~2.88 bits/param) with an 11GB draft model doing speculative decoding — the big model verifies 4 positions per fused pass, which is how it beats the naive bandwidth ceiling per accepted token. It’s real (measured on gfx1151, ROCm 7.2.4, July 2026) but it’s a fork with its own quant, not a flag you flip in stock DS4.
And the one-card outlier: a single RTX PRO 6000 Blackwell holds the whole quant in its 96GB of VRAM and decodes at 43 tok/s. At $13,998–$19,999 street it’s not a value pick — it’s what the ceiling looks like when memory bandwidth stops being the constraint.
The problem you will actually hit on a 128GB Mac
The quant is ~96.5GB. A 128GB Mac sounds comfortable until macOS enforces its default GPU wired-memory limit — roughly 96GB on a 128GB machine — and the model fails to load or the loader gets killed. The fix is raising the ceiling:
sudo sysctl iogpu.wired_limit_mb=112640
# expected output:
# iogpu.wired_limit_mb: 0 -> 112640
That pins up to 110GB for the GPU, leaving the 96.5GB of weights plus KV cache and compute buffers inside the wired budget with ~18GB left for macOS. DS4 also reads an environment override (DS4_PREFILL_METAL_PHASES_WIRED_LIMIT_MIB=112640). Two warnings from people who’ve run this config: keep the limit strictly below total RAM, and don’t expect to run an IDE, a browser full of tabs, and a video call next to a wired 110GB — the OS has almost nothing left and can seize. The setting resets on reboot, which is the safe default.
There’s also a confirmed report (ds4 issue #46) of the stock quant running on a 96GB Mac — it fits with almost zero headroom, so treat 96GB as “technically works for benchmarking” and 128GB as the real minimum for daily use.
What this does to the buy-vs-API math
Be honest about the other side of the ledger before spending $3,500–$5,100: the API got cheaper while all of this was happening. DeepSeek launched V4.1 Flash on September 10 and cut Flash-series prices by up to 60%, moving to time-of-day pricing — off-peak output runs RMB 4 per million tokens (roughly $0.56), doubling at Chinese peak hours, with cache-hit input at RMB 0.02. The “significant price increase” DeepSeek warned about in August — the one we built trigger tables for — resolved into a price cut. If raw cost per token is your only criterion, the API still wins by a mile, and no hardware purchase on this page changes that.
What local V4 Flash buys you is the usual trio — privacy, offline operation, and immunity from rate limits and pricing whiplash — plus one thing that’s new: V4.1 Flash is a 552B MoE that needs a 4-node Spark cluster locally, and as of September 14 DeepSeek routes hosted V4 Pro requests to V4.1 Flash. The hosted target keeps moving. The 284B V4 Flash weights on your SSD don’t, and for agentic coding against a codebase you can’t legally upload, “a fixed model you control” is the feature.
Also note what DS4’s disk-backed KV cache does to the practical economics: the expensive part of long-context local work is prefill, and DS4 makes prefill a one-time cost per session instead of a per-restart cost. On a Spark, prefilling 131K tokens at ~165 tok/s incremental takes ~13 minutes — paying that once and resuming from disk afterward is what makes a 13.8 tok/s machine livable as an agent host.
Scaling past one box
A single Spark is the weakest option in the table, but it’s the only one with a sanctioned multi-node story today. On the NVIDIA developer forums, a 4-node Spark cluster running V4 Flash over DSpark measured 2,500 tok/s prefill and 90 tok/s decode — faster than the single RTX PRO 6000, for about $18,800 of hardware. The same forum crowd has V4.1 Flash (552B) running on 4× Spark at 72–77 tok/s on warm coding runs. On the Mac side, DSpark’s tensor parallelism pairs two Macs over Thunderbolt 5 RDMA; antirez has been explicit that the two-Mac path is a first-class target, and October’s updates improved it. We haven’t seen a stable two-Mac decode number worth printing yet — when one lands, it goes here.
What to actually buy
Prices as of October 2026, all verified against this site’s current price table:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| You want V4 Flash as a daily driver, fastest single box | Mac Studio M5 Max 128GB | ~$5,099 | Check price |
| Same, but you need it portable | MacBook Pro M5 Max 128GB | $6,699–$6,999 | Check price |
| $3,500 is the hard ceiling | GMKtec EVO-X2 128GB or Framework Desktop 128GB ($3,449, frame.work) | $3,449–$3,649 | Check price |
| CUDA required, or you’ll cluster later | ASUS Ascent GX10 (GB10, from $3,099) / DGX Spark ($4,699) | $3,099–$4,699 | Check price |
| Money is no object, one slot, 43 tok/s | RTX PRO 6000 Blackwell 96GB | $13,998+ | Check price |
| Not ready for $3,500 — run great 27–35B models instead | Used RTX 3090 24GB | $1,150–$1,350 | Check price |
| Undecided — test a local workflow before buying anything | Rented RTX 3090, from $0.07/hr | pay per hour | Vast.ai |
If you land in the last two rows, start with the best models for 24GB — a 24GB card running Qwen3.6-35B-A3B at 107 tok/s is a better daily experience than a 284B model at 14 tok/s for most tasks, and it costs a third as much. The 128GB machines earn their price only when the thing you need is specifically a frontier-adjacent model with your data staying home. And if you’re choosing between the Mac and the EVO-X2 on grounds beyond this one model, our Mac Studio M5 Max vs EVO-X2 comparison covers the full picture — thermals, ecosystem, and what each runs best across model sizes.
FAQ
Does DS4 replace Ollama or llama.cpp? No, and it doesn’t try to. DS4 runs DeepSeek V4 Flash, V4 PRO, and GLM 5.2 — that’s the catalog. For everything else (Qwen, Gemma, Llama, your GGUF collection), you still want llama.cpp or Ollama. DS4 is what you run when the whole point of the machine is one very large MoE done right.
Will it run on my 64GB machine? No. The smallest working setup is a ~96.5GB quant on a 96GB Mac with zero headroom (one confirmed report), and 128GB is the practical floor. A 64GB Strix Halo or 64GB Mac cannot host V4 Flash under DS4; see what 128GB unified memory actually unlocks before upgrading for this one model.
Is 2-bit quantization of the experts actually usable, or is this a benchmark toy? The asymmetric recipe is the point: only the routed experts are 2-bit, while shared experts, attention, and routing stay 8-bit. Community reception — 23k stars, people running it as their daily coding agent through Claude Code and opencode — suggests it holds up for real work, but there’s no formal perplexity study yet. If your workload is quality-critical, run your own evals before committing hardware money.
Why is Strix Halo faster than DGX Spark here when Spark costs more? Decode is bandwidth-bound and the two are nearly tied on paper (256 vs 273 GB/s). Stock-vs-stock they trade blows (15.6 vs 13.8 tok/s); Strix Halo’s 32 tok/s needs the ds4fa fork’s speculative decoding. The Spark’s case was never single-unit speed — it’s the ConnectX networking and the 4-node, 90 tok/s cluster path.
Can I use DS4 as a coding agent backend today?
Yes — that’s antirez’s own daily use. It ships ds4-agent and integrates with Claude Code, opencode, and Pi. The disk-backed KV cache matters more than raw tok/s for this: agent sessions resume without re-prefilling the repo context. For the editor-side setup of local backends generally, our sister site aicoderscope.com covers the tooling.
Sources
- A few words on DS4 — antirez
- antirez/ds4: DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm — GitHub
- ds4 Benchmarks: DeepSeek V4 Prefill and Token Speed — dwarfstar.sh
- Tried running DeepSeek V4 Flash 284B on DGX Spark with DwarfStar 4 — Classmethod
- 4-node DGX Spark Cluster with DeepSeek V4 Flash DSpark benchmark — NVIDIA Developer Forums
- DeepSeek V4 Flash: 284B model, up to 32 tok/s on AMD Ryzen AI MAX+ 395 — Framework Community
- Local LLM benchmarks on AMD Strix Halo — slb350
- DwarfStar 4 × RTX PRO 6000 Blackwell: DeepSeek V4 Flash Q2 Reaches 43 tok/s — Loft LLC
- DeepSeek formally launches V4.1 Flash, routes V4 Pro requests to Flash — TechNode
- DeepSeek-V4-Flash 0731: 284B Agent Fits a 128GB MacBook — modelfit.io
- iogpu.wired_limit_mb on Mac: Raising the Metal Memory Ceiling — ModelPiper
- DwarfStar 4 is a compact native inference engine designed specifically for DeepSeek V4 Flash — GIGAZINE
Last updated October 5, 2026. Prices and specs change; verify current rates before purchasing. Benchmark figures are reported by the sources above, not independently measured by us.
Recommended Gear
Products linked in this guide:
- Mac Studio M5 Max 128GB — fastest single-box V4 Flash decode under $6,000
- GMKtec EVO-X2 128GB — the budget 128GB seat
- ASUS Ascent GX10 — the cheaper GB10/CUDA entry
- RTX PRO 6000 Blackwell 96GB — single-card ceiling, 43 tok/s
- Used RTX 3090 24GB — the honest alternative if $3,500 is too much
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.