DGX Spark 64GB vs 128GB vs Strix Halo in 2026: NVIDIA's New $4,999 Box Can't Run the Model It's Famous For

dgx-sparknvidiagb10strix-halomini-pclocal-llmhardware

TL;DR: NVIDIA announced a 64GB DGX Spark on October 2 at $4,999 (on sale October 23) and raised the 128GB Founders Edition to $6,950 — 74% above its original $3,999 list. The 64GB model cannot hold gpt-oss 120B, the model nearly every Spark benchmark is built on. For pure local inference, a 128GB Strix Halo box at $2,959–$3,649 remains the value pick.

DGX Spark 64GBDGX Spark 128GBStrix Halo 128GB
Best forCUDA dev work on ≤50GB modelsCUDA dev + 120B-class MoEPure inference per dollar
Price$4,999 (Oct 23, partner SKUs)$6,950 FE$2,959–$3,649
Memory / bandwidth64GB @ 273 GB/s128GB @ 273 GB/s128GB @ 256 GB/s
The catchgpt-oss 120B (~65GB) doesn’t fitSecond price hike in a year~5× slower prefill, no CUDA

Honest take: Buy the 64GB Spark only if you need the CUDA/DGX software stack and your models stay under ~50GB. If you just want to run big MoE models at home, a Minisforum MS-S1 Max gives you double the memory for $2,000 less — and if you genuinely need a Spark, the $6,950 128GB is the only configuration that runs what the platform is marketed on.

What NVIDIA announced on October 2

NVIDIA’s DGX Spark line got its first down-market variant: a 64GB unified-memory configuration at $4,999, shipping October 23 through Acer, ASUS, Dell, Gigabyte, HP, and MSI rather than as a Founders Edition. Everything else about the GB10 platform carries over unchanged — the 20-core Arm CPU, the Blackwell GPU with up to 1 petaFLOP of sparse FP4 compute, ConnectX-7 networking, DGX OS, and the same 273 GB/s of LPDDR5X bandwidth. NVIDIA pitches it for models up to roughly 100B parameters, and explicitly positions two clustered units as the path back to 128GB.

One spec did not carry over: storage. The 128GB Founders Edition ships a 4TB self-encrypting NVMe drive; for the 64GB model NVIDIA published no standard capacity, and partners will differentiate on the SSD. Treat $4,999 as the floor, not the price.

The bigger news is what happened to the existing model. The 128GB Founders Edition now lists at $6,950:

DateDGX Spark 128GB list priceChange
Launch (2025)$3,999—
February 23, 2026$4,699+18%
October 2, 2026$6,950+48% (+74% cumulative)

NVIDIA attributes both hikes to memory supply — the same DRAM squeeze that has DDR5 kits selling at 4× their 2025 prices (a 64GB DDR5 kit now runs $680–$1,070 per Tom’s Hardware’s RAM index). The 64GB Spark isn’t a cheaper Spark so much as a hedge: NVIDIA halving the most expensive component to keep an entry price under $5K.

The 65GB model vs the 64GB box

Here’s the problem with that halving. The model that made the Spark’s reputation — gpt-oss 120B, the MoE that NVIDIA demos lean on and that every independent benchmark thread runs first — weighs about 65GB in its native MXFP4 quantization, and ~68.5GB resident once llama.cpp allocates full context. On a 64GB machine that still needs several GB for the OS and the rest of the stack, it does not fit. Not “fits with tuning” — it doesn’t load. NVIDIA’s own “up to 100B parameters” guidance quietly concedes this.

What does fit in a ~55GB usable budget:

  • Dense 70B at Q4_K_M (42.5GB) — fits, and is unusable. Decode on the 128GB Spark measures ~2.6 tok/s on dense 70B because all 42.5GB of weights cross the 273 GB/s bus for every token. The 64GB model has identical bandwidth, so it inherits the identical crawl. Human reading speed is 7–10 tok/s.
  • gpt-oss 20B (~13GB) — runs great: ~60 tok/s decode, ~2,000 tok/s prefill. It also runs great on a used RTX 3090 that costs $1,150–$1,350.
  • Qwen3-Coder 30B at Q8 (~32GB) — ~44 tok/s on the GB10, a genuinely good fit.
  • 30–50GB MoE models generally — this is the 64GB Spark’s real lane: models too big for a 24GB GPU, small enough to leave headroom, with few active parameters per token.

That lane is real, but notice what’s missing from it: everything between roughly 55GB and 128GB — gpt-oss 120B, Qwen3.8-Flash-Next 125B at 4-bit, 70B at Q8 — the exact capacity class that justified buying a unified-memory box instead of a graphics card in the first place. Run your target model through our VRAM calculator before you order; the 64GB tier punishes guessing.

Paying $1,951 less buys the exact same speed

The one honest virtue of the 64GB SKU: it is precisely as fast as the $6,950 model on anything that fits. Decode speed in this class is memory-bandwidth-bound, and both Sparks move 273 GB/s. The ceiling is simple arithmetic — bandwidth divided by bytes read per token — and measured results land at 70–95% of it:

ModelDGX Spark (either config)Ryzen AI Max+ 395 (Strix Halo)
gpt-oss 20B MXFP4~60 tok/s30–33 tok/s
Qwen3-Coder 30B Q8~44 tok/s~40 tok/s class
gpt-oss 120B MXFP438.5 tok/s (doesn’t fit 64GB)31–56 tok/s depending on stack
Llama 3.3 70B dense Q4~2.6 tok/s4–6 tok/s
gpt-oss 120B prefill~1,723 tok/s~340 tok/s

Two things jump out. First, Strix Halo’s 256 GB/s is within 7% of the Spark’s bandwidth, so on decode — the part you spend your time watching — the $2,959 box and the $4,999 box are effectively tied. Second, the Spark’s genuine advantage is prefill: roughly 5× faster prompt processing, because prefill is compute-bound and Blackwell’s tensor cores plus mature CUDA kernels crush an RDNA 3.5 iGPU there. If your workload is long-context RAG or agentic loops that re-read 50K-token prompts, that 5× is real money. If you’re chatting with a model, it’s invisible. We walked through this dynamic in detail in Ryzen AI Halo vs DGX Spark back when both cost $3,999 — the performance picture hasn’t changed, only the price gap has tripled.

The clustering pitch costs more than not needing it

NVIDIA’s answer to “64GB isn’t enough” is ConnectX-7: link two 64GB Sparks into a 128GB pool. The arithmetic is unkind — two 64GB units run $9,998 against $6,950 for one 128GB machine that needs no network hop, no tensor-parallel configuration, and no second power brick. Clustering only makes sense if you’re buying the second unit later, or you want two independent dev boxes most of the time.

If you do cluster, budget an afternoon for the known ConnectX-7 gotcha on GB10 systems. Fresh DGX OS images have shipped with a driver that throttles the fabric:

$ ibstat | grep Rate        # expected: Rate: 200
Rate: 13

A link negotiating at 13 Gbps instead of ~200 Gbps makes cluster inference slower than one box. The fix that worked on the ASUS GB10 units we covered in the Ascent GX10 vs DGX Spark teardown: sudo apt full-upgrade to move the driver stack from 580.126.09 to 580.142, which restored measured throughput from 13 to 196 Gbps. NVIDIA has said a DGX OS update later in October targets cluster setup specifically.

What $4,999 buys elsewhere

The uncomfortable table, prices as of October 2026 (street, not MSRP — see each machine’s linked coverage for verification):

MachineMemoryPrice$/GB
Minisforum MS-S1 Max128GB @ 256 GB/s~$2,959 (2TB)$23
Framework Desktop128GB @ 256 GB/s$3,449$27
GMKtec EVO-X2128GB @ 256 GB/s$3,499–$3,649$28
GMKtec EVO-X2 64GB64GB @ 256 GB/s$2,199$34
DGX Spark 128GB128GB @ 273 GB/s$6,950$54
DGX Spark 64GB64GB @ 273 GB/s$4,999$78

The 64GB Spark is the most expensive unified memory per gigabyte you can buy in this class — more than double the EVO-X2 64GB, which matches its capacity and decodes within 7% of its speed. What the premium buys is entirely software and prefill: first-class CUDA, cuDNN and the PyTorch training ecosystem (QLoRA on 70B has been measured north of 5,000 tok/s throughput on the GB10), NVFP4 kernels, the DGX OS container stack, and that 5× prompt-processing lead. For a developer whose employer reimburses hardware and whose code targets CUDA in production, that’s a defensible $4,999. For a home lab that wants to run big models tonight, it isn’t — the $3K/$6K/$10K build guide covers what the same money assembles from parts, and the 64GB VRAM model guide shows exactly which models the 64GB tier serves well.

One more alternative deserves naming: if what you actually want is CUDA for occasional fine-tuning runs, renting covers it. A 3090 starts around $0.07/hr on Vast.ai and even 5090s start near $0.25/hr — at $0.25/hr, the $4,999 the Spark costs buys roughly 20,000 GPU-hours before the purchase breaks even on rental.

What to actually buy

Prices as of October 2026, all taken from the comparison above:

Your situationThe machinePriceWhere
You want 120B-class MoE at home, cheapest routeMinisforum MS-S1 Max (128GB/2TB)~$2,959Check price
Same, but you want current firmware support and 3× M.2GMKtec EVO-X2 (128GB)$3,499–$3,649Check price
You need CUDA + DGX stack and run 120B-class modelsDGX Spark 128GB$6,950Check price
You need CUDA and your models stay under ~50GBDGX Spark 64GB (from Oct 23)$4,999+Partner stores (Acer/ASUS/Dell/Gigabyte/HP/MSI)
Undecided — test your workload before buying anythingRented GPU, 3090 from $0.07/hrpay per hourVast.ai

For coding-assistant workloads specifically — where prefill speed matters because the agent re-reads your repo constantly — pair whichever box you pick with the local-first setups in Cline and the local-backend coding tools, and if you’re standing up the serving layer, Ollama’s 2026 state of play covers which runtime to put on it.

FAQ

Can the DGX Spark 64GB run gpt-oss 120B? No. The model is ~65GB in native MXFP4 and ~68.5GB resident with full context in llama.cpp — more than the machine’s total memory. You’d need two clustered units ($9,998) or the 128GB model ($6,950). On the 64GB box, the biggest comfortable targets are 30–50GB MoE models.

Is the 64GB Spark slower than the 128GB one? No. Same GB10 chip, same 273 GB/s bandwidth, same ~1 petaFLOP FP4 — identical speed on anything that fits in 64GB. You’re paying $1,951 more for memory capacity, not speed.

Why did the 128GB DGX Spark jump to $6,950? NVIDIA cites memory supply costs — LPDDR5X competes for the same fab capacity as the HBM going into datacenter accelerators. It’s the second hike: $3,999 at launch, $4,699 in February 2026, $6,950 now.

Is Strix Halo really as fast as DGX Spark for inference? On token generation, effectively yes: 256 vs 273 GB/s bandwidth means decode lands within ~7% (34 vs 38.5 tok/s measured on gpt-oss 120B). The Spark wins prompt processing by ~5× and everything CUDA-shaped (fine-tuning, video gen, NVFP4). Pure chat/inference buyers are paying the NVIDIA premium for a phase of the workload they barely notice.

Should I wait for the October 23 partner SKUs before buying anything? Only if you specifically want the 64GB Spark — partner pricing and SSD configs aren’t final until listings go live. The Strix Halo boxes and the 128GB Spark are shipping today, and nothing about the 64GB launch changes their math.

Sources

Prices as of October 2026. Hardware prices move weekly — verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.