RTX Spark Pricing Is Real: Every Surface Laptop Ultra Config From $2,599 to $5,899, and the Only One That Makes Sense for Local AI
TL;DR: RTX Spark pricing is no longer a rumor — Surface Laptop Ultra preorders opened October 7 at $2,599 to $5,899 across eight configs, shipping October 16. Only the $5,899 128GB config matters for local AI, and it costs $2,250 more than a GMKtec EVO-X2 with the same capacity. Buy it for CUDA-in-a-backpack or not at all.
| Surface Laptop Ultra 128GB | GMKtec EVO-X2 128GB | Mac Studio M5 Max 128GB | |
|---|---|---|---|
| Best for | 128GB + CUDA in a laptop | Cheapest 128GB that runs 120B models | Fastest 128GB machine under $5,500 |
| Memory / bandwidth | 128GB / ~300 GB/s | 128GB / 256 GB/s | 128GB / 614 GB/s |
| Price | $5,899 (preorder, ships Oct 16) | $3,499–$3,649 | $5,099 |
| The catch | v1 drivers, Windows takes a memory cut | No CUDA, slow prompt processing | No CUDA, it’s a desktop |
Honest take: If the question is “cheapest way to run gpt-oss-120b at home,” the Surface loses to three machines that already shipped. The $5,899 config exists for exactly one buyer: someone who needs 100GB+ of model weights and the CUDA stack and a battery. Everyone else should read the price ladder below, notice that five of the eight configs can’t hold more model than a $1,300 used RTX 3090, and close the preorder tab.
Yesterday we told you not to preorder the Surface Laptop Ultra because Microsoft hadn’t named a price and the only public benchmark came from a leaked prototype whose CUDA path didn’t work. Half of that changed within hours: preorders went live October 7 at $2,599 to $5,899, with shipping starting October 16 (TechPowerUp, VideoCardz). The driver question is still open — nobody outside NVIDIA has published a retail-hardware llama.cpp run yet — but the money question now has exact answers, and they change the buying math enough to deserve their own breakdown.
This is that breakdown: what each of the eight configs costs, what each one can actually hold, why the memory bandwidth is identical from $2,599 to $5,899, and where the one interesting config lands against the 128GB machines you can already buy.
The full price ladder
Microsoft listed eight configurations at launch (TweakTown, cross-checked against Windows Central). Two N1X chip tiers exist — an 18-core CPU with a 5,120-CUDA-core GPU, and a 20-core CPU with 6,144 CUDA cores (HWBusters) — and the memory ceiling is what you’re really paying for:
| Config | Memory | SSD | Price | $/GB of memory | Local AI verdict |
|---|---|---|---|---|---|
| 18-core / 5,120 CUDA | 24GB | 512GB | $2,599 | $108 | Pass — a used RTX 3090 holds the same models at 3× the bandwidth |
| 18-core / 5,120 CUDA | 24GB | 1TB | $2,999 | $125 | Pass — same ceiling, $400 for storage |
| 18-core / 5,120 CUDA | 32GB | 1TB | $3,299 | $103 | Pass — 27B-class ceiling |
| 20-core / 6,144 CUDA | 32GB | 1TB | $3,699 | $116 | Pass — faster prefill, same 32GB wall |
| 20-core / 6,144 CUDA | 48GB | 1TB | $3,999 | $83 | Weak — 70B Q4 fits but decodes at ~5-7 tok/s |
| 20-core / 6,144 CUDA | 64GB | 1TB | $4,299 | $67 | Marginal — gpt-oss-120b doesn’t fit after Windows takes its cut |
| 20-core / 6,144 CUDA | 64GB | 2TB | $4,699 | $73 | Marginal — the 2TB is the only path to big storage |
| 20-core / 6,144 CUDA | 128GB | 1TB | $5,899 | $46 | The only local-AI config |
Three things jump out of that table.
The 64GB-to-128GB step costs $1,600. That sounds like Apple-grade memory pricing until you remember what DDR5 costs in late 2026: a 64GB desktop kit runs $680–$1,070 retail right now, and LPDDR5X soldered on a package is the more expensive kind. The upcharge is steep, not insane — but it does mean the 128GB config carries the best per-gigabyte price on the whole ladder, $46/GB versus $108/GB at the bottom. Microsoft priced the ladder to pull AI buyers upward, and for once the top rung is the rational one.
The 128GB config is 1TB-only. If you keep a few quantized 70B–120B models plus a ComfyUI model folder, 1TB fills fast, and there’s no 128GB/2TB option at any price. Budget for an external drive.
And pricing came in under the leaks. Earlier retail chatter pegged N1X systems “above $2,899” at entry and floated $7,000+ for the top config (VideoCardz). $2,599 at the bottom and $5,899 at the top undercut both — against the MacBook Pro 16” M5 Max at $6,999 for 128GB, the Surface is $1,100 cheaper for the same capacity (at less than half the bandwidth; more on that below).
The Surface isn’t alone anymore, either. MSI’s Prestige N16 Flip AI+ is listed for preorder at $3,299 with the full 20-core chip, 32GB, and 1TB, with a November 6 release (per Newegg’s own preorder listing — citation in Sources) — $400 less than Microsoft charges for the same silicon and memory. It doesn’t change the local-AI verdict (32GB is 32GB), but it signals that OEM competition on the same chip will compress prices fast. Microsoft also announced a Surface RTX Spark Dev Box — a mini PC with a fixed 128GB and a 100W thermal budget, bundling WSL2 GPU passthrough and full CUDA out of the box (Engadget, Tom’s Hardware). One outlet reports Dev Box preorders at $5,999 shipping in November (Pulse2), but Microsoft’s own page showed no price at announcement — treat that figure as unconfirmed until it’s on microsoft.com.
Every config has the same bandwidth, so the ladder only buys capacity
All eight configs use LPDDR5X at roughly 300 GB/s — the $2,599 machine and the $5,899 machine move weights at the same speed. (The DGX Spark desktop, built on the same GB10 silicon family, is rated at 273 GB/s.) Decode speed on local LLMs is bandwidth-bound, so the ceiling math is fixed across the whole ladder:
- A dense 70B at Q4_K_M reads ~42.5GB per token: 300 ÷ 42.5 ≈ 7 tok/s ceiling, before real-world losses. The 48GB config can hold that model; nothing on this platform runs it pleasantly.
- gpt-oss-120b is a MoE that reads ~3GB of active experts per token: 300 ÷ 3 ≈ 100 tok/s ceiling. This is the model class the 128GB config exists for.
How close does GB10-family hardware get to that MoE ceiling in practice? The best public data is the llama.cpp community benchmark thread for the DGX Spark, the Surface’s 273 GB/s desktop cousin (llama.cpp discussion #16578). On the stock DGX OS kernel, build 6816:
$ llama-bench -m gpt-oss-120b-mxfp4.gguf -fa 1 -ub 2048
| model | size | test | t/s |
| ----------------------- | --------- | ------ | --------------- |
| gpt-oss 120B MXFP4 MoE | 59.02 GiB | pp2048 | 1737.17 ± 81.66 |
| gpt-oss 120B MXFP4 MoE | 59.02 GiB | tg32 | 45.87 ± 0.74 |
That’s 46 tok/s decode on a 120B-class model — about half the theoretical ceiling, which tells you the software stack still has headroom. The same thread proves the point: one user swapped the stock kernel for a mainline 6.17 build on Fedora 43 and the same benchmark jumped to 60.57 tok/s decode and 1,956 tok/s prefill — a 32% speedup from a kernel change alone. Expect the Surface’s Windows numbers to start below that and climb with driver updates, the same trajectory. For reference, the only leaked Surface prototype ran Qwen3.5 9B at 22.45 tok/s on Vulkan because CUDA wasn’t working at all — we covered that leak in detail yesterday. Nothing published since changes that caution: preordering still means betting the retail drivers land in better shape than the September prototype.
Windows takes a cut: 128GB isn’t 128GB
There’s a subtlety buried in NVIDIA’s own developer documentation that matters more on this machine than any spec-sheet number: Windows splits the unified pool, and the GPU can’t touch all of it (NVIDIA RTX Spark Porting Guide — Unified Memory Architecture).
The pool is carved into three regions: a dedicated carveout reserved for the GPU (what Windows reports as “dedicated GPU memory” — it’s ordinary DRAM, not separate VRAM), a shared region both CPU and GPU can use, and a CPU-only remainder. The shared region is sized by formula: post-carveout capacity minus 16GB, clamped between 50% and 80% of post-carveout capacity. On a 128GB machine with a hypothetical 16GB carveout, that’s 112GB post-carveout, and the 80% clamp caps the shared region at ~89.6GB — roughly 105GB GPU-touchable in total, not 128GB. Microsoft says it’s raising the GPU-accessible limit on high-memory unified systems in current Windows builds (Windows Experience Blog), but no exact retail numbers exist yet.
The practical fallout, config by config: on the 128GB machine, a ~60GB gpt-oss-120b plus KV cache fits with room to spare even after the split — fine. On the $4,299 64GB config, it doesn’t: apply the same formula and the GPU-accessible pool lands in the high-40s of GB, short of the model’s 59GB of weights before you allocate a single token of context. That’s why the table above calls 64GB “marginal” — the config that looks like the budget path to 120B-class models probably isn’t one on launch-day Windows. If a 70B Q4 at ~7 tok/s doesn’t excite you (it shouldn’t), the 64GB tiers buy nothing the 32GB tiers don’t. Check your exact model-plus-context arithmetic in our VRAM calculator before you pick a tier.
This is the same class of problem Strix Halo owners hit with Windows GTT limits — and the fix there was Linux. On RTX Spark, Linux support exists (DGX OS on the sibling desktops), but the Surface ships with Windows 11 and an agentic-Windows pitch; buying it to immediately install Linux defeats the point of buying a Surface.
What $5,899 buys against the 128GB machines that already shipped
So the real decision is the top config or nothing. Here’s where $5,899 lands in the 128GB-class field we’ve been benchmarking all year:
| Machine | Price | Bandwidth | gpt-oss-120b decode | CUDA | Portable |
|---|---|---|---|---|---|
| Surface Laptop Ultra 128GB | $5,899 | ~300 GB/s | unproven (sibling DGX Spark: 46–61 tok/s) | Yes | Yes |
| GMKtec EVO-X2 128GB | $3,499–$3,649 | 256 GB/s | ~31 tok/s | No (ROCm/Vulkan) | No |
| ASUS Ascent GX10 | $3,099–$4,150 | 273 GB/s | 46–61 tok/s (GB10, llama.cpp #16578) | Yes | No |
| NVIDIA DGX Spark | $4,699 | 273 GB/s | 46–61 tok/s | Yes | No |
| Mac Studio M5 Max 128GB | $5,099 | 614 GB/s | 65–88 tok/s | No (Metal/MLX) | No |
| MacBook Pro 16” M5 Max 128GB | $6,999 | 614 GB/s | 65–88 tok/s class | No | Yes |
Read that table cold and the Surface’s problem is obvious: it’s the second-most-expensive machine in the class, with the second-slowest proven silicon family, and the slowest option — the EVO-X2 — costs $2,250 less for the same capacity. The GX10 runs the same GB10-family numbers for $2,800 less if you can live without a screen and battery. And if raw decode speed per dollar is the metric, the Mac Studio M5 Max wins outright: $800 cheaper than the Surface, double the bandwidth, measured 65–88 tok/s on the same model class.
What the table can’t show is the two-word combination no other row has: CUDA, portable. The MacBook Pro is portable and fast but locks you out of the CUDA ecosystem — fine for llama.cpp and MLX, a wall the moment your workflow touches vLLM features, NVFP4 ComfyUI pipelines, CUDA-only fine-tuning stacks, or local coding stacks that assume an NVIDIA backend. The EVO-X2, GX10, and DGX Spark have no battery. If your actual requirement is “develop against CUDA with a 100GB model in the seat next to me,” the Surface Laptop Ultra 128GB is — for now — the only product on Earth that does it, and $5,899 is what a monopoly config costs. That’s a real buyer. It’s just a much rarer buyer than Microsoft’s launch-event framing suggests.
For everyone else, the machines above it in value already exist, ship today, and have months of community benchmarks behind them. Our $3K/$6K/$10K build guide covers how the desktop paths slot into real budgets.
What to actually buy
Prices as of October 2026, all verified in the comparison above:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| Need 128GB + CUDA + a battery, accept v1 drivers | Surface Laptop Ultra 128GB | $5,899 | Microsoft Store (no affiliate relationship — go direct) |
| Cheapest 128GB that runs 120B-class MoE today | GMKtec EVO-X2 128GB | ~$3,499–$3,649 | Check price |
| Fastest 128GB machine under $5,500 | Mac Studio M5 Max 128GB | $5,099 | Check price |
| CUDA desktop on the same GB10 silicon, shipping now | ASUS Ascent GX10 | $3,099+ | Check price |
| Everything you run fits in 24GB | Used RTX 3090 | $1,190–$1,400 | Check price |
| Want to test 120B-class models before spending $3,500+ | Rented GPU | from $0.07/hr (3090), $0.25/hr (5090) | Vast.ai |
FAQ
Is the $2,599 base config good value for local AI? No. 24GB of capacity at ~300 GB/s is strictly worse for inference than a used RTX 3090 at 936 GB/s for around $1,300 — the 3090 holds the same models and decodes roughly three times faster. The base Surface is a nice Arm laptop that happens to run small models; it is not an AI purchase.
Why skip the 48GB and 64GB middle tiers? Bandwidth. Every config decodes at the same ~300 GB/s, so the middle tiers only add capacity for dense 70B models that decode at ~7 tok/s — below comfortable reading speed. And after the Windows memory split, the 64GB tier likely can’t fit gpt-oss-120b at all. The ladder’s useful rungs are the bottom (as a general laptop) and the top (as an AI machine); the middle buys capacity you can’t enjoyably use.
Should I preorder the 128GB config now or wait? Wait for the first retail reviews showing llama.cpp or Ollama running on CUDA — not Vulkan — at a sane fraction of the 100 tok/s MoE ceiling. The last public data point (a September prototype) had the CUDA path failing outright, and the DGX Spark’s post-launch kernel-swap speedup shows this platform’s software is still moving fast. Shipping starts October 16; reviews will exist within days of that. Two weeks of patience against $5,899 is cheap insurance.
Does the MSI Prestige N16 Flip AI+ change anything? It undercuts Microsoft by $400 for the same 20-core chip at 32GB/1TB ($3,299, November 6 per Newegg’s listing), which is good news for RTX Spark pricing generally and irrelevant for local AI specifically — no announced MSI config reaches 128GB yet.
What about the Surface RTX Spark Dev Box instead? If the reported ~$5,999 price holds, you’d be paying Surface-laptop money for a desktop with the same ~300 GB/s ceiling — while the DGX Spark ($4,699) and GX10 (from $3,099) do the same job on the same silicon family for less. Its one distinctive feature is shipping Windows 11 Pro with WSL2 GPU passthrough and CUDA preconfigured. Wait for a confirmed price.
Sources
- Microsoft opens Surface Laptop Ultra preorders at $2,599 with RTX Spark N1X inside — TweakTown
- Microsoft Surface Laptop Ultra Starts at $2,599, Ships October 16 — TechPowerUp
- NVIDIA RTX Spark laptops launch October 16, prices start at $2,599 — VideoCardz
- Surface Laptop Ultra with RTX Spark preorders are live — Windows Central
- NVIDIA RTX Spark laptops may start above $1,799, N1X systems reportedly above $2,899 — VideoCardz
- RTX Spark Porting Guide: Unified Memory Architecture — NVIDIA Docs
- Windows memory management changes for unified-memory systems — Windows Experience Blog
- Performance of llama.cpp on NVIDIA DGX Spark — ggml-org/llama.cpp discussion #16578
- MSI Prestige N16 Flip AI+: the RTX Spark laptop with the full N1X chip for $3,299 — Newegg Insider
- Microsoft’s Surface RTX Spark Dev Box will handle tougher AI workloads — Engadget
- Microsoft debuts Surface RTX Spark Dev Box — Tom’s Hardware
- Microsoft Opens Pre-Orders For Surface Laptop Ultra And Surface RTX Spark Dev Box — Pulse2
- NVIDIA’s first Windows-on-Arm GeForce driver confirms N1X specs — HWBusters
- NVIDIA confirms RTX Spark configurations and availability — Windows Central
Last updated October 8, 2026. Prices and specs change; verify current rates before purchasing.
Recommended Gear
- GMKtec EVO-X2 128GB — cheapest 128GB unified-memory box, $3,499–$3,649
- Mac Studio M5 Max 128GB — fastest 128GB machine under $5,500, $5,099
- ASUS Ascent GX10 — CUDA GB10 desktop from $3,099
- Used RTX 3090 24GB — the 24GB value answer, $1,190–$1,400
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.