ASUS Ascent GX10 vs GMKtec EVO-X2 128GB for Local AI in 2026: The CUDA Box Became the Cheap One

asus-ascent-gx10gmktec-evo-x2gb10strix-halomini-pclocal-llmhardware

TL;DR: Both are 128GB unified-memory boxes and both decode gpt-oss 120B in the low-to-mid 30s tok/s, because 273 GB/s and 256 GB/s are nearly the same bandwidth. The GX10 adds CUDA and ~5× faster prompt processing; the EVO-X2 is a normal x86 PC with far lower idle power. The old reason to pick AMD — price — is gone: the GX10 1TB now streets at $3,099–$3,400 versus the EVO-X2 128GB’s $3,499–$3,649.

ASUS Ascent GX10 (1TB)GMKtec EVO-X2 (128GB)Used RTX 3090 tower
Best forCUDA, fine-tuning, agentic prefillAll-day MoE inference, normal-PC dutyEverything that fits in 24GB
Price (Sep 2026)$3,099–$3,400 street, patchy stock$3,499–$3,649~$1,400 complete build
The catchArm-only DGX OS box, single-purpose~340 tok/s prefill, ROCm rough edgesHard 24GB ceiling

Honest take: If a GX10 1TB is in stock near $3,100, buy it over the EVO-X2 — same capacity, same decode speed, 5× the prefill, and the CUDA ecosystem, for less money. The EVO-X2 still wins one specific buyer: the person who wants a general-purpose x86 machine that also serves 100B-class MoE models all day at 8–14W idle.

The ASUS Ascent GX10 is NVIDIA’s GB10 platform — the same silicon as the $4,699 DGX Spark — in ASUS clothing. The GMKtec EVO-X2 is the best-known AMD Strix Halo mini PC. Cross-shopping them is the real decision most $3,000–$4,000 buyers face in late 2026, and until now this site only compared each against the DGX Spark (GX10 vs Spark, Strix Halo vs Spark). Head-to-head, the answer has changed since summer — because the prices did. Run your target models through the VRAM calculator first; if everything you run fits in 24GB, skip both boxes and read the last section.

The price flip nobody noticed

When Strix Halo machines launched, their pitch against the GB10 platform was brutal and simple: the same bandwidth-bound decode speed for $1,700 less than a $4,699 DGX Spark. That framing died quietly over the summer, from both directions.

The GX10 undercut the Spark from day one and kept sliding. NVIDIA’s own list price for the platform went up — the Spark jumped to $4,699 in February 2026 — but ASUS’s US list price of $3,999 for the GX10 hasn’t held either: Amazon has stocked the 1TB model at $3,099.99, price trackers logged a $3,461–$3,999 range across 2026, and business resellers list it at ~$3,384 as of September. Stock is the catch — the 1TB GX10 goes in and out of availability, and the 4TB version runs ~$4,150.

The EVO-X2 moved the other way. The DRAM supercycle repriced soldered LPDDR5X like everything else, and the 128GB/2TB configuration that once promo-priced under $2,000 now sits at $3,649 on Amazon, with GMKtec’s newer 128GB/1TB variant at ~$3,499. (The 64GB model at $2,199 is a different decision — we covered why 64GB vs 128GB is a one-shot call on this soldered-RAM platform.)

Net result as of September 2026: the CUDA box is $100–$550 cheaper than the AMD box when stocked at street price, and at worst a few hundred dollars more if you catch a bad listing week. The decision is no longer about price. It’s about what the two architectures are actually good at.

Same capacity, different silicon

ASUS Ascent GX10GMKtec EVO-X2 (128GB)
ChipNVIDIA GB10: 20-core Arm CPU + Blackwell GPUAMD Ryzen AI Max+ 395: 16 Zen 5 cores + Radeon 8060S iGPU
Memory128GB LPDDR5X @ 273 GB/s128GB LPDDR5X-8000 @ 256 GB/s
ComputeUp to 1 petaFLOP sparse FP440 RDNA 3.5 CUs + 50 TOPS XDNA 2 NPU
OSDGX OS (Ubuntu-based, Arm)Windows 11 or Linux, ordinary x86
Software stackFull CUDA, TensorRT, NIMROCm / Vulkan / llama.cpp
Rated draw240W box147–160W measured under dense-70B load
Storage1TB or 4TB (soldered-adjacent; CID warranty limits swaps)User-accessible M.2 slots

Two numbers in that table decide almost everything. The first is memory bandwidth: 273 vs 256 GB/s, a 7% difference. LLM decode is bandwidth-bound, so token generation speed on identical models lands within a few percent on these machines — no driver update changes that. The second is FP4 compute: the GB10’s petaFLOP-class tensor hardware doesn’t help decode at all, but it demolishes the Radeon 8060S at prefill, which is compute-bound.

Inference: a decode tie, a 5× prefill gap

The Register’s head-to-head testing of the Strix Halo platform against the GB10 platform put gpt-oss 120B decode at 34.13 tok/s (AMD) vs 38.55 tok/s (NVIDIA) — a 13% gap that tracks the 7% bandwidth difference plus software polish. On the EVO-X2 specifically, independent testing lands around 31 tok/s on the same model (ServeTheHome measured the identical-chip Beelink GTR9 Pro at 31.41 tok/s and 125–128W; same chip, same speed). StorageReview’s vLLM testing confirmed the GX10 performs identically to other GB10 systems, so DGX Spark numbers transfer directly.

WorkloadGX10 (GB10)EVO-X2 (Strix Halo)
gpt-oss 120B decode~38 tok/s stock, up to ~60 tok/s (LMSYS tuned SGLang)~31–34 tok/s stock, mid-40s–low-50s with llama.cpp Linux tuning
gpt-oss 120B prefill~1,720 tok/s~340 tok/s
Dense Llama 3.3 70B decode~2.6 tok/s~5 tok/s (Q6_K, 147–160W)
Qwen3-235B-A22Bfits, bandwidth-bound~11 tok/s
Fine-tuningLoRA Llama 3 8B at 53K peak tok/s (StorageReview)works via ROCm, rougher and slower

Read that table by row, not by column. Decode: effectively a tie — both feel comfortably interactive on MoE models (human reading speed is ~7–10 tok/s). Prefill: not a tie at all. If your workload is agentic — long system prompts, RAG contexts, coding assistants re-reading a repo every turn — the GX10 starts responding roughly 5× sooner, and that difference is felt on every single request. That’s the same conclusion our Spark vs Strix Halo comparison reached, and it’s the workload split that should drive this purchase. If you’re pairing either box with a local coding stack, the prefill row is the one that matters — see the editor-side setups at aicoderscope.com for what those context sizes look like in practice.

Dense 70B models are bad on both (~2.6–5 tok/s). Nothing at 256–273 GB/s fixes that; a Mac Studio M5 Max at 614 GB/s roughly doubles it, and that’s still not fast.

Software: a DGX appliance vs a normal PC

The GX10 runs DGX OS on an Arm CPU. That buys you the full CUDA stack — day-one model support, TensorRT, NIM containers, PyTorch training that just works — with updates that flow from NVIDIA through ASUS validation. It also means the machine is an appliance: no Windows, no x86 binaries, no gaming, no second life as a family desktop. You’re buying it to do AI work, full stop. One more GB10 quirk to know about: the SSD sits under a warranty-void CID sticker, so buy the storage tier you’ll actually need.

The EVO-X2 is an ordinary x86 PC that happens to have 128GB of unified memory. It ships with Windows 11, dual-boots Ubuntu, runs Steam, and can be repurposed the day you upgrade. The AI cost of that flexibility: ROCm on the gfx1151 iGPU still has rough edges (llama.cpp’s Vulkan backend is often the pragmatic choice, and if rocminfo can’t see the GPU you’re in HSA_OVERRIDE territory), and the best decode numbers on Strix Halo come from community Linux tuning, not the out-of-box Windows experience.

Power is the EVO-X2’s quiet win. Reviewers measured 8–14W idle and 147–160W under sustained dense-70B inference; at $0.12/kWh, an always-on EVO-X2 idles for about a dollar a month. The GB10 box is rated at 240W and isn’t built around idle frugality. For a 24/7 home server, idle power dominates the annual bill — we ran that math in the power-bill breakdown.

The GX10 bug you should know about (and its fix)

The GB10 platform’s headline expansion trick — pooling two boxes over the 200 Gbps ConnectX-7 link — shipped broken on early GX10 firmware. Owners measured inter-box bandwidth capped at ~13 Gbps because the ConnectX-7 was being power-throttled at 27W on the PCIe bus, which NVIDIA’s forums traced to driver 580.126.09. The fix is a standard package upgrade to 580.142:

$ sudo apt update && sudo apt full-upgrade
# NVIDIA driver 580.126.09 → 580.142, then reboot

$ iperf3 -c 192.168.100.2
# before: ~13 Gbits/sec
# after:  ~196 Gbits/sec

Two takeaways. First, if you’re buying a GX10 planning to stack a second one later for 256GB, run apt full-upgrade before benchmarking anything. Second, this is what the ASUS update path looks like in practice: the fix arrived, but through ASUS’s validated channel rather than NVIDIA-direct — the trade we detailed in the GX10 vs Spark piece.

The EVO-X2’s equivalent gotcha is smaller: out of the box on Windows, the default runtime allocates a fraction of the 128GB to the GPU; you set the GPU memory split in BIOS (up to 96GB) and use Linux + llama.cpp for the published-best numbers.

What neither box fixes

At 256–273 GB/s, dense models above ~30B are slow on both machines, and models that fit in 24GB are faster on a used discrete GPU. A used RTX 3090 at $1,150–$1,350 pushes 936 GB/s — roughly 3× the decode throughput of either box on 7B–14B models — and a complete build lands near half the price of either machine (why it’s still the value king). These 128GB boxes only make sense once your models physically don’t fit on a real GPU. The VRAM calculator settles which side of that line you’re on in about a minute.

What to actually buy

Prices as of September 2026, all verified in the comparison above:

Your situationThe machinePriceWhere
Agentic/coding workloads, fine-tuning, or you live in CUDAASUS Ascent GX10 (1TB)$3,099–$3,400 when stockedCheck price
All-day MoE inference server that’s also a normal PCGMKtec EVO-X2 128GB$3,499–$3,649Check price
Everything you run fits in 24GBUsed RTX 3090$1,150–$1,350 (card)Check price
Want first-party NVIDIA support and 4TB, price no objectNVIDIA DGX Spark$4,699Check price
Undecided — test the workload before spending $3K+Rented GPU, RTX 4090 from ~$0.14/hrpay per hourVast.ai

That last row is the honest first step if you’ve never run a 100B-class MoE: rent a big-VRAM instance for an evening, measure your actual prompt sizes and tok/s tolerance, and you’ll know which column of this table you’re in before committing to soldered, non-upgradable memory.

FAQ

Is the ASUS Ascent GX10 really the same as a DGX Spark? Same GB10 chip, same 128GB at 273 GB/s, and StorageReview measured identical vLLM performance across GB10 systems. You’re choosing storage tier, chassis, and update path — we compared them in full here.

Which is faster for chat-style inference? Nearly a tie. gpt-oss 120B decodes at ~31–34 tok/s on the EVO-X2 and ~38 tok/s on the GB10 platform stock — both well past reading speed. Tuned stacks raise both (mid-40s–50s on Strix Halo Linux llama.cpp, ~60 tok/s on GB10 via LMSYS’s SGLang build).

Which is faster for agents, RAG, and coding assistants? The GX10, decisively. Prefill runs ~1,720 tok/s vs ~340 tok/s — about 5× — so long contexts start answering much sooner. Prefill is compute-bound, and that’s where Blackwell’s FP4 hardware works.

Can the EVO-X2 run CUDA software? No. It’s ROCm/Vulkan territory. Most inference stacks (llama.cpp, LM Studio, Ollama) run fine, but CUDA-only tooling — much of the fine-tuning and TensorRT ecosystem — needs the NVIDIA box.

Why not just buy a used RTX 3090 instead? If your models fit in 24GB, you should — it’s ~3× the decode speed for a third of the price. These 128GB boxes exist for the 70B+ MoE class that no consumer GPU can hold.

Sources

Last updated September 30, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.