Threadripper 9970X vs PRO 9965WX for Multi-GPU AI in 2026: When Is WRX90 Actually Worth It?

threadripperamdcpumulti-gpuworkstationlocal-llmhardware

TL;DR: The Threadripper 9970X ($2,499 list, 32 cores) on TRX50 now exposes 80 usable PCIe 5.0 lanes — enough for four GPUs at full or near-full width — so the classic “you need PRO for the lanes” rule died this generation. The Threadripper PRO 9965WX ($2,699.99 on Newegg, September 2026) buys 128 lanes, seven x16 slots, and 8-channel memory — but populating those channels during the worst RDIMM shortage in a decade is where the real cost hides.

Threadripper 9970X (TRX50)Threadripper PRO 9965WX (WRX90)Ryzen 9 9950X (AM5)
Best for2–4 GPU AI rigs5–7 GPUs, huge MoE CPU-offload, >1TB RAM1–2 GPUs at x8/x8
Cores / PCIe 5.0 lanes32C/64T, 80 usable24C/48T, 12816C/32T, 24 usable
Memory4-ch DDR5-6400, ~205 GB/s peak, 1TB max8-ch DDR5-6400, ~410 GB/s peak, 2TB max2-ch, ~90 GB/s
CPU price (Sep 2026)$2,499 list, dipped to $2,000$2,699.99 Newegg ($2,899 list)~$485
The catch4 memory channels cap CPU-offload speedBoard + 8 RDIMMs add ~$2,000+ over TRX5024 usable lanes wall at 2 GPUs

Honest take: Buy the 9970X on TRX50 for any rig up to four GPUs — the PRO premium now buys memory channels, not GPU lanes, and unless you’re CPU-offloading 200GB+ MoE models or racking more than four cards, that’s $2,000+ of platform cost you’ll never feel.

The Threadripper decision used to be simple: non-PRO chips gave you 48 PCIe 5.0 lanes, so anyone planning more than two GPUs paid the PRO tax. AMD quietly broke that logic with the 9000 series. Per TechSpot’s launch review, the non-PRO Threadripper 9000 chips expose up to 80 usable PCIe 5.0 lanes on TRX50 — up from 48 on the 7000 series. Four GPUs at x16 is 64 lanes. A non-PRO chip now covers that with 16 lanes to spare for NVMe.

So the question this article answers is narrower and more expensive than it looks: what exactly does the extra money buy when you step from a $2,499 32-core 9970X to a $2,899-list 24-core PRO 9965WX — and who building a local AI machine actually needs it?

The spec sheet, and the trap hiding in it

Both chips are Zen 5 “Shimada Peak” silicon on the same sTR5 socket, both 350W TDP, both 128MB of L3. The differences are all platform:

Spec9970XPRO 9965WX
Cores / threads32 / 6424 / 48
Base / boost clock4.0 / 5.4 GHz4.2 / 5.4 GHz
PCIe 5.0 lanes (usable)80128
Memory channels4 × DDR5-64008 × DDR5-6400
Theoretical bandwidth204.8 GB/s409.6 GB/s
Max RAM1TB2TB
ChipsetTRX50WRX90 (TRX50 compatible, with cuts)
List price$2,499$2,899

Notice the trap: at this price point, going PRO means losing eight cores. The PRO chip that matches the 9970X’s 32 cores is the 9975WX at $4,099 list — a $1,600 gap for identical core counts, per AMD’s official WX-series pricing. So the honest comparison isn’t “$400 more for PRO.” It’s either $400 more for fewer cores, or $1,600 more for the same cores. Either way, you’re paying for the platform, not the processor.

One compatibility note that trips people up: PRO chips physically work in TRX50 boards, but they drop to 4-channel memory and TRX50’s lane count when you do it — you pay the PRO premium and get none of it back. The reverse doesn’t work at all: WRX90 boards reject non-PRO chips outright. Pick the platform first, then the chip follows.

What the platforms actually cost in September 2026

CPU list prices are the least interesting part of this math. Here’s what the full platform looks like.

TRX50 side. Boards launched at $599 (Gigabyte TRX50 AERO D), $799 (ASRock TRX50 WS), and $899 (ASUS Pro WS TRX50-SAGE WIFI), per Wccftech’s launch pricing roundup. For multi-GPU builds the standout is the Gigabyte TRX50 AI TOP, which puts four double-spaced PCIe 5.0 x16 slots on one board — Gigabyte markets it specifically as a quad-GPU AI platform.

The best real-world anchor we found: Micro Center’s in-store bundle — 9970X + ASUS TRX50-SAGE Pro WS WiFi A + Kingston Fury Renegade Pro 128GB DDR5-5600 ECC RDIMM kit — for $3,299.99 (marked down from $4,284.97) as of September 2026. CPU, workstation board, and all four memory channels populated, for less than some WRX90 boards plus a PRO chip alone. The 9970X has also dipped to $2,000 on Amazon as a standalone part, per PC Guide’s deal tracking.

WRX90 side. There is effectively one board everyone buys: the ASUS Pro WS WRX90E-SAGE SE, with seven PCIe 5.0 x16 slots. It listed around $1,266 in August 2026, with retailer pricing running $1,299 and up. Add the 9965WX at $2,699.99 (Newegg, September 2026 — listing in Sources) and you’re at ~$3,970 before a single stick of RAM.

And the RAM is the ambush. WRX90’s whole value proposition is 8-channel bandwidth, which only exists if you populate all eight slots. You’d be doing that during a DRAM supply crunch in which a 64GB DDR5 RDIMM that contracted near $255 in Q3 2025 now crosses $900, with DDR5 ECC RDIMM pricing up 100–116% year-over-year and no meaningful relief expected before 2027. Eight channels means eight modules, bought at the worst possible time. A modest 8 × 32GB population runs well past what the entire TRX50 bundle above costs in memory alone at current street pricing; go 8 × 64GB for MoE offload headroom and the RAM bill alone can exceed $7,000.

Realistic totals for CPU + board + 128–256GB RAM, September 2026:

  • 9970X / TRX50: ~$3,300–$4,500
  • 9965WX / WRX90: ~$5,500–$8,000+

That $2,000–$3,500 platform gap is the number to hold in your head — not the $400 sticker difference between the chips.

The part spec sheets won’t tell you: cores don’t make tokens

If you’re eyeing the 32-core chip because more cores should mean faster local inference, the benchmark data says stop. In the llama.cpp CPU results from Phoronix’s launch testing, the 32-core 9970X and the 64-core 9980X landed within a single token per second of each other — 118 vs 119 tok/s in the same test, as compiled in popularai.org’s CPU inference ranking. Doubling the cores bought less than 1%, because token generation is memory-bandwidth-bound: quad-channel DDR5-6400 delivers roughly 180 GB/s in practice, and every core past the point of saturating that is idle silicon during decode.

This is exactly where the PRO’s 8 channels stop being a spec-sheet flex and start mattering — in one specific workload. If you run huge MoE models (DeepSeek-class, Kimi K2-class) with expert weights offloaded to system RAM via llama.cpp’s --n-cpu-moe, the offloaded experts are read at system memory speed, and doubling host bandwidth directly moves generation speed. 409.6 GB/s theoretical vs 204.8 GB/s is the difference between a 300GB-of-experts model being usable and being a slideshow.

But be honest about whether that’s you. We’ve made this argument before in our CPU-for-inference guide: if maximum host bandwidth per dollar is the goal, a used 8- or 12-channel EPYC platform undercuts any new Threadripper PRO build — Threadripper’s case has to rest on the GPUs, the single-threaded clocks for tool-heavy agentic work, and warranty, or it doesn’t hold at all.

GPU topology: what each platform physically holds

The TRX50 AI TOP’s four double-spaced x16 slots handle four two-slot cards — four used RTX 3090s for a pooled 96GB, the build we costed at $5,600–$7,200 total against a single RTX PRO 6000. Most other TRX50 boards give you two or three usable full-length slots once you account for triple-slot consumer card coolers, so check physical clearance before assuming the lane budget is your constraint.

The WRX90E-SAGE’s seven x16 slots only fully matter with blower-style or single/dual-slot cards — seven three-slot consumer GPUs don’t physically fit any EEB board. Its real audience runs five to seven workstation cards (RTX PRO 6000 Max-Q, R9700-class blowers) or wants four GPUs plus a stack of NVMe carriers and 100GbE NICs with zero lane compromises.

And keep lane width in perspective: for single-stream local inference, PCIe width barely moves tokens per second — we measured the NVLink/PCIe question already — it’s tensor-parallel serving (vLLM) and multi-GPU training where x16 vs x8 starts to show. A four-GPU llama.cpp rig on TRX50 at x16/x16/x16/x16 leaves nothing meaningful on the table vs WRX90.

What to actually buy

Prices as of September 2026, all taken from the comparison above:

Your situationThe machinePriceWhere
Building a 2–4 GPU local AI towerThreadripper 9970X + TRX50 (Micro Center bundle w/ 128GB ECC)~$3,300Check price
5+ GPUs, 200GB+ MoE CPU-offload, or >1TB RAMThreadripper PRO 9965WX + WRX90E-SAGE SE~$3,970 + RAMCheck price
Two GPUs max, x8/x8 is fineRyzen 9 9950X on AM5~$485Check price
Not sure your workload needs any of this yetRented GPU, test first3090 from $0.07/hr, 5090 from $0.25/hrVast.ai

The AM5 row deserves one sentence of respect: 24 usable PCIe 5.0 lanes runs two cards at x8/x8, which for inference costs you almost nothing — the entire Threadripper conversation only starts at GPU number three. And if what you’re actually sizing is which models fit in how much VRAM before committing to any of this, run your numbers through the VRAM calculator first.

FAQ

Can I put the PRO 9965WX in a TRX50 board and upgrade to WRX90 later? It boots, but it runs with 4 memory channels and TRX50’s lane budget — you’ve paid for 8 channels and 128 lanes you can’t use. If WRX90 is in your future, buy WRX90 now; the board is ~$1,300, not the expensive part anymore. The RAM is.

Why not the 24-core 9960X instead of the 9970X? Same 80-lane, 4-channel platform for about $1,000 less list ($1,499). For a GPU-centric rig it’s a legitimate saving — decode speed won’t change, as the 118-vs-119 tok/s result above shows. We spec the 9970X here because agentic workloads (compiles, tool calls, RAG indexing) do scale with cores, and it’s the bundle Micro Center discounts hardest.

Is 4-channel RDIMM enough to CPU-offload MoE models at all? Yes, with expectations set: ~205 GB/s theoretical is roughly half of WRX90 but still 2.3× a desktop AM5 board. Offloading a fraction of experts from a model that mostly fits your VRAM works fine; running a 300GB model mostly from system RAM is where 8 channels earns its cost.

Does either chip beat a GPU for inference? No. A used RTX 3090’s 936 GB/s is 4.5× the 9970X’s peak host bandwidth, and it costs ~$1,050. The CPU here is plumbing: lanes, RAM capacity, and tool-execution speed for agents. Spend the savings on VRAM — our sister site aicoderscope.com covers the coding-agent stacks that turn those GPUs into a local Copilot.

Products linked in this guide:

Sources

Last updated September 18, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.