NVIDIA RTX PRO 5000 72GB for Local AI in 2026: The $9,100 Card That Finally Fits gpt-oss-120b Whole
TL;DR: The RTX PRO 5000 Blackwell 72GB sells for $9,099–$9,575 in August 2026 — the first single card under five figures that holds gpt-oss-120b at full 131K context (68.5GB) with no offload flags. It’s also 43% less than the repriced $16,000 RTX PRO 6000. The catch: three used RTX 3090s deliver the same 72GB for $3,792.
| RTX PRO 5000 72GB (~$9,100) | 3× used RTX 3090 (~$3,792) | RTX PRO 6000 96GB ($16,000) | |
|---|---|---|---|
| Best for | One-slot-pair, one-plug 72GB; quiet office box | Cheapest 72GB in existence | 96GB tier, 196 tok/s on gpt-oss-120b |
| VRAM / bandwidth | 72GB / 1,344 GB/s | 72GB / 936 GB/s per card | 96GB / 1,792 GB/s |
| Power | 300W, one 16-pin | ~1,050W load, 1,600W PSU | 600W |
| The catch | $5,300 over the triple-3090 rig | Three x8+ slots, PCIe spelunking, no warranty | The Aug 13 repricing |
Honest take: If the machine lives under your desk, has to stay quiet, and the budget exists, this is the cleanest 72GB you can buy — one card, 300W, done. If you’re optimizing dollars per gigabyte, stop reading and buy three 3090s: same capacity, $5,300 cheaper, and faster on MoE models split across cards. The PRO 5000 72GB is a convenience purchase, not a value one.
Nine days ago we wrote that 72GB is where gpt-oss-120b stops needing excuses — the tier where the default local reasoning model’s full-context footprint (68.5GB) finally goes fully resident. At the time, the only way to own that number was three graphics cards and a motherboard that tolerates them. NVIDIA apparently agrees that 72GB is a real tier now: the RTX PRO 5000 Blackwell 72GB, launched quietly in December 2025 with 50% more memory than the 48GB workstation card it’s based on, has settled into US retail stock at $9,099.99 at Micro Center (marked down from an $11,999.99 list), $9,575 at ServerSupply, with stock also at B&H, CDW, and Newegg.
That price lands in a strange gap: more than double a used RTX 4090 ($2,268), about the same as a pair of RTX 5090s at their inflated $4,329–$4,699 street prices, and $6,900 under the RTX PRO 6000 Blackwell since NVIDIA repriced it to $16,000 on August 13. So the question the card actually has to answer isn’t “is 72GB useful” — it is. The question is who should pay $126 per gigabyte for it when the used market sells the same gigabytes at $53.
First, though, two corrections to the framing this card launched with — because the “what can it run” claims circulating on Reddit are partly wrong.
What 72GB actually runs (and the two headline models it doesn’t)
The viral pitch for this card asks whether it can run Llama 4 Maverick and Qwen3 72B on one GPU. The honest answers are no and that model doesn’t exist.
Llama 4 Maverick does not fit. Maverick has 17B active parameters, but it’s a 402B-total MoE — the Q4_K_M GGUF is 243GB, and even the quality-wrecking IQ1_78 ultra-quant is ~122GB. Nothing about 72GB changes Maverick math; the smallest machine that runs it whole is still a 128GB unified-memory box, and it runs it badly. Any listicle telling you Maverick fits this card in ~40GB has confused it with a 70B dense model.
“Qwen3 72B” isn’t a model you can download. Qwen3’s dense lineup tops out at 32B — 0.6B through 32B, with everything larger being MoE. The 72B people remember is Qwen2.5-72B, the 2024-generation model that Qwen3-32B already matches. You can run Qwen2.5-72B Q4_K_M (~47GB) on this card with room to spare. There’s just little reason to in August 2026.
What the card genuinely unlocks, from the official llama.cpp gpt-oss sizing tables and Unsloth’s published GGUF sizes:
| Model | Footprint | Fits 72GB? | Notes |
|---|---|---|---|
| gpt-oss-120b (MXFP4) | 64.0GB @ 8K ctx, 68.5GB @ full 131K | Yes — fully resident, full context | The reason this tier exists |
| Llama 3.3 70B Q5_K_M | 49.9GB + cache | Yes, with ~20GB of context room | Q6_K (~58GB) also fits |
| DeepSeek R1 Distill 70B Q4_K_M | ~40GB + cache | Yes, easily | ~25 tok/s measured (below) |
| Qwen2.5-72B Q4_K_M | ~47GB + cache | Yes | Superseded by Qwen3-32B |
| Qwen3.6-35B-A3B Q8_0 | ~37GB + cache | Yes, with 30GB+ for context | Fast MoE daily driver |
| GLM-5.3-Flash UD-IQ1_S | 93GB | No | Smallest usable quant is 93GB |
| Llama 4 Maverick Q4_K_M | 243GB | No | Not close, at any usable quant |
The pattern: 72GB perfects the 70B class and fully domesticates gpt-oss-120b, but the 2026 frontier MoEs — GLM-5.3-Flash at 93GB, everything above it — skipped straight past this tier to 96GB-and-up. That’s exactly the one-gigabyte-short story of the 64GB tier repeating one rung higher, and it’s worth knowing before you spend $9,100 that the next model generation may treat 96GB as the floor. Run your own model-plus-context numbers in the VRAM calculator before committing.
The spec sheet, and the one number that matters
Per PNY’s product page and datasheet and the B&H listing: Blackwell silicon with 14,080 CUDA cores (same GPU as the 48GB version — only the memory doubled, via 3GB GDDR7 modules), 72GB of ECC GDDR7 at 1,344 GB/s, 300W total board power through a single 16-pin connector, dual-slot full-height 10.5” board with an active blower cooler, four DisplayPort 2.1b outputs, PCIe 5.0 x16, and Multi-Instance GPU support — it can partition into isolated virtual GPUs, which none of the consumer cards do.
The number that matters is 1,344 GB/s, because decode speed is bandwidth-bound. Two useful reference points:
- It’s the same 1,344 GB/s as the 48GB version. You’re paying the ~$4,900 step over the 48GB card’s $4,199.99 street price purely for capacity, not speed.
- It’s exactly 75% of the RTX PRO 6000’s 1,792 GB/s. Every dense-model tok/s figure you’ve seen for the PRO 6000, scale by 0.75 and you’re close.
For home-lab practicalities, the card is unusually undemanding: 300W is RTX 3070 territory — a quality 750W PSU runs the whole box, and the blower exhausts out the back rather than dumping heat in the case. Compare that to ~1,050W of GPU load and a 1,600W PSU for the triple-3090 build, or 600W for the PRO 6000. This is the only 72GB configuration that fits in an ordinary mid-tower on an ordinary power circuit without further thought.
Real speeds: measured where they exist, scaled where they don’t
Labeled honestly — “measured” numbers are third-party benchmarks of this card; “estimated” numbers are bandwidth math from the PRO 6000’s measured results.
| Workload | RTX PRO 5000 72GB | Basis |
|---|---|---|
| DeepSeek R1 Distill 70B Q4 | ~25 tok/s | Measured — Every Local AI benchmark page |
| Llama 3.3 70B Q4_K_M | ~22–25 tok/s | Estimated — same size class as above |
| gpt-oss-120b MXFP4 | ~145 tok/s | Estimated — 75% of the PRO 6000’s 196 tok/s measured |
| Llama 3.1 8B FP16, vLLM batch serving | ~2,460 tok/s total throughput | Measured — DatabaseMart vLLM benchmark, 50 concurrent prompts, 4K ctx |
| GPT-OSS 20B, vLLM batch serving | ~3,359 tok/s total throughput | Measured — same DatabaseMart run |
Two readings of that table. For single-user chat on a dense 70B, ~25 tok/s is comfortably past reading speed — this card runs the 70B class well, not marginally. For gpt-oss-120b, the MoE structure (5.1B active of 117B total) means the estimated ~145 tok/s isn’t just usable, it’s faster than most cloud endpoints feel — and unlike the 64GB builds, there’s no --n-cpu-moe offload flag, no trimmed context, no asterisk:
./llama-server -hf ggml-org/gpt-oss-120b-GGUF \
--ctx-size 131072 --flash-attn on --jinja
On a 64GB rig that exact command dies at allocation time — ggml_backend_cuda_buffer_type_alloc_buffer: allocating 68512.00 MiB on device 0: cudaMalloc failed: out of memory — and the fix is either offloading expert layers to CPU (and eating the speed penalty) or capping context at 8K, one gigabyte from the finish line. On this card the same command loads 68.5GB into VRAM, leaves ~3.5GB free, and serves the full 131K window. That single no-flags command is what the $9,100 buys.
The batch-serving numbers matter for a different reader: with MIG partitioning and 2,400+ tok/s of aggregate 8B throughput, this is a legitimate one-card inference server for a small team — the kind of thing you’d otherwise rent an A100 for. If your models are coding assistants, our sister site covers wiring local endpoints into Cursor and Cline — a 70B at 25 tok/s is a plausible daily backend, and gpt-oss-120b at estimated triple digits absolutely is.
The math against everything else that reaches 64–96GB
August 2026 prices, using ResalePrices’ $1,264 used-3090 average and our previously verified street prices:
| Path | VRAM | Cost | $/GB | Power (load) | The tax |
|---|---|---|---|---|---|
| RTX PRO 5000 72GB | 72GB | $9,100 | $126 | 300W | Price |
| 3× used RTX 3090 | 72GB | $3,792 | $53 | ~1,050W | Slots, PSU, tuning, no warranty |
| 2× RTX 5090 | 64GB | ~$8,660 | $135 | ~1,150W | No NVLink — PCIe only; 8GB less |
| RTX PRO 6000 96GB | 96GB | $16,000 | $167 | 600W | The Aug 13 repricing |
| RunPod A100 80GB | 80GB | $1.39/hr | — | — | Nothing local, nothing private |
The dual-5090 row is the one to internalize: for effectively the same money as the PRO 5000, you get less VRAM, four times the power draw, and a tensor-parallel software stack to babysit — consumer Blackwell dropped NVLink entirely, so the two cards talk over PCIe. Unless you also game on the machine, the dual-5090 path lost this comparison the day 5090 street prices crossed $4,300.
The triple-3090 row is the one that keeps the PRO 5000 honest. $5,300 of savings buys a lot of tolerance for IOMMU groups and ACS overrides, and on MoE models llama.cpp splits layers across cards well enough that the speed difference is smaller than the bandwidth gap suggests. The 3090 rig loses on power (roughly 3.5× the wall draw — at $0.12/kWh and load, ~$0.13/hour versus $0.036/hour), on noise, on the motherboard it demands, and on the fact that every one of those cards is a 2020 GPU with no warranty in a market that has already priced 24GB cards like appreciating assets. It wins on the only number most people rank first.
And the rental row frames the whole decision: $9,100 at $1.39/hour is about 6,500 hours of A100 80GB time on RunPod — 4.5 hours a day, every day, for four years, with 8GB more VRAM than the card you’d have bought. If your 70B+ usage is sessions rather than a service, rent first and let your actual hours tell you whether this purchase exists.
Who this card is actually for
Strip away the spec sheet and three buyer profiles remain. The quiet-office builder: someone who wants gpt-oss-120b or a 70B resident on a machine that sits in a bedroom or office, sips 300W, and never sounds like a server — this card has no competition at 72GB, because the alternative is three blowers and a kilowatt. The small-team inference host: MIG partitioning plus the vLLM batch numbers make one card serve several developers’ local models, ECC and a warranty included, in ordinary IT-department packaging. And the workstation buyer with a real budget line: for whom the comparison isn’t against used 3090s at all, but against the $16,000 PRO 6000 — where “75% of the bandwidth and 75% of the capacity for 57% of the price” is simply a rational trade.
Everyone else — the value hunter, the tinkerer, the person who read this far hoping the verdict would justify the purchase — already owns the answer: the used market sells these same 72 gigabytes for $3,792, and the 24GB tier plus rented cloud hours covers most workloads for a fifth of either number.
FAQ
Is the RTX PRO 5000 72GB faster than the 48GB version? No. Same GPU, same 14,080 CUDA cores, same 1,344 GB/s bandwidth — only the memory capacity changed. You pay ~$4,900 more purely for capacity. If your target models fit in 48GB (a 70B at Q4 does, tightly), the 48GB card at $4,199 is the better buy.
Can it run GLM-5.3-Flash or the other 2026 frontier MoEs? No. GLM-5.3-Flash’s smallest usable quant is 93GB; the frontier MoE class starts above this card’s capacity. 72GB perfects the 70B/120B-MoE class — it does not reach the 300B+ class. That’s 96GB-and-up territory.
Does it work with Ollama and llama.cpp out of the box? Yes — it’s standard Blackwell (sm_120), the same architecture as the RTX 50-series consumer cards, so anything built with CUDA 12.8+ wheels runs unmodified. It also carries ECC memory and MIG, which the consumer cards lack.
Is there a used market for it yet? Barely. The card launched December 2025; eBay listings in August 2026 start around $9,440 — above Micro Center’s new price. Buy new, from a retailer, until that inverts.
Should I wait for prices to drop? The direction of travel argues the opposite: NVIDIA repriced the PRO 6000 to $16,000 on August 13, DRAM contract prices are still climbing, and the used 24GB market has risen for four straight months. Micro Center’s $9,099 (off an $11,999 list) is already the discount.
Recommended Gear
- NVIDIA RTX PRO 5000 Blackwell 72GB — the one-card 72GB path
- RTX PRO 6000 Blackwell 96GB — the step up, if the budget survives $16,000
- Used RTX 3090 24GB — three of these is the $3,792 value alternative
- RTX 5090 32GB — the single-card choice if 32GB is enough
Sources
- PNY NVIDIA RTX PRO 5000 Blackwell 72GB, $9,099.99 — Micro Center
- NVIDIA RTX PRO 5000 Blackwell 72GB (900-5G153-2270-000), $9,575 — ServerSupply
- NVIDIA RTX PRO 5000 72GB Blackwell product page — PNY
- RTX PRO 5000 Blackwell 72GB specs — B&H Photo
- NVIDIA RTX PRO 5000 Blackwell 72GB — CDW
- PNY RTX PRO 5000 Blackwell 48GB, $4,199.99 — Newegg
- Nvidia’s new RTX Pro 5000 Blackwell GPU with 72GB GDDR7 — Tom’s Hardware
- RTX Pro 5000 Blackwell 72GB launch coverage — TechRadar Pro
- gpt-oss GGUF memory sizing and PRO 6000 benchmarks — llama.cpp discussion #15396
- RTX Pro 5000 Blackwell vLLM inference benchmark — DatabaseMart
- RTX PRO 5000 72GB Blackwell local AI benchmarks — Every Local AI
- A100 PCIe 80GB pricing — RunPod
- Used RTX 3090 pricing — ResalePrices
- Qwen3 dense lineup — Hugging Face
Last updated August 30, 2026. Prices and specs change; verify current rates before purchasing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →