RTX 5090 Gone From US Online Retail: What Local AI Builders Should Buy Instead in October 2026
TL;DR: The RTX 5090 has disappeared from US online retail — Newegg, Amazon, Best Buy, and B&H show no first-party stock, and third-party sellers are asking $6,500 to $9,500. This is a supply collapse, not a restock cycle, so stop waiting for it. A used RTX 4090 at $2,300–$2,800 is the speed pick; a Radeon AI PRO R9700 at $1,550–$1,700 is the 32GB pick.
| Used RTX 4090 24GB | Radeon AI PRO R9700 32GB | RTX 5090 (Micro Center, in store) | |
|---|---|---|---|
| Best for | Fastest sane-money CUDA card | Most VRAM per dollar, new with warranty | NVFP4 and 1,792 GB/s, if you can physically get one |
| Price (Oct 2026) | $2,300–$2,800 used | $1,550–$1,700 street | from $4,299, in-store only |
| The catch | No warranty, 12VHPWR history to inspect | 640 GB/s bandwidth — capacity, not speed | Stock is per-store luck; online it’s $6,500+ |
Honest take: Most readers who were saving for a 5090 should buy a used RTX 4090 and stop watching stock trackers. If your model list needs more than 24GB in one card, the R9700 is the only 32GB card still selling near list price. Pay $4,299 at a Micro Center if one is in stock near you and you genuinely need Blackwell; pay a scalper $6,500+ never.
Before you re-plan a build around any of these cards, run your actual model and context targets through the VRAM calculator — the right replacement depends entirely on whether your workload fits 24GB.
What actually happened
In mid-September 2026, Tom’s Hardware’s price tracker logged the RTX 5090 effectively vanishing from US online retail. Newegg, Amazon, Best Buy, and B&H all stopped showing first-party stock, leaving only third-party marketplace sellers asking between $6,500 and $9,500 for a card with a $1,999 MSRP. The slide was fast: the tracker’s June median was $4,299, the lowest online price at the start of September was $5,199, and within two weeks first-party inventory was gone. In the EU, Guru3D logged street prices above €5,000 (about $5,700) over the same period.
The one working exception is Micro Center, which sells GPUs in-person only. Tom’s Hardware found in-store 5090 stock starting at $4,299 at checked locations — still 2.15× MSRP, but $2,200 to $5,200 under what online third parties want.
The market reacted the way you’d expect. At Mindfactory, Germany’s big component retailer, AMD took roughly 56% of GPU unit sales in the first week of September while NVIDIA slipped to about 40% — a snapshot of what happens when the halo card is functionally unpurchasable and RX 9070 XT-class cards are sitting at the checkout. Keep it in proportion: Jon Peddie Research still puts NVIDIA around 90% of global discrete shipments. But for a buyer standing in front of a parts list this month, the global share doesn’t matter. The shelf does.
Why waiting won’t work this time
If this were a normal restock cycle, the answer would be “wait four weeks.” It isn’t, for three reasons this site has been tracking all year:
The memory crisis is the root cause, and nobody forecasts relief before late 2027. Samsung, SK Hynix, and Micron are running at full capacity, but the wafers are going to HBM for datacenter AI accelerators. The 5090’s 32GB of GDDR7 is one of the most supply-constrained components in the PC industry — the same squeeze that roughly quadrupled DDR5 prices since summer 2025. The deeper structural picture is in our NVIDIA-as-central-bank analysis.
There is no next card to wait for. Supply-chain reporting says no new consumer RTX gaming GPUs in 2026, with the RTX 60 series expected around 2028. The 5090 that exists now is the consumer flagship for the foreseeable future.
The alternatives are appreciating while you wait. GPU Poet’s used-market tracker has the RTX 4090’s lowest monthly average up 12.4% from July ($2,300) to September ($2,776). Used RTX 3090s are up double digits over 90 days. Every month of waiting for a 5090 restock that isn’t coming makes the fallback more expensive.
First, be honest about what the 5090 was buying you
The 5090 is three things in one package: 1,792 GB/s of memory bandwidth (the highest of any consumer card, per TechPowerUp), 32GB of VRAM, and Blackwell’s NVFP4 format for 4-bit image generation and inference. Decode speed for local LLMs is bandwidth-bound, so those translate directly: 282.5 tok/s on gpt-oss-20b in llama.cpp community benchmarks, and the only single consumer card that holds a 34B dense model at Q4 with real context headroom.
Almost nobody needs all three. Pick the one you were actually paying for, and the replacement picks itself.
If you were buying speed: used RTX 4090
The RTX 4090 has 1,008 GB/s of bandwidth — 56% of a 5090’s, but 8% more than a 3090’s — and the same 24GB of VRAM. On the same llama.cpp gpt-oss-20b benchmark thread, it decodes at 225 tok/s against the 5090’s 282.5. For dense 7B–13B models you’re well past reading speed on either card; the 5090’s margin only matters when you’re batch-serving or chasing maximum tokens per watt-hour of your own time.
October 2026 reality: GPU Poet’s tracker put the lowest monthly average at $2,776 in September, up from $2,317 in August — call it $2,300–$2,800 depending on cooler, condition, and whether the seller takes returns. That is painful for a five-year warranty-less card, and we said so in detail in the used 4090 breakdown and the used-4090-vs-5090 head-to-head: buy from a seller with returns, and inspect the 12VHPWR connector for heat discoloration before money changes hands.
One habit worth keeping in a market full of mining-era and AI-farm pulls: benchmark any used card the day it arrives, while the return window is open. A thirty-second llama.cpp sanity check against this site’s canon numbers catches a throttling or lane-starved card immediately:
$ ./llama-bench -m gpt-oss-20b-Q4_K_M.gguf -ngl 99
# healthy used RTX 4090: tg128 in the ~200–225 tok/s range
# healthy used RTX 3090: tg128 in the ~145–161 tok/s range
If you land far below that, check the PCIe link before blaming the card — nvidia-smi --query-gpu=pcie.link.gen.current,pcie.link.width.current --format=csv should report the full x16 at Gen 4. A card stuck at x4 (wrong slot, bad riser) loses prompt-processing speed and looks “defective” when it isn’t.
If you were buying 24GB on a budget: used RTX 3090, or the AMD discount
The RTX 3090 remains this site’s default recommendation for a reason: 24GB at 936 GB/s for $1,190–$1,400 (ResalePrices logs a $1,246 used average with a $1,189–$1,292 fair range across 301 listings; active eBay asks start around $1,397). It decodes gpt-oss-20b at 161 tok/s and 8B Q4_K_M at ~112 tok/s — 70% of a 4090 for half the money. The full case is in Used RTX 3090: still the value king.
The quiet bargain is still the RX 7900 XTX: 24GB at 960 GB/s — nominally more bandwidth than the 3090 — for a $806–$854 used fair range (ResalePrices, 106 listings; GPU Poet’s September lowest average was $817). Community llama.cpp numbers put it at 167–177 tok/s on 7B Q4_0 under Vulkan, class-parity with the 3090. The $400 you save is the CUDA tax you’re accepting: weaker diffusion and fine-tuning support, and day-one model support that usually lands on NVIDIA first. The full 7900 XTX breakdown covers who should and shouldn’t take that trade.
If you were buying 32GB: the R9700 is the only card still near list
Here’s the uncomfortable part: if the 5090 appealed to you specifically because 24GB wasn’t enough, no NVIDIA consumer card replaces it. The honest options:
Radeon AI PRO R9700 32GB — $1,299 MSRP, and in October 2026 actually buyable at $1,550 (PowerColor, Amazon, in stock) to $1,700 (ASRock, Amazon/Newegg). The catch is bandwidth: 640 GB/s, roughly a third of a 5090. On MoE models that barely matters — community llama-bench threads measured 127.4 tok/s decode on Qwen3.5-35B-A3B via Vulkan (147.8–156.3 with driver tuning), against 194 tok/s for a 5090 on the identical model. That’s 65–80% of the 5090’s MoE speed for roughly a third of Micro Center money and a quarter of scalper money. Dense 27B models run at an honest 29–33 tok/s. Full details and setup caveats in the R9700 hardware guide and the 5090-vs-dual-R9700 comparison.
Dual used RTX 3090s (48GB pooled) — about $2,400–$2,800 at current per-card prices, and the only consumer path that exceeds the 5090’s capacity: 48GB holds a 70B Q4_K_M dense model entirely in VRAM at 7–10 tok/s. You take on multi-GPU complexity — PCIe lane planning, power, and the IOMMU/NCCL pitfalls we documented — and a 700W continuous draw. The $3K/$6K/$10K build tiers guide shows where this path fits.
If what you really wanted was capacity for 100B+ MoE models rather than a fast 32GB card, that’s not a GPU decision anymore — unified-memory boxes (Strix Halo 128GB, Mac Studio) are covered across our mini-PC comparisons and win that fight on capacity per dollar.
The Micro Center exception, and when $4,299 is actually rational
A $4,299 in-store 5090 is 2.15× MSRP, and for most local AI workloads the math above says spend half that on a used 4090. The cases where driving to a Micro Center still makes sense: you need NVFP4 for 3× faster image generation on RTX 50-series, you’re batch-serving multiple users where 1,792 GB/s compounds, or your model list genuinely needs 32GB and CUDA in a single card. Stock is per-store and moves daily; call before driving. And since Micro Center is in-person only, treat any “Micro Center” listing you see online as a third-party reseller.
If you only need 5090-class hardware occasionally, rent it instead: 5090s on Vast.ai’s marketplace start around $0.25/hr — the $1,500–$2,000 premium between a used 4090 and an in-store 5090 buys roughly 6,000–8,000 rental hours before owning wins.
What not to do
Don’t pay a marketplace third party $6,500–$9,500. At $6,500 you’re within $7,500 of an RTX PRO 6000 Blackwell with 96GB and a real warranty — and at $9,500 you’ve bought two-thirds of one. Scalper pricing on a 575W consumer card with no transferable support is the worst deal in this entire market.
Don’t treat the RTX 5080 as the fallback. It shares the Blackwell name and almost nothing else that matters here: 16GB at 960 GB/s for $1,517–$1,799 street (September tracking). For local AI, 16GB is a different tier — it can’t hold the 20–24GB models that justify flagship money. A used 3090 costs less and runs more.
Don’t wait for a restock announcement. NVIDIA has announced nothing, the memory supply picture points the other way, and every tracked alternative got more expensive each month of 2026 you waited.
What to actually buy
Prices as of October 2026, all verified in the sections above:
| Your situation | The card | Price | Where |
|---|---|---|---|
| Want the fastest card still at sane money | Used RTX 4090 24GB | $2,300–$2,800 | Check price |
| Want 24GB + CUDA at the lowest buy-in | Used RTX 3090 24GB | $1,190–$1,400 | Check price |
| Pure llama.cpp inference, cheapest 24GB | Used RX 7900 XTX | $900–$1,000 | Check price |
| Need 32GB in one new card, warranty included | Radeon AI PRO R9700 | $1,550–$1,700 | Check price |
| Need 70B dense fully in VRAM | 2× used RTX 3090 (48GB) | $2,400–$2,800 | Check price |
| Truly need Blackwell/NVFP4 today | RTX 5090, Micro Center | from $4,299 | In store only — call ahead |
| Not sure yet — test the workload first | Rented 5090 | from $0.25/hr | Vast.ai |
FAQ
Will the RTX 5090 return to online shelves at normal prices? Nothing in the supply chain says yes. The DRAM/GDDR7 squeeze that caused this has no forecast relief before late 2027, NVIDIA is reported to be shipping no new consumer GPUs in 2026, and the June→September price curve ($4,299 median → $5,199 floor → gone) moved in one direction. Plan around the card’s absence.
Is the RTX 5080 a reasonable substitute? Not for local AI. Its 16GB ceiling excludes the model tier that makes flagship spending rational, and at $1,517–$1,799 it costs more than a used 3090 that runs larger models. See RTX 5090 vs 5080: 32GB vs 16GB.
Does the AMD Mindfactory number mean I should switch to AMD? It means AMD is what’s left on shelves, not that the software gap closed. For llama.cpp/Ollama inference, AMD is genuinely good now — the 7900 XTX and R9700 numbers above are real. For diffusion, fine-tuning, and day-one model support, CUDA still earns its premium. Match the card to your actual stack; our coding-tool sister site aicoderscope.com covers which local backends the big AI coding tools support.
I already own a 5090. Should I sell it? If it’s earning its keep in your rig, no — you can’t replace it at what you’d net. If it’s idle, $5,000+ of recoverable value against a $2,500 used 4090 that covers most workloads is a trade worth doing the math on.
What about renting until this blows over? Entirely reasonable, and cheaper than people assume: used-3090-class rentals start around $0.07/hr and 5090s around $0.25/hr on Vast.ai’s market. Our rent-vs-buy breakdown has the full math.
Recommended Gear
- Used RTX 4090 24GB — the speed pick, $2,300–$2,800
- Used RTX 3090 24GB — the value pick, $1,190–$1,400
- Used RX 7900 XTX 24GB — the budget 24GB pick, ~$820
- Radeon AI PRO R9700 32GB — the 32GB pick, $1,550–$1,700
Sources
- Nvidia’s RTX 5090 vanishes from online retail in the US — third-party sellers now demand as much as $9,500 — Tom’s Hardware
- The Nvidia RTX 5090 has vanished from retailer shelves in the US — TechRadar
- RTX 5090 Stock Dries Up Online as Third Party Prices Reach $9,500 — Digital Citizen
- GeForce RTX 5090 Street Prices in EU Now Exceed €5,000 — Guru3D
- AMD Overtakes Nvidia in GPU Sales (Mindfactory, first week of September 2026) as RTX 5090 Hits $9,500 — Tech-Insider
- NVIDIA GeForce RTX 4090 price history, September 2026: lowest monthly average $2,776 — GPU Poet
- NVIDIA GeForce RTX 3090 used market: $1,246 average, $1,189–$1,292 fair range — ResalePrices
- AMD Radeon RX 7900 XTX used market: $839 average, $806–$854 fair range — ResalePrices
- AMD Radeon RX 7900 XTX price, September 2026: lowest monthly average $817 — GPU Poet
- ASRock Radeon AI PRO R9700 Creator 32GB, $1,699.99 Amazon/Newegg — BuildCores
- PowerColor Radeon AI PRO R9700 32GB, $1,550 Amazon in stock — Computer Orbit
- gpt-oss-20b llama.cpp benchmarks: 282.5 tok/s (RTX 5090), 225 (RTX 4090), 161 (RTX 3090) — llama.cpp discussion #15396
- RTX 5090 (CUDA) vs R9700 (Vulkan), Qwen3.5-35B-A3B llama-bench: 194 vs 127.4 tok/s decode — llama.cpp discussion #19890
- GPU benchmarks on LLM inference (RTX 3090: 111.74 tok/s, 8B Q4_K_M) — XiongjieDai, GitHub
- NVIDIA GeForce RTX 5090 specifications: 32GB GDDR7, 1,792 GB/s — TechPowerUp
- RAM Price Index: DDR5 roughly 4× since summer 2025 — Tom’s Hardware
Last updated October 6, 2026. GPU prices and stock are moving weekly in this market; verify current listings before purchasing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.