RTX 5090 vs RTX 4090 vs Used RTX 3090 for Local AI in 2026: Which Is Actually Worth Buying
TL;DR: Each step up this ladder roughly doubles the price: used RTX 3090 at ~$1,050, used RTX 4090 at ~$2,268, new RTX 5090 at $4,700 street. The 3090 wins price-per-token by a landslide; the 4090 buys speed, not capacity; the 5090’s extra 8GB still can’t hold a 70B at Q4.
| Used RTX 3090 | Used RTX 4090 | RTX 5090 (new) | |
|---|---|---|---|
| Street price (Sep 2026) | ~$1,000–$1,050 sold | ~$2,268 ($2,150–$2,350) | $4,699.99 median |
| VRAM / bandwidth | 24GB / 936 GB/s | 24GB / 1,008 GB/s | 32GB / 1,792 GB/s |
| GPT-OSS 20B speed | 161 tok/s | 225 tok/s | 282 tok/s |
| Our $/tok/s math | $6.50 | $10.10 | $16.70 |
| The catch | 2020 card, no warranty | Same 24GB as the 3090 | Costs more than 4× a 3090 |
Honest take: Buy the used RTX 3090 unless a coding agent is hammering your GPU for hours a day — then the 4090 earns its premium. Almost nobody should buy a $4,700 RTX 5090; the workloads that genuinely need more than 24GB are served better by renting an 80GB card by the hour.
The three-way flagship question gets asked constantly, and the answer changed twice this year because the prices moved — not the silicon. Back in May, our 5090 vs 4090 comparison could honestly say a used 4090 cost more than a new 5090 at MSRP. That’s dead. The 5090’s median US street price hit $4,699.99 in September 2026 (videocardprices tracking, up from $4,299.99 in June), Amazon listings run $5,069–$5,799, and one European retailer is literally asking $5,090 for an RTX 5090. The $1,999 MSRP is a number that exists on NVIDIA’s website and nowhere else.
So this comparison uses street prices, because that’s what your card will actually cost. Before reading further, it’s worth 30 seconds in our VRAM calculator to check what your target models need — the whole decision hinges on whether 24GB covers you.
What each card costs right now
Used RTX 3090, ~$1,000–$1,050. That’s the eBay sold-price average tracked by Best Value GPU and GPUDojo in September 2026. Fair warning on a spread you’ll notice while shopping: asking prices on active listings currently run $1,287–$1,411 (ResalePrices), and the used average is up 11.3% over the past 90 days. Sold prices are what matter — patient buyers still land near $1,050, but the window is drifting upward, not down.
Used RTX 4090, ~$2,268. Best Value GPU’s August 2026 eBay average, with ResalePrices tracking a fair range of $2,150–$2,350. Four years after launch, the card trades 40% above its $1,599 MSRP — 24GB of CUDA-capable VRAM simply doesn’t depreciate in this market.
RTX 5090, $4,699.99 median street. That’s 135% over the $1,999 launch price, and it crossed $5,000 at several US retailers in early September per Tom’s Hardware coverage. The driver is datacenter spillover — H100/H200 supply is so constrained that businesses buy 5090s as substitutes, and home-lab buyers get outbid.
Notice what happened to the ladder: each rung is roughly 2× the one below it. That framing makes the rest of this article easy.
Same model, same software: the three-way benchmark
The cleanest apples-to-apples numbers come from the llama.cpp community benchmark thread (Discussion #15396), which tested gpt-oss-20b (MXFP4, tg128) across all three cards:
| GPU | GPT-OSS 20B generation | Price paid | $ per tok/s (our math) |
|---|---|---|---|
| Used RTX 3090 | 161 tok/s | ~$1,050 | $6.52 |
| Used RTX 4090 | 225 tok/s | ~$2,268 | $10.08 |
| RTX 5090 | 282 tok/s | ~$4,700 | $16.67 |
The dollars-per-token column is our arithmetic on the verified prices and benchmarks above, and it’s the most honest summary of this market: the 5090 delivers 1.75× the 3090’s speed for 4.5× the money. The 4090 sits in between — 1.4× the speed for 2.2× the price.
Two things the table understates. First, generation speed scales with memory bandwidth, and the gaps here (936 → 1,008 → 1,792 GB/s) show up consistently across model sizes — on 8B Q4 models, Local AI Master measured the 5090 at 213 tok/s where a 4090 lands around 127. Second, the 5090 and 4090 pull far ahead on prompt processing: Hardware Corner clocked the 5090 at roughly 10,000 tok/s prefill on gpt-oss-20b. If your workload is a coding agent re-reading 30K-token contexts all day, prefill speed is what you actually feel, and it’s the strongest genuine argument for the newer silicon.
But calibrate against human reading speed (~7–10 tok/s): all three cards generate chat responses faster than anyone reads them. The speed difference only converts to saved time in agentic and batch workloads — which is exactly the local coding-agent setup crowd, and roughly nobody running a chat UI in the evenings.
The 32GB myth: what the 5090 does and doesn’t unlock
The intuitive case for the 5090 is “8 more gigabytes = bigger models.” Here’s where that breaks down, and it’s the trap that catches the most buyers:
Problem: You buy a 5090 to run Llama 3.3 70B at Q4_K_M, the quantization everyone recommends. You load it and llama.cpp starts offloading to system RAM anyway. The 70B Q4_K_M GGUF is ~42GB — it doesn’t fit in 32GB any more than it fits in 24GB. With offload, ModelFit’s testing puts 70B Q4 on a 5090 at 14–22 tok/s, CPU-bound and stuttery.
Fix: Run 70B at Q3/IQ3 instead (fits native, with quality loss that’s measurable in benchmarks), cap your context, or accept that dense-70B-at-good-quant is a 48GB problem — see our 48GB tier guide for the dual-3090 path that costs ~$2,100, less than half a 5090.
What the 32GB genuinely buys you over 24GB:
- The 27B–35B class with full context headroom. A 3090 runs Qwen3.6-35B-A3B at 107 tok/s but has to be stingy with context; the 5090 runs the same model with room for 60K+ contexts.
- gpt-oss-20b at its full 128K context without offload — the first consumer card that can, per Hardware Corner. Worth knowing the catch: at 128K, generation collapses to ~9 tok/s on any GPU, because that’s an attention-compute wall, not a VRAM wall.
- 70B at Q3, natively. Usable, but a quality compromise on a $4,700 card.
That’s a real but narrow slice. For most local-AI workloads in 2026 — the 20B–35B sweet spot where the best current models live — 24GB is enough, which is precisely why the 24GB used market refuses to get cheaper.
Power, heat, and what the wall socket sees
Rated board power: 350W (3090), 450W (4090), 575W (5090). At the $0.12/kWh US-average rate we’ve used in our home-server power math, that’s $0.042, $0.054, and $0.069 per hour of full load respectively — our arithmetic. Run four hours of inference a day and the 5090 costs about $3.30/month more than the 3090. Electricity won’t decide this; infrastructure might:
- The 3090 is happy on a quality 850W PSU. The 4090 wants 850–1,000W. The 5090 needs 1,000W minimum with 1,200W recommended — an upgrade that adds $150–$250 if you don’t already own one.
- The 4090 and 5090 use the 12V-2x6 connector (successor to the early-4090 12VHPWR melting saga); seat it fully, no adapters daisy-chained.
- Used-3090 buyers: run a day-one soak test. Load a model and watch for thermal throttling or ECC-style glitches:
$ nvidia-smi --query-gpu=temperature.gpu,power.draw,memory.used --format=csv -l 5
temperature.gpu, power.draw [W], memory.used [MiB]
68, 347.21 W, 21843 MiB
71, 351.44 W, 21843 MiB
Thirty minutes of sustained load with temps under ~83°C and no crashes de-risks most of the “2020 card, no warranty” objection. VRAM on 3090s runs hot (GDDR6X on the back of the board); if temps creep past 100°C on memory junctions, repad or repaste — a $20 fix that’s well documented for this card.
Verdicts, one per buyer
Buy the used RTX 3090 (~$1,050) if you’re like most readers of this site: running 8B–35B models, experimenting with local coding models, image generation, learning the stack. It’s the best price-per-token on the market by a factor of 1.5 over the next option, it runs the same model catalog as the 4090, and six years after launch it’s still the value king. The $1,200 you save over a 4090 buys 64GB of DDR5, a good PSU, and a fast SSD — a better system, not just a card.
Buy the used RTX 4090 (~$2,268) if the GPU is a tool you bill hours against: agentic coding runs, batch document processing, anything where 1.4× generation and ~2× prefill compound over thousands of calls a day. Same models as the 3090, materially less waiting. The full worth-it math comes down to utilization — above a few hours of active generation daily, it pays for itself in time.
The RTX 5090 ($4,700) is the right buy for almost no one at this price. The honest case requires all three: you need 25–32GB resident (not 24, not 40), you need it many hours a day, and you can’t tolerate cloud latency or data egress. That’s a real person — a startup prototyping on 32B fine-tunes, maybe — but if you’re reading a buying guide to decide, you’re not them. The card is 4.5× a 3090’s price for 1.75× its speed, and the one capacity tier it “unlocks” (70B) it only runs at reduced quantization.
If you occasionally need more than 24GB, don’t buy it — rent it. An 80GB A100 or H100 on RunPod handles the rare 70B-at-Q8 or fine-tuning job for single-digit dollars per session, and our rent-vs-buy breakdown shows the crossover: at $4,700, the 5090 needs years of heavy utilization to beat renting the big-VRAM jobs while a $1,050 3090 handles the daily ones.
For the budget tiers below this flagship conversation, the buyer’s guide by budget covers $500–$2,500 with one pick per tier.
FAQ
Isn’t buying a 6-year-old GPU with no warranty risky? Less than it sounds. The failure modes are known (fan bearings, VRAM thermal pads), the soak test above catches most bad cards inside the return window, and at ~$1,050 you’re risking a quarter of a 5090. Buy from a seller with real feedback and pay via a platform with buyer protection.
Will the 5090’s price come back down to $1,999? Nothing in the September 2026 data suggests it soon: the median rose $400 between June and September, and the driver (datacenter demand spillover) is a structural shortage, not a launch spike. Treat the street price as the price.
Two used 3090s instead of one 4090 or 5090? For capacity, yes — ~$2,100 gets you 48GB and the dense-70B class at Q4, which no single card here runs properly. You take on multi-GPU complexity and ~700W of draw; the 48GB tier guide walks through it.
Does the 3090 support the same software as the newer cards? CUDA compute capability 8.6 is fully supported by llama.cpp, vLLM, ExLlama, and ComfyUI in 2026. You miss FP8/FP4 hardware paths (Ada/Blackwell features) — relevant for some quant formats like NVFP4, not for the GGUF/Q4 workflows most people run.
What about a new RTX 5080 instead of a used card? It’s a 16GB card at ~$1,400+ — it loses the entire 27B–35B class to save $350 versus… nothing, since the 3090 is cheaper. We compared the 5090 and 5080 directly; for local AI, 16GB at that price is parked money.
Recommended Gear
The cards compared in this guide (affiliate links — at no extra cost to you, purchases support the site):
- Used RTX 3090 24GB — ~$1,000–$1,050 sold, the value pick
- Used RTX 4090 24GB — ~$2,268, the speed pick
- RTX 5090 32GB — $4,700+ street, rent before you buy this
Sources
- RTX 5090 Price Tracker, September 2026 — videocardprices.com
- RTX 5090 on sale for $5,090 — Aroged, Sep 1 2026
- RTX 5090 price surge on AI demand — shattered.io
- RTX 3090 Price Tracker US, Sep 2026 — Best Value GPU
- RTX 3090 24GB Used Price & History, Sep 2026 — GPUDojo
- RTX 3090 Used Fair Asking Range — ResalePrices
- RTX 4090 Price Tracker US — Best Value GPU
- RTX 4090 Used Price & Fair Asking Range — ResalePrices
- gpt-oss benchmark thread — llama.cpp Discussion #15396
- RTX 5090 LLM benchmarks: prompt processing and 128K context — Hardware Corner
- RTX 5090 model fit: 32B Q4 yes, 70B Q4 no — ModelFit
- RTX 5090 vs 5080: 213 vs 132 tok/s on 8B — Local AI Master
- RTX 4090 specs, 1,008 GB/s bandwidth — RunPod docs
- RTX 5090 specifications — NVIDIA
Last updated September 5, 2026. GPU prices move weekly — verify current listings before purchasing. Dollars-per-tok/s and electricity figures are our own arithmetic on the sourced prices, benchmarks, and board-power ratings above.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →