Used RTX 4090 vs New RTX 5090 for Local AI in 2026: The Price Gap Just Doubled

rtx-4090rtx-5090used-gpugpulocal-llmbuying-guide

TL;DR: A used RTX 4090 averages $2,527 over 30 days (latest eBay solds near $3,010) while the RTX 5090 has no first-party stock at Amazon, Newegg, or Best Buy — marketplace listings run $5,199–$6,900, with Micro Center in-store around $4,299. The 4090 delivers roughly 80% of the 5090’s decode speed for about half the money. Pay for the 5090 only if you need 32GB or native FP4.

Used RTX 4090New RTX 5090Used RTX 3090
Best forDiffusion + fine-tuning + fast LLM in one box32GB capacity, FP4 speed, warrantyPure LLM inference on a budget
Price (Sep 2026)~$2,500–$3,010$4,299 in-store / $5,199+ online~$1,264
VRAM / bandwidth24GB / 1,008 GB/s32GB / 1,792 GB/s24GB / 936 GB/s
The catchExpired warranty + 12VHPWR historyCosts two used 4090s; still can’t fit 70B Q4~2× slower diffusion, no FP8/FP4

Honest take: At September’s prices the used 4090 wins this matchup for most buyers — but buy it like a used car: tested, returnable, connector inspected. The 5090 is only rational at Micro Center’s ~$4,299 walk-in price and only if 32GB or FP4 actually appears in your week. At $5,200+ marketplace pricing, rent the 5090 by the hour instead and keep the cash.

Three months ago this was a close call. In June 2026 the RTX 5090’s median street price was $4,299 and a used RTX 4090 averaged around $2,268 — a $2,000 gap that bought you 33% more VRAM, 78% more bandwidth, and a warranty. Then the memory supercycle ate both prices, and it ate the 5090’s faster. The RTX 5090 is now effectively out of stock at every first-party US retailer, and the decision has changed shape: it’s no longer “is the upgrade worth $2,000” but “is it worth $2,700–$4,400 — against a used card that itself got more expensive and riskier to buy.”

Here’s the September 2026 math, workload by workload.

The September 2026 price board

Both sides of this comparison moved hard in the last 60 days. Verified tracker data as of mid-September:

CardPrice (Sep 2026)Source basis
Used RTX 4090$2,527 30-day avg; latest eBay solds ~$3,010 (Sep 12), up 36.8% in 30 daysGetPCParts sold-listing tracker
Used RTX 4090 (fair asking)~$2,500 market average; ~$2,482 ceiling for a tested card with returnsResalePrices
New RTX 5090 (in-store)~$4,299 at Micro Center, when local stock existsPCGamesN retail check
New RTX 5090 (online)$5,199 best street price early Sep; cheapest orderable $6,449 (Newegg marketplace, Sep 14); tracker median $6,899.99 (Sep 15)Tech Insider / PCGamesN / videocardprices
Used RTX 3090~$1,264 averageResalePrices (Aug–Sep)

Two things jump out. First, the 4090’s used market is accelerating — a 36.8% 30-day move means the 90-day average ($2,364) badly understates what you’ll pay this week. Second, the 5090’s “price” is really two prices: if you live near a Micro Center with stock, you pay ~$4,299; if you don’t, the online marketplace premium starts at $5,199 and runs to $6,900. We covered why nothing in this market is getting cheaper in our GDDR7 price-hike analysis and the H200 supply story — the short version is that AI datacenter demand has swallowed the memory supply, and consumer cards are collateral damage.

At the June gap ($2,000), the 5090 was arguable. At today’s gap, you can buy a used 4090 and a used 3090 for less than one marketplace 5090 — 48GB of pooled VRAM against 32GB.

Decode speed: 25% faster, 100% more expensive

LLM generation speed follows memory bandwidth, and on paper the 5090’s 1,792 GB/s is 78% more than the 4090’s 1,008 GB/s. But measured single-stream decode doesn’t scale with the full spec gap on models that fit comfortably in both cards. The cleanest apples-to-apples numbers come from the llama.cpp community benchmark thread (Discussion #15396), testing gpt-oss-20b (MXFP4, tg128):

GPUgpt-oss-20b decodePrice paid (Sep 2026)$ per tok/s (our math)
Used RTX 3090161 tok/s~$1,264$7.85
Used RTX 4090225 tok/s~$2,527$11.23
New RTX 5090282 tok/s$5,199 (online)$18.44
New RTX 5090282 tok/s$4,299 (in-store)$15.24

On smaller dense models the gap widens toward the bandwidth ratio — Local AI Master measured a 5090 at 213 tok/s on an 8B Q4 where a 4090 lands around 127 — and on prompt processing the 5090 pulls furthest ahead, with Hardware Corner clocking roughly 10,000 tok/s prefill on gpt-oss-20b. If you run a coding agent that re-reads 30K-token contexts all day (the local BYOK setup crowd), prefill is what you actually feel, and it’s the strongest honest argument for Blackwell silicon.

But for interactive chat, calibrate against reading speed (~7–10 tok/s): a used 4090 at 225 tok/s and a 5090 at 282 tok/s are both generating text 20–30× faster than you read it. The 25% decode advantage converts to saved wall-clock time only in agentic and batch workloads.

The 32GB question: what 8 extra gigabytes actually unlock

The capacity argument is real but narrower than the spec sheet suggests. What 32GB gets you over 24GB, per our 32GB tier guide:

  • 27B–35B models with big context. A 24GB card runs Qwen3.6-35B-A3B but has to be stingy with context; 32GB runs the same model with 60K+ token headroom.
  • gpt-oss-20b at its full 128K context without offload — with the known catch that generation collapses to ~9 tok/s at 128K on any GPU, because that’s an attention-compute wall.
  • Q5/Q6 quants of the 27B class instead of Q4, if you’re chasing quality margins.

What it does not get you: the 70B tier. Llama 3.3 70B at Q4_K_M is a ~42GB GGUF — it doesn’t fit in 32GB any more than it fits in 24GB, and with CPU offload the 5090 manages a stuttery 14–22 tok/s. If 70B-class models are the goal, the honest paths are dual cards or unified memory, not a single 5090. Check your exact model + context combination in our VRAM calculator before you let 32GB be the deciding factor.

FP4: the one gap that grows every quarter

The 4090’s Ada silicon has FP8 tensor cores; Blackwell adds native FP4. That difference was a footnote in early 2026 and gets less footnote-shaped every month:

If your workload is image/video generation or you expect to ride the NVFP4 release wave (more models ship FP4-first every month), the 5090’s advantage is structural, not incremental. If you run GGUF Q4 chat models, it barely matters.

The risk file: what “used, out of warranty” means for this card

This is where the 4090’s price advantage picks up an asterisk. Three problems are specific to buying a 2022–2024 card in late 2026:

Warranties are expiring right now. Most AIB partners (ASUS, MSI, Gigabyte, Zotac) sold the 4090 with a 3-year warranty from the original purchase date — XDA’s coverage this year makes exactly this point: the fleet is crossing its third anniversary. A card bought at launch is out of coverage; most of the rest have months left, and several AIBs don’t honor transfers to second owners in the US at all.

The 12VHPWR history didn’t end. The 4090’s connector-melting saga is not a 2022 memory: a GPU repair shop owner told Wccftech he still receives hundreds of 4090s with burned 12VHPWR connectors, and warranty denials over third-party cables are documented. On a used card, you inherit whatever seating, bending, and thermal cycling the connector already survived. Before money changes hands: pull the connector, look for browned or deformed pins, and budget for a new ATX 3.1 native cable rather than trusting the included adapter.

You don’t know its duty cycle. At $2,500+, plenty of used 4090s coming to market now are exiting small AI farms and render fleets — sustained-load lives, not gaming evenings. That’s not disqualifying (datacenter-style steady load is arguably gentler than thermal cycling), but it strengthens the case for buying only from sellers with return windows, at the ~$2,482 “tested with returns” fair ceiling ResalePrices tracks, rather than chasing the cheapest private listing.

The new 5090 makes all three paragraphs disappear: 3-year warranty starting today, a revised 12V-2x6 connector (not incident-free, but current-generation support), and zero mystery hours. Whether that peace of mind is worth $1,800–$4,400 depends entirely on how the earlier sections landed for you.

Power, PSU, and running cost

Rated board power is 450W (4090) versus 575W (5090). At the $0.12/kWh US average we use in our home-server power math, that’s $0.054 versus $0.069 per full-load hour — our arithmetic. Four hours of daily inference costs about $1.80/month more on the 5090. Electricity won’t decide this.

The PSU might: NVIDIA recommends 850W for the 4090 and 1000W for the 5090, and Blackwell’s transient spikes make quality ATX 3.1 units the safe call. If your current build has an 850W PSU, the 5090 quietly adds a ~$180–$250 PSU replacement to its sticker price.

What to actually buy

Prices as of September 2026, all taken from the comparison above:

Your situationThe machinePriceWhere
Pure LLM inference — tokens per dollar is the metricUsed RTX 3090~$1,264Check price
Diffusion + fine-tuning + fast LLM in one boxUsed RTX 4090 (tested, with returns)~$2,500Check price
Need 32GB or FP4 today, near a stocked Micro CenterNew RTX 5090~$4,299 in-storeCheck price
Want 5090-class speed without paying the panic premiumRented 5090, from ~$0.25/hrpay per hourVast.ai

The rental row deserves one number: at Vast.ai’s market rates (RTX 5090 from ~$0.25/hr, RTX 4090 from ~$0.14/hr as of September 2026), a $5,199 marketplace 5090 buys roughly 20,000 hours of rented 5090 time. If your need for 32GB is a project rather than a lifestyle, rent it first and let the purchase decision wait out the price spike.

FAQ

Is a used RTX 4090 safe to buy in 2026? Reasonably — if you buy from a seller with a return window, at or below the ~$2,482 tested-card fair ceiling, and you inspect the 12VHPWR connector for browning or deformed pins before installing. Assume no warranty: most 3-year AIB coverage from 2022–2023 purchases has expired or doesn’t transfer.

Why not just buy a new RTX 4090? Production ended; remaining sealed stock lists at $3,400+, which is $900 from Micro Center’s 5090 price with none of the 5090’s VRAM, bandwidth, or warranty runway. It’s the worst square on this board.

Does the RTX 5090 run 70B models? Not in VRAM. Llama 3.3 70B Q4_K_M is ~42GB against 32GB of VRAM; expect 14–22 tok/s with CPU offload. For 70B ambitions, look at dual 24GB cards or unified-memory machines instead — the 5090’s real tier is 27B–35B with huge context.

Will 5090 prices come back down? Nothing in the current supply picture says soon: the card climbed from $4,299 (June) to $4,699 (August) to $5,199+ (September) with official retail out of stock, and memory contract prices are still rising. Waiting has been a losing strategy all year — which is an argument for the used 4090 or for renting, not for paying $6,900.

What about the RTX 5080 as the budget path to Blackwell? It’s 16GB. For local AI that’s a different (lower) tier than either card here — the model list that fits changes completely. See our 16GB tier guide before considering it.

Sources

Last updated September 17, 2026. GPU prices are moving weekly in this market — verify current listings before purchasing. Dollars-per-tok/s, electricity, and rental-hour figures are our own arithmetic on the sourced prices, benchmarks, and board-power ratings above.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.