RTX 50 SUPER in Late 2026: The 24GB Cards Are Built, the Launch Window Just Closed — What to Buy Instead
TL;DR: The RTX 50 SUPER refresh — including the 24GB RTX 5070 Ti SUPER that local AI builders have been waiting on for a year — is physically built and sitting at board partners with no launch date. The Q3 2026 window NVIDIA gave its partners ends September 30 with nothing shipped, and supply-chain chatter now points to CES 2027 at the earliest. If you need more VRAM, buy it on today’s used market.
| Buy used 24GB now | Buy new 16GB Blackwell | Keep waiting for SUPER | |
|---|---|---|---|
| Best for | Most local-AI builders | Models fit 16GB; want warranty + FP4 | Gamers with no deadline |
| Price / Cost | RTX 3090 ~$1,150–$1,350; RX 7900 XTX ~$900–$1,130 | RTX 5070 Ti ~$1,000–$1,200 (MSRP $749) | $0 today; rising used prices while you wait |
| The catch | No warranty | 16GB ceiling — 27B Q4 doesn’t load | No date, no guarantee, leaked MSRPs won’t survive this market |
Honest take: The SUPER refresh was the last plausible “just wait” argument for 24GB under $1,000. It lost to memory economics. A used RTX 3090 at ~$1,250 today beats a phantom $749 card that has spent two months in a warehouse with no ship date.
The RTX 5070 Ti SUPER was supposed to be the card this site never got to recommend: 24GB of VRAM, new, with a warranty, at a leaked $749–$799 MSRP. That’s the exact spec that runs Qwen3.8-27B at Q4_K_M with room to spare — the thing a $1,000–$1,200 RTX 5070 Ti can’t do at any price. If you’ve been holding your upgrade budget for it, the past two weeks produced the clearest signal yet that the wait has no payoff date. Before deciding anything, check what your target models actually need in our VRAM calculator.
The Q3 2026 window closes tomorrow, and nothing shipped
Here’s the full arc, because the individual headlines only make sense in sequence:
- October 2025 — leaked lineup and pricing: RTX 5080 SUPER 24GB ($999–$1,199), RTX 5070 Ti SUPER 24GB ($749–$799), RTX 5070 SUPER 18GB ($549–$599), all built on 3GB GDDR7 modules (TweakTown leak coverage).
- January 2026 — the refresh misses its expected Q1 window. NVIDIA tells board partners the plan is postponed to Q3 2026, not canceled. We covered that stage in our June supply analysis.
- July 18, 2026 — the cards are real. At least one board partner has finished RTX 50 SUPER units in hand, and NVIDIA tells AIBs to hold the launch over 3GB GDDR7 pricing — no new date attached. Tom’s Hardware confirms the same picture: finished GPUs “stuck in limbo” because a 3GB module costs triple a 2GB one. Even the RTX 5050 9GB, a budget SKU that needed the same chips, went on hold.
- Mid-July 2026 — supply-chain reports slide the window again, to CES 2027 at the earliest.
- August 2026 — analysis pieces start writing the obituary: an indefinite hold is worse than a cancellation for board partners, who can’t reallocate cooler inventory or PCB tooling while they wait for a green light that may never come.
- September 30, 2026 — tomorrow — the Q3 window NVIDIA gave its partners in January expires with zero SUPER cards shipped.
NVIDIA has never officially announced the SUPER series, so there is nothing to officially cancel. But “finished cards in warehouses, no ship date, next rumored window is a trade show four months out” is not a launch trajectory. It’s a holding pattern.
Why NVIDIA won’t ship a finished card
The math is brutally simple. The SUPER cards get their extra VRAM from 3GB GDDR7 modules instead of 2GB ones — same number of chips, same bus width, 50% more capacity per chip. As of the July supply-chain reports, a 3GB GDDR7 module costs $60–$70 against roughly $20 for the 2GB part.
A 5070 Ti SUPER on a 256-bit bus carries eight modules. Eight 3GB chips at $60–$70 is $480–$560 of memory on a card that leaked at a $749–$799 MSRP. The current 16GB card’s eight 2GB chips cost about $160. The VRAM upgrade alone adds roughly $320–$400 to the bill of materials — before the die, board, cooler, or anyone’s margin. Those modules are scarce because the same fab capacity is being soaked up by HBM and server DRAM for AI accelerators, where the margins are multiples higher. And with no new AMD flagship expected in 2026, NVIDIA faces no competitive pressure to eat the difference.
None of this is improving on the timescale a buyer cares about. TrendForce’s Q3 2026 outlook has server DRAM contract prices still rising 13–18% quarter-over-quarter, with increases expected to continue into 2027 and a server DRAM shortage already anticipated for next year. We’ve tracked what that squeeze did to system RAM and SSD prices and to the RTX 5090’s street price all year. The SUPER series is the same story from the supply side: the memory it needs is worth more in a server than in your tower.
September’s twist: Micron just walked away from the current cards’ memory too
While everyone watched the 3GB modules, the 2GB ones got quietly worse. On September 22, UNIKO’s Hardware spotted that Micron’s parts catalog now lists both of its 2GB GDDR7 chips — the 28Gbps MT68A512M32DF-28:A and the 32Gbps MT68A512M32DF-32:A — as End of Life, confirmed over the following days by TrendForce and other outlets. Micron is exiting the exact memory that the existing RTX 50 lineup is built on, leaving Samsung and SK hynix as the only suppliers.
Micron was never the primary GDDR7 source, so this doesn’t crater supply overnight. The direction is what matters: a third of the supplier base just decided consumer graphics memory isn’t worth the fab time. That is the opposite of the setup you’d want if you’re betting prices on current 16GB cards will drift down while you wait for the 24GB refresh.
What the SUPER would have fixed — and what it wouldn’t have
For local AI, the SUPER refresh was never about speed. The leaked 5070 Ti SUPER runs its 24GB at 28Gbps on the same 256-bit bus — that’s 896 GB/s, identical to the current 16GB card. Decode speed on a fully-loaded GPU is bandwidth-bound, so a 5070 Ti SUPER would generate tokens at the same rate the 5070 Ti does today. What you were waiting for was capacity:
| Model | Weights (Q4-class) | Fits 16GB? | Fits 24GB? |
|---|---|---|---|
| gpt-oss-20b (MXFP4) | 13.3GB | Yes — 111–189 tok/s on 5070 Ti (llama.cpp #15396) | Yes |
| Qwen3.8-27B Q4_K_M | 16.8GB | No | Yes — ~41 tok/s on a 3090 (our 27B guide) |
| Qwen3.6-35B-A3B Q4 (MoE) | ~22GB | No | Yes — 107 tok/s on a 3090 (tier breakdown) |
That table is the whole argument. The 16GB ceiling cuts you off exactly where current-generation open models get good, and the SUPER was going to move the ceiling for $749. Waiting for it was rational for most of 2026. It stops being rational when the card has no date, the memory market it depends on is still inflating, and the used 24GB cards you’d buy instead climbed 11.3% in 90 days (ResalePrices tracking) while you held out.
The 16GB wall, in practice
If you’re on a 16GB card today wondering what the fuss is about, try pulling a 27B-class model. Ollama won’t refuse — it silently splits the model between GPU and CPU, and your throughput falls off a cliff. The check:
$ ollama run qwen3.8:27b "test" --verbose
$ ollama ps
NAME ID SIZE PROCESSOR UNTIL
qwen3.8:27b a1b2c3d4e5f6 19 GB 23%/77% CPU/GPU 4 minutes from now
Anything other than 100% GPU in that PROCESSOR column means part of every forward pass is running from system RAM at a tenth of the bandwidth, and decode speed collapses with it — the failure mode we broke down in the shared-GPU-memory fallback guide. Your real options on 16GB are a lower quant with measurable quality loss, the Blackwell-only NVFP4 path that squeezes a 27B into ~14GB with a cramped context budget, or hardware with more VRAM. The SUPER was supposed to be option three. It isn’t coming on any schedule you can plan around.
What to buy instead
Every card here is measured against the phantom $749 24GB SUPER, because that’s the mental anchor everyone waiting has. Prices are September 2026 street, sources in the table and Sources below.
Used RTX 3090 24GB, ~$1,150–$1,350. Still the default answer. 936 GB/s of bandwidth — 4% more than the 5070 Ti SUPER would have had — running ~41 tok/s on Qwen3.8-27B Q4_K_M and 107 tok/s on the 35B-A3B MoE. Yes, it costs $400–$600 more than the SUPER’s leaked MSRP. The SUPER’s leaked MSRP is fiction (more on that below). Full analysis in the value-king piece and the head-to-head against the 5070 Ti.
Used RX 7900 XTX 24GB, ~$900–$1,130. The cheapest 24GB you can buy, with 960 GB/s of bandwidth and 167–177 tok/s on 7B Q4_0 under llama.cpp’s Vulkan backend. It gives up 10–25% on dense 27B decode versus a 3090 and you inherit the ROCm/Vulkan software tax — the honest trade-offs are in our used-XTX breakdown. If your stack is llama.cpp or LM Studio and you run Linux, this is the closest thing to the SUPER’s price point that exists.
RTX 5070 Ti 16GB, ~$1,000–$1,200. Only if your models genuinely fit in 16GB — coding models like gpt-oss-20b, 8–14B assistants, SDXL/Flux image work — and you value the warranty and native FP4. It matches the 3090’s decode speed inside 16GB and beats it decisively on prompt processing. It is not a “wait it out” substitute for 24GB.
Used RTX 4090 24GB, ~$2,150–$2,350. The same 24GB ceiling as the 3090 at roughly 40% more decode speed (225 vs 161 tok/s on gpt-oss-20b, llama.cpp #15396) and far stronger prefill. Buy it for throughput, not capacity — the worth-it math is here.
Skip the RTX 5080 16GB at $1,517–$1,799. It’s 24GB money for a 16GB card. The 5080 SUPER was the SKU that fixed it; see the ceiling problem in our 5070 Ti vs 5080 comparison.
Will the SUPER ever ship — and should you care if it does?
The current rumor picture says CES 2027 at the earliest, and the same reports note the refresh could be dropped entirely, with NVIDIA’s engineering attention already on the RTX 60 series for late 2027–2028. Assume a card you can’t buy for at least another quarter, probably longer.
Then apply the only lesson 2026 has taught consistently: MSRP is not the price you pay. The current 5070 Ti carries a $749 MSRP and streets at $1,000–$1,200. The 5080 lists at $999 and streets at $1,517–$1,799. A 5070 Ti SUPER launched into a worse memory market than either of those would clear well north of $1,000 — right where a used 3090 sits today, except the 3090 is available now and will have been generating tokens for you for six months by the time the SUPER’s first restock sells out. If the SUPER ships at a sane price, the used 24GB card you bought today resells; the reverse insurance doesn’t exist.
For AI-coding workloads specifically, our sister site covers running local backends in Cursor and Cline — the use case where the 27B-class models that need 24GB earn their keep. For the serving stack itself, see aifoss.dev’s Ollama review.
What to actually buy
Prices as of September 2026, all taken from the comparison above:
| Your situation | The card | Price | Where |
|---|---|---|---|
| Want 24GB for the least money, run Linux/llama.cpp | Used RX 7900 XTX 24GB | ~$900–$1,130 | Check price |
| Want 24GB with the CUDA ecosystem — the default pick | Used RTX 3090 24GB | ~$1,150–$1,350 | Check price |
| Models fit 16GB; want new, warranty, FP4 | RTX 5070 Ti 16GB | ~$1,000–$1,200 | Check price |
| Want 24GB plus double the decode speed | Used RTX 4090 24GB | ~$2,150–$2,350 | Check price |
| Undecided — test your workload on 24GB first | Rented 3090, from $0.07/hr | pay per hour | Vast.ai |
FAQ
Is the RTX 50 SUPER officially canceled? No — and it was never officially announced, so there’s nothing to cancel. The verifiable facts: finished cards reached board partners by mid-July 2026, NVIDIA told partners to hold with no new date, the Q3 2026 internal window expired, and the latest supply-chain rumor says CES 2027 at the earliest.
Would the RTX 5070 Ti SUPER really have cost $749? $749–$799 is the leaked MSRP from October 2025 — before 3GB GDDR7 settled at $60–$70 per module. Eight of those modules alone approach $560. Every recent Blackwell card streets 33–60% above MSRP; there’s no reason a memory-heavy SKU launched into a worse market would behave better.
Should I wait for the RTX 60 series instead? That’s a late-2027-to-2028 event on current reporting, and TrendForce expects memory prices to keep rising into 2027. Eighteen-plus months of not running the models you want is the real cost of that plan.
Is 16GB enough for local AI in late 2026? For 8–14B models, gpt-oss-20b, and image generation — comfortably. The wall is the 27B-class dense models (16.8GB at Q4_K_M) and 30B+ MoE models, which is exactly the tier where open models now compete with cloud offerings. See the 16GB tier guide for the full menu.
What about just buying an RTX 5090? It’s the fastest consumer card and its 32GB clears the 27B tier easily, but at $3,822–$5,000 street it costs more than three used 3090s. That’s a different conversation — the 3K/6K/10K build tiers walk through when it makes sense.
Recommended Gear
- Used RTX 3090 24GB — the default 24GB pick, ~$1,150–$1,350
- Used RX 7900 XTX 24GB — cheapest 24GB, ~$900–$1,130
- RTX 5070 Ti 16GB — best new card under $1,200 if 16GB fits
- Used RTX 4090 24GB — 24GB at twice the speed, ~$2,150–$2,350
Sources
- NVIDIA’s RTX 50 Super GPUs have reached board partners, but launch is on hold over 3GB GDDR7 pricing — TweakTown, Jul 18 2026
- Nvidia RTX 50 Super GPUs are reportedly ready, but stuck in limbo due to excessive GDDR7 pricing — Tom’s Hardware
- Micron Reportedly Ends 2GB GDDR7, Narrowing Supply Options for NVIDIA’s RTX 50 Series — TrendForce, Sep 24 2026
- Micron Has Reportedly Ceased Production of 2GB GDDR7 Modules — The FPS Review, Sep 23 2026
- NVIDIA told AIB partners it has NOT canceled RTX 50 SUPER plans, but postponed until Q3 2026 — TweakTown
- GeForce RTX 50 SUPER Series is now rumored to be coming in early 2027 — TweakTown
- Full leaked details on RTX 5080 SUPER, RTX 5070 Ti SUPER, RTX 5070 SUPER — TweakTown, Oct 2025
- Why Nvidia abandoned the RTX 50 Super before it even launched — How-To Geek, Aug 2026
- Server DRAM Contract Prices Expected to Rise 13–18% QoQ in 3Q26 — TrendForce, Jul 9 2026
- RTX 3090 Used GPU Price & Fair Asking Range — ResalePrices
- RTX 3090 Price History, September 2026 — BestValueGPU
- RTX 5070 Ti Price Tracker, September 2026 — videocardprices.com
- Guide: running gpt-oss with llama.cpp — ggml-org discussion #15396
Last updated September 29, 2026. GPU and memory prices are moving weekly in the current supply crunch; verify current listings before purchasing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.