Used RTX 2080 Ti for Local AI in 2026: $250 for 616 GB/s — and the $500 Mod That Doubles Its VRAM
TL;DR: A used RTX 2080 Ti averages about $250 in August–September 2026 — and its 616 GB/s of bandwidth beats a $509 used RTX 4070 (504 GB/s). The catch is the odd 11GB buffer, one gigabyte short of where the 12GB tier starts to breathe. The genuinely interesting version is the ~$500 pre-modded 22GB card: 27B models at 37 tok/s for less than half the price of a used RTX 3090.
| Used RTX 2080 Ti 11GB (~$250) | Modded 2080 Ti 22GB (~$499–$529) | Used RTX 3090 24GB (~$1,286) | |
|---|---|---|---|
| Best for | Fast 7B–14B on a strict budget | 27B-class models at the lowest buy-in | The clean 24GB answer, 936 GB/s |
| Bandwidth / VRAM | 616 GB/s / 11GB | 616 GB/s / 22GB | 936 GB/s / 24GB |
| The catch | 11GB is the wrong number in 2026 | Grey-market mod, no warranty, 2018 silicon | 2.6× the price of the modded card |
Honest take: At $250 the stock card is a legitimate speed bargain but a capacity dead end — a $286 RTX 3060 12GB is slower yet fits more. If you’re tempted at all, be tempted by the 22GB mod at ~$500: it’s the cheapest ticket to the 27B class that exists, provided you accept that a resistor-strapped 2018 flagship comes with zero warranty and an unknown fan history.
Why a 2018 flagship is back in the conversation
The RTX 2080 Ti launched September 20, 2018 at $999 ($1,199 for the Founders Edition) as the first consumer flagship of the RTX era. Eight years later, BestValueGPU’s August 2026 tracker puts the used price at about $250, with live eBay listings running from $265 for a clean PNY card to $425 for premium AIB models like the AORUS Xtreme.
Two things dragged it back into local-AI relevance this year. First, the DRAM crisis pushed every modern card up a tier in price — a new RTX 5070 runs ~$871–$900 and even the used market inflated — while the 2080 Ti, too old to game fashionably, mostly didn’t move. Second, Hong Kong and UAE sellers began listing pre-modded 22GB versions on eBay at $499–$529, and a Japanese repair shop will convert your own card for about $282 by swapping the eleven 1GB GDDR6 modules for 2GB chips and adjusting the strap resistors. Double the VRAM on a card whose 616 GB/s bus still out-runs most of the modern midrange.
That bandwidth number deserves a second look. Token generation is memory-bound, and 616 GB/s is more than a used RTX 4070 ($509, 504 GB/s), more than a used RTX 3070 ($228, 448 GB/s), and 1.7× a used RTX 3060 12GB ($286, 360 GB/s). Among sub-$300 cards, nothing else moves bytes this fast except the 10GB RTX 3080. Before you buy anything in this bracket, sanity-check your target model and context against our VRAM calculator — at 11GB, the margins are thin enough that the calculator earns its keep.
One more point the spreadsheet doesn’t show: unlike the $260 Tesla P40, the 2080 Ti has no driver expiration date looming. NVIDIA’s 580-branch cutoff moves Maxwell, Pascal, and Volta to legacy status — Turing (the RTX 20 series) explicitly survives it and keeps riding the mainline driver. A P40 is 24GB with a mid-2028 clock ticking; a 2080 Ti is a smaller buffer with no deadline.
What a used RTX 2080 Ti actually delivers
Specs that matter for inference: TU102 with 4,352 CUDA cores, 11GB GDDR6 on a 352-bit bus at 616 GB/s, 250W TDP, two 8-pin connectors on most boards. Turing has first-generation tensor cores — FP16 and INT8 are accelerated, but BF16 arrived with Ampere, which occasionally matters for training and fine-tuning stacks. For GGUF inference it’s business as usual: llama.cpp and Ollama treat compute capability 7.5 as a fully supported CUDA citizen.
Measured numbers from LocalScore’s community benchmark database for the RTX 2080 Ti:
| Model | Quant class | VRAM fit on 11GB | Speed |
|---|---|---|---|
| Llama 3.2 1B | Q4 | Trivial | 222 tok/s |
| Llama 3.1 8B | Q4 | ~5–6GB, comfortable | 72.9 tok/s |
| Qwen2.5 14B | Q4 | ~9GB — fits, tight | 39.5 tok/s |
| Qwen3.8-27B (22GB mod) | Q4_K | ~17GB, needs the mod | 37.1 tok/s |
That 14B number is the tell: 39.5 tok/s is faster than the same class on a used RTX 4070 (~33–36 tok/s) and roughly 40% ahead of the RTX 3060 12GB (25–32 tok/s). The 616 GB/s bus does exactly what the spec sheet promises. Where the card gives ground is prompt processing — 4,352 Turing cores at 2018 clocks trail Ada and Ampere on prefill, so giant system prompts and RAG contexts take noticeably longer to chew through than on a 4070, even though generation is quicker once it starts. Speed like that is only useful while your models still fit in 11GB — the GPU buyer’s guide covers what to buy for when they stop.
The 27B row comes from ai-muninn’s test of the modded 22GB cards: a single card runs Qwen3.8-27B at Q4_K at 37.1 tok/s, and two of them under llama.cpp’s tensor-split mode reach 59.6 tok/s. For reference, our 24GB tier guide puts a used RTX 3090 in the mid-40s on the same class — the modded Turing card lands within striking distance at 40% of the price.
What you cannot run — the honest section
On the stock 11GB card, the ceiling is lower than every number above makes it feel:
- Qwen3 14B at long context — the weights fit at ~9GB, but 16K context pushes total VRAM past ~11.5GB. On a 12GB card that squeaks by; on 11GB it spills. This tier’s daily driver is really “14B at 8K context.”
- GPT-OSS 20B — 12.8GB of MXFP4 weights before KV cache. Doesn’t fit, full stop.
- Gemma 4 26B-A4B QAT — ~15GB. Needs 16GB minimum.
- Qwen3.6-27B — 16.8GB at Q4_K_M. This is the model that makes 24GB (or the 22GB mod) worth having.
- Dense 70B — 43GB at Q4_K_M; see running a 70B on a 24GB card for why even 24GB is a compromise.
The 22GB mod rewrites most of that list: Qwen3.6/3.8-27B fits with room for real context, Gemma 4 31B QAT (~18GB) fits, and the 20B class becomes comfortable. What it doesn’t buy you is the 35B-A3B MoE sweet spot at speed (fits at ~21GB but with almost no KV headroom) or anything in the 48GB conversation — for that, two modded cards or a dual-3090 build are the next rungs. Worth knowing: the 2080 Ti was the first consumer card with NVLink (2-way, up to 100 GB/s over a $79 bridge), but llama.cpp’s tensor-split runs over plain PCIe anyway — the ai-muninn dual-card result needed no bridge.
The trap: 11GB looks like 12GB until the KV cache arrives
The most common way this card disappoints: you pull a 14B, it benchmarks at ~39 tok/s, and a week later a longer chat session is crawling at 6 tok/s. Nothing broke — the KV cache grew past the last gigabyte, and Ollama silently started splitting the model between GPU and system RAM. On this card the check is:
$ ollama ps
NAME ID SIZE PROCESSOR UNTIL
qwen3:14b a3f... 11 GB 12%/88% CPU/GPU 4 minutes from now
Anything other than 100% GPU in the PROCESSOR column means the 616 GB/s bus is now waiting on your DDR4. The fix, in order:
- Cap context to what you use:
/set parameter num_ctx 8192, or bake it into a Modelfile. - Set
OLLAMA_FLASH_ATTENTION=1andOLLAMA_KV_CACHE_TYPE=q8_0on the service (systemd drop-in or Windows service environment, not your shell) — quantizing the KV cache roughly halves its footprint, which is precisely the gigabyte this card is missing. - Close Chrome; hardware acceleration routinely holds 1–2GB of VRAM on Windows.
Re-run ollama ps, confirm 100% GPU, and the mid-30s come back. Deeper diagnosis in Ollama not using your GPU.
The modded card has its own fine print, and it’s less about software: these are eight-year-old boards resold through grey-market channels after board-level rework. The 2GB Micron modules and strap changes present to the driver as a native 22GB card — no hacked BIOS, standard drivers, coverage from TechPowerUp and Tom’s Hardware confirms they behave normally — but there is no warranty, no RMA path, and no way to know whether the fans spent 2019–2022 in a mining rack. Budget for a $25–$40 fan kit or repaste, buy from a seller with a real return window, and stress-test the full 22GB on day one (llama-bench with a model that actually fills it, then watch nvidia-smi for ECC-less memory errors showing up as gibberish output).
Used 2080 Ti vs the alternatives, September 2026 prices
| Card | Street price | Bandwidth | VRAM | Best fit | Verdict |
|---|---|---|---|---|---|
| Used RTX 3070 | $228 | 448 GB/s | 8GB | 8B class | Cheaper, but 8GB stings |
| Used RTX 2080 Ti | ~$250 | 616 GB/s | 11GB | Fast 8B–14B @ 8K ctx | Speed bargain, capacity dead end |
| Used RTX 3080 | $285 | 760 GB/s | 10GB | Fastest sub-$300 | More speed, even less room |
| Used RTX 3060 12GB | $286 | 360 GB/s | 12GB | 14B with headroom | The sane default |
| Modded 2080 Ti 22GB | $499–$529 | 616 GB/s | 22GB | 27B class at 37 tok/s | The value play, eyes open |
| Used RTX 3090 | $1,286 | 936 GB/s | 24GB | 27B–35B, no asterisks | The clean answer |
Three of those match-ups deserve a sentence:
vs the RTX 3060 12GB at $286: the 3060 is the boring, correct pick for most people — one more gigabyte exactly where this tier needs it, half the power draw, modern warranty-era stock, and our full 3060 write-up stands. The 2080 Ti earns its place only if 14B-at-8K feeling fast (39 vs 27 tok/s) matters more to you than 14B-at-16K fitting at all.
vs the RTX 3080 at $285: the $285 3080 is 23% faster on paper and loses a gigabyte. Both are speed-over-capacity picks; the 3080 does it with newer silicon. If you’re choosing between the two stock cards, take the 3080 — the 2080 Ti’s real argument was never the 11GB version.
Modded 22GB vs used 3090: $500 vs $1,286 for 22GB vs 24GB. The 3090 is 52% faster, newer, warrantied by volume (343 listings, liquid market), and the safe recommendation we’ve made since 2026 began. The modded card exists for exactly one buyer: you want the 27B class, $1,300 is genuinely out of reach, and you’re comfortable owning grey-market hardware. That buyer gets 90% of the capacity for 39% of the money — a trade nothing else on this list offers.
Power and running cost
At its 250W TDP under sustained load, a 2080 Ti costs about $0.046/hour at the September 2026 US residential average of 18.44¢/kWh — call it $1.10 for a 24-hour batch day, pennies for interactive use. You’ll want a quality 650W PSU with two 8-pins, and note that Turing-era boards predate the transient-spike drama of later generations. If it runs around the clock, power-limiting to ~180W costs a few percent of tok/s and keeps 2018-vintage VRM and fans well inside their comfort zone — cheap insurance on silicon this old. Full 24/7 math in what a home AI server actually costs.
Buy or skip: the price thresholds
Buy the stock 11GB card at $230–$270 only if:
- You want the fastest possible 8B (72.9 tok/s) and 14B-at-8K experience under $300, and you’ll actually set the KV-cache flags above
- You already own one and are deciding whether to sell — at these prices, keeping it as a dedicated Ollama endpoint beats the ~$220 you’d clear
Buy the modded 22GB card at $480–$530 if:
- The 27B class is your target and a $1,286 RTX 3090 is out of budget — 37 tok/s on Qwen3.8-27B is a real daily-driver speed
- You’re comfortable stress-testing on arrival and eating the loss if the board dies in year two
Skip both if:
- You want the sane budget pick → used RTX 3060 12GB at $286, warranty-era silicon, same model list as the stock 2080 Ti minus the speed. Our 12GB tier guide covers what it runs.
- You can stretch to $1,286 → the used 3090 deletes every asterisk in this article: more VRAM, 52% more bandwidth, no mod risk, liquid resale market.
- You just want to try the 27B class before committing → rent a 24GB pod on RunPod for $0.30–$0.70/hr first; the rent-vs-buy math says hardware only wins at daily usage.
Software support is a non-issue either way: Ollama v0.32.x and llama.cpp treat Turing as standard CUDA, and if you’re wiring the card into a local coding stack, Continue.dev or Cline against an Ollama endpoint runs the same on compute 7.5 as on anything newer.
FAQ
Is a used RTX 2080 Ti good for local AI in 2026? Within limits. Its 616 GB/s bandwidth delivers 72.9 tok/s on 8B models and 39.5 tok/s on 14B Q4 — faster than a used RTX 4070 on generation. The 11GB buffer is the problem: 14B models only fit at short context, and nothing in the 20B+ class fits at all without the 22GB mod.
What is the 22GB RTX 2080 Ti mod, and is it safe? Repair shops replace the eleven 1GB GDDR6 modules with 2GB chips and adjust strap resistors; the card then reports 22GB to standard drivers with no BIOS hacks. Pre-modded cards sell for $499–$529 on eBay (Hong Kong and UAE sellers), or a Japanese shop converts your card for ~$282. They test as fully functional, but there’s no warranty — treat it as enthusiast hardware, not an appliance.
Can the modded 22GB 2080 Ti run 27B models? Yes — that’s its entire reason to exist. Qwen3.8-27B at Q4_K runs at 37.1 tok/s on a single modded card, and two cards under llama.cpp tensor-split hit 59.6 tok/s. A used RTX 3090 is roughly 50% faster on the same class but costs 2.6× as much.
Will NVIDIA drop driver support for the RTX 2080 Ti soon? No announced deadline. The 580-branch cutoff that moves cards to legacy status covers Maxwell, Pascal, and Volta — Turing survives it and stays on the mainline driver, unlike the Tesla P40, which is why the 2080 Ti doesn’t carry the P40’s mid-2028 expiration risk.
Used RTX 2080 Ti vs RTX 3060 12GB — which should I buy? The 3060 for most people: 12GB beats 11GB exactly where this tier is tightest, it draws 170W to the 2080 Ti’s 250W, and the used stock is younger. The 2080 Ti wins only on raw speed (39.5 vs 25–32 tok/s on 14B). If speed is really the priority, a $285 RTX 3080 beats both.
Recommended Gear
- RTX 2080 Ti (used) — ~$250 for 616 GB/s; check seller history, assume it needs a repaste
- RTX 3060 12GB (used) — the sane $286 alternative with the extra gigabyte
- RTX 3090 24GB (used) — $1,286 average; the no-asterisk 24GB answer
Sources
- RTX 2080 Ti Price Tracker US, August 2026 — BestValueGPU
- 22 GB Modded GeForce RTX 2080 Ti Cards Listed on eBay, $499 per Unit — TechPowerUp
- Pre-Modded 22GB RTX 2080 Ti Cards Surface on eBay for $500 — Tom’s Hardware
- Japanese Repair Shop Sells GPU VRAM Upgrades: RTX 2080 Ti to 22GB for $282 — Tom’s Hardware
- Modded NVIDIA RTX 2080 Ti with 22GB VRAM Is Selling on eBay for $500 — TechSpot
- NVIDIA GeForce RTX 2080 Ti Benchmark Results — LocalScore
- Two Modded 2080 Tis Reach 59.6 tok/s on Qwen3.8-27B with llama.cpp Tensor Parallel — ai-muninn
- NVIDIA’s v580 Driver Branch Ends Support for Maxwell, Pascal, and Volta GPUs — TechPowerUp
- NVIDIA to Axe Maxwell, Pascal, and Volta with End of Driver Support — Tom’s Hardware
- Official NVIDIA RTX 2080 Ti Specs, Price, Release Date — GamersNexus
- NVIDIA GeForce RTX 2080 Ti Graphics Card Specs — VideoCardz
- NVIDIA GeForce RTX 2080 Ti NVLink SLI Scaling Explored — HotHardware
- Electricity Rates by State, September 2026 — ElectricChoice
Last updated September 3, 2026. Prices and specs change; verify current rates before purchasing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →