Used RTX 2080 Ti for Local AI in 2026: $250 for 616 GB/s — and the $500 Mod That Doubles Its VRAM

rtx-2080-tiused-gpugpulocal-llmbuying-guidevram-mod

TL;DR: A used RTX 2080 Ti averages about $250 in August–September 2026 — and its 616 GB/s of bandwidth beats a $509 used RTX 4070 (504 GB/s). The catch is the odd 11GB buffer, one gigabyte short of where the 12GB tier starts to breathe. The genuinely interesting version is the ~$500 pre-modded 22GB card: 27B models at 37 tok/s for less than half the price of a used RTX 3090.

Used RTX 2080 Ti 11GB (~$250)Modded 2080 Ti 22GB (~$499–$529)Used RTX 3090 24GB (~$1,286)
Best forFast 7B–14B on a strict budget27B-class models at the lowest buy-inThe clean 24GB answer, 936 GB/s
Bandwidth / VRAM616 GB/s / 11GB616 GB/s / 22GB936 GB/s / 24GB
The catch11GB is the wrong number in 2026Grey-market mod, no warranty, 2018 silicon2.6× the price of the modded card

Honest take: At $250 the stock card is a legitimate speed bargain but a capacity dead end — a $286 RTX 3060 12GB is slower yet fits more. If you’re tempted at all, be tempted by the 22GB mod at ~$500: it’s the cheapest ticket to the 27B class that exists, provided you accept that a resistor-strapped 2018 flagship comes with zero warranty and an unknown fan history.

Why a 2018 flagship is back in the conversation

The RTX 2080 Ti launched September 20, 2018 at $999 ($1,199 for the Founders Edition) as the first consumer flagship of the RTX era. Eight years later, BestValueGPU’s August 2026 tracker puts the used price at about $250, with live eBay listings running from $265 for a clean PNY card to $425 for premium AIB models like the AORUS Xtreme.

Two things dragged it back into local-AI relevance this year. First, the DRAM crisis pushed every modern card up a tier in price — a new RTX 5070 runs ~$871–$900 and even the used market inflated — while the 2080 Ti, too old to game fashionably, mostly didn’t move. Second, Hong Kong and UAE sellers began listing pre-modded 22GB versions on eBay at $499–$529, and a Japanese repair shop will convert your own card for about $282 by swapping the eleven 1GB GDDR6 modules for 2GB chips and adjusting the strap resistors. Double the VRAM on a card whose 616 GB/s bus still out-runs most of the modern midrange.

That bandwidth number deserves a second look. Token generation is memory-bound, and 616 GB/s is more than a used RTX 4070 ($509, 504 GB/s), more than a used RTX 3070 ($228, 448 GB/s), and 1.7× a used RTX 3060 12GB ($286, 360 GB/s). Among sub-$300 cards, nothing else moves bytes this fast except the 10GB RTX 3080. Before you buy anything in this bracket, sanity-check your target model and context against our VRAM calculator — at 11GB, the margins are thin enough that the calculator earns its keep.

One more point the spreadsheet doesn’t show: unlike the $260 Tesla P40, the 2080 Ti has no driver expiration date looming. NVIDIA’s 580-branch cutoff moves Maxwell, Pascal, and Volta to legacy status — Turing (the RTX 20 series) explicitly survives it and keeps riding the mainline driver. A P40 is 24GB with a mid-2028 clock ticking; a 2080 Ti is a smaller buffer with no deadline.

What a used RTX 2080 Ti actually delivers

Specs that matter for inference: TU102 with 4,352 CUDA cores, 11GB GDDR6 on a 352-bit bus at 616 GB/s, 250W TDP, two 8-pin connectors on most boards. Turing has first-generation tensor cores — FP16 and INT8 are accelerated, but BF16 arrived with Ampere, which occasionally matters for training and fine-tuning stacks. For GGUF inference it’s business as usual: llama.cpp and Ollama treat compute capability 7.5 as a fully supported CUDA citizen.

Measured numbers from LocalScore’s community benchmark database for the RTX 2080 Ti:

ModelQuant classVRAM fit on 11GBSpeed
Llama 3.2 1BQ4Trivial222 tok/s
Llama 3.1 8BQ4~5–6GB, comfortable72.9 tok/s
Qwen2.5 14BQ4~9GB — fits, tight39.5 tok/s
Qwen3.8-27B (22GB mod)Q4_K~17GB, needs the mod37.1 tok/s

That 14B number is the tell: 39.5 tok/s is faster than the same class on a used RTX 4070 (~33–36 tok/s) and roughly 40% ahead of the RTX 3060 12GB (25–32 tok/s). The 616 GB/s bus does exactly what the spec sheet promises. Where the card gives ground is prompt processing — 4,352 Turing cores at 2018 clocks trail Ada and Ampere on prefill, so giant system prompts and RAG contexts take noticeably longer to chew through than on a 4070, even though generation is quicker once it starts. Speed like that is only useful while your models still fit in 11GB — the GPU buyer’s guide covers what to buy for when they stop.

The 27B row comes from ai-muninn’s test of the modded 22GB cards: a single card runs Qwen3.8-27B at Q4_K at 37.1 tok/s, and two of them under llama.cpp’s tensor-split mode reach 59.6 tok/s. For reference, our 24GB tier guide puts a used RTX 3090 in the mid-40s on the same class — the modded Turing card lands within striking distance at 40% of the price.

What you cannot run — the honest section

On the stock 11GB card, the ceiling is lower than every number above makes it feel:

  • Qwen3 14B at long context — the weights fit at ~9GB, but 16K context pushes total VRAM past ~11.5GB. On a 12GB card that squeaks by; on 11GB it spills. This tier’s daily driver is really “14B at 8K context.”
  • GPT-OSS 20B — 12.8GB of MXFP4 weights before KV cache. Doesn’t fit, full stop.
  • Gemma 4 26B-A4B QAT — ~15GB. Needs 16GB minimum.
  • Qwen3.6-27B — 16.8GB at Q4_K_M. This is the model that makes 24GB (or the 22GB mod) worth having.
  • Dense 70B — 43GB at Q4_K_M; see running a 70B on a 24GB card for why even 24GB is a compromise.

The 22GB mod rewrites most of that list: Qwen3.6/3.8-27B fits with room for real context, Gemma 4 31B QAT (~18GB) fits, and the 20B class becomes comfortable. What it doesn’t buy you is the 35B-A3B MoE sweet spot at speed (fits at ~21GB but with almost no KV headroom) or anything in the 48GB conversation — for that, two modded cards or a dual-3090 build are the next rungs. Worth knowing: the 2080 Ti was the first consumer card with NVLink (2-way, up to 100 GB/s over a $79 bridge), but llama.cpp’s tensor-split runs over plain PCIe anyway — the ai-muninn dual-card result needed no bridge.

The trap: 11GB looks like 12GB until the KV cache arrives

The most common way this card disappoints: you pull a 14B, it benchmarks at ~39 tok/s, and a week later a longer chat session is crawling at 6 tok/s. Nothing broke — the KV cache grew past the last gigabyte, and Ollama silently started splitting the model between GPU and system RAM. On this card the check is:

$ ollama ps
NAME            ID          SIZE     PROCESSOR          UNTIL
qwen3:14b       a3f...      11 GB    12%/88% CPU/GPU    4 minutes from now

Anything other than 100% GPU in the PROCESSOR column means the 616 GB/s bus is now waiting on your DDR4. The fix, in order:

  1. Cap context to what you use: /set parameter num_ctx 8192, or bake it into a Modelfile.
  2. Set OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0 on the service (systemd drop-in or Windows service environment, not your shell) — quantizing the KV cache roughly halves its footprint, which is precisely the gigabyte this card is missing.
  3. Close Chrome; hardware acceleration routinely holds 1–2GB of VRAM on Windows.

Re-run ollama ps, confirm 100% GPU, and the mid-30s come back. Deeper diagnosis in Ollama not using your GPU.

The modded card has its own fine print, and it’s less about software: these are eight-year-old boards resold through grey-market channels after board-level rework. The 2GB Micron modules and strap changes present to the driver as a native 22GB card — no hacked BIOS, standard drivers, coverage from TechPowerUp and Tom’s Hardware confirms they behave normally — but there is no warranty, no RMA path, and no way to know whether the fans spent 2019–2022 in a mining rack. Budget for a $25–$40 fan kit or repaste, buy from a seller with a real return window, and stress-test the full 22GB on day one (llama-bench with a model that actually fills it, then watch nvidia-smi for ECC-less memory errors showing up as gibberish output).

Used 2080 Ti vs the alternatives, September 2026 prices

CardStreet priceBandwidthVRAMBest fitVerdict
Used RTX 3070$228448 GB/s8GB8B classCheaper, but 8GB stings
Used RTX 2080 Ti~$250616 GB/s11GBFast 8B–14B @ 8K ctxSpeed bargain, capacity dead end
Used RTX 3080$285760 GB/s10GBFastest sub-$300More speed, even less room
Used RTX 3060 12GB$286360 GB/s12GB14B with headroomThe sane default
Modded 2080 Ti 22GB$499–$529616 GB/s22GB27B class at 37 tok/sThe value play, eyes open
Used RTX 3090$1,286936 GB/s24GB27B–35B, no asterisksThe clean answer

Three of those match-ups deserve a sentence:

vs the RTX 3060 12GB at $286: the 3060 is the boring, correct pick for most people — one more gigabyte exactly where this tier needs it, half the power draw, modern warranty-era stock, and our full 3060 write-up stands. The 2080 Ti earns its place only if 14B-at-8K feeling fast (39 vs 27 tok/s) matters more to you than 14B-at-16K fitting at all.

vs the RTX 3080 at $285: the $285 3080 is 23% faster on paper and loses a gigabyte. Both are speed-over-capacity picks; the 3080 does it with newer silicon. If you’re choosing between the two stock cards, take the 3080 — the 2080 Ti’s real argument was never the 11GB version.

Modded 22GB vs used 3090: $500 vs $1,286 for 22GB vs 24GB. The 3090 is 52% faster, newer, warrantied by volume (343 listings, liquid market), and the safe recommendation we’ve made since 2026 began. The modded card exists for exactly one buyer: you want the 27B class, $1,300 is genuinely out of reach, and you’re comfortable owning grey-market hardware. That buyer gets 90% of the capacity for 39% of the money — a trade nothing else on this list offers.

Power and running cost

At its 250W TDP under sustained load, a 2080 Ti costs about $0.046/hour at the September 2026 US residential average of 18.44¢/kWh — call it $1.10 for a 24-hour batch day, pennies for interactive use. You’ll want a quality 650W PSU with two 8-pins, and note that Turing-era boards predate the transient-spike drama of later generations. If it runs around the clock, power-limiting to ~180W costs a few percent of tok/s and keeps 2018-vintage VRM and fans well inside their comfort zone — cheap insurance on silicon this old. Full 24/7 math in what a home AI server actually costs.

Buy or skip: the price thresholds

Buy the stock 11GB card at $230–$270 only if:

  • You want the fastest possible 8B (72.9 tok/s) and 14B-at-8K experience under $300, and you’ll actually set the KV-cache flags above
  • You already own one and are deciding whether to sell — at these prices, keeping it as a dedicated Ollama endpoint beats the ~$220 you’d clear

Buy the modded 22GB card at $480–$530 if:

  • The 27B class is your target and a $1,286 RTX 3090 is out of budget — 37 tok/s on Qwen3.8-27B is a real daily-driver speed
  • You’re comfortable stress-testing on arrival and eating the loss if the board dies in year two

Skip both if:

  • You want the sane budget pick → used RTX 3060 12GB at $286, warranty-era silicon, same model list as the stock 2080 Ti minus the speed. Our 12GB tier guide covers what it runs.
  • You can stretch to $1,286 → the used 3090 deletes every asterisk in this article: more VRAM, 52% more bandwidth, no mod risk, liquid resale market.
  • You just want to try the 27B class before committing → rent a 24GB pod on RunPod for $0.30–$0.70/hr first; the rent-vs-buy math says hardware only wins at daily usage.

Software support is a non-issue either way: Ollama v0.32.x and llama.cpp treat Turing as standard CUDA, and if you’re wiring the card into a local coding stack, Continue.dev or Cline against an Ollama endpoint runs the same on compute 7.5 as on anything newer.

FAQ

Is a used RTX 2080 Ti good for local AI in 2026? Within limits. Its 616 GB/s bandwidth delivers 72.9 tok/s on 8B models and 39.5 tok/s on 14B Q4 — faster than a used RTX 4070 on generation. The 11GB buffer is the problem: 14B models only fit at short context, and nothing in the 20B+ class fits at all without the 22GB mod.

What is the 22GB RTX 2080 Ti mod, and is it safe? Repair shops replace the eleven 1GB GDDR6 modules with 2GB chips and adjust strap resistors; the card then reports 22GB to standard drivers with no BIOS hacks. Pre-modded cards sell for $499–$529 on eBay (Hong Kong and UAE sellers), or a Japanese shop converts your card for ~$282. They test as fully functional, but there’s no warranty — treat it as enthusiast hardware, not an appliance.

Can the modded 22GB 2080 Ti run 27B models? Yes — that’s its entire reason to exist. Qwen3.8-27B at Q4_K runs at 37.1 tok/s on a single modded card, and two cards under llama.cpp tensor-split hit 59.6 tok/s. A used RTX 3090 is roughly 50% faster on the same class but costs 2.6× as much.

Will NVIDIA drop driver support for the RTX 2080 Ti soon? No announced deadline. The 580-branch cutoff that moves cards to legacy status covers Maxwell, Pascal, and Volta — Turing survives it and stays on the mainline driver, unlike the Tesla P40, which is why the 2080 Ti doesn’t carry the P40’s mid-2028 expiration risk.

Used RTX 2080 Ti vs RTX 3060 12GB — which should I buy? The 3060 for most people: 12GB beats 11GB exactly where this tier is tightest, it draws 170W to the 2080 Ti’s 250W, and the used stock is younger. The 2080 Ti wins only on raw speed (39.5 vs 25–32 tok/s on 14B). If speed is really the priority, a $285 RTX 3080 beats both.

Sources

Last updated September 3, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?