The $5,000 Local AI Workstation in 2026: Full Parts List, and Why Half the Budget Goes to Two Used GPUs

gpulocal-llmworkstationbuild-guidertx-3090hardware

TL;DR: $5,000 in September 2026 buys a dual used RTX 3090 workstation — 48GB of VRAM that runs Llama 3.3 70B Q4_K_M fully resident at 7–10 tok/s. The parts list lands at ~$4,600, and the memory supercycle means your RAM and SSD cost triple what they did in 2024. A single used RTX 4090 build is the faster-but-smaller alternative.

This build (2× used 3090)1× used RTX 4090Rented GPU
Best for70B models fully in VRAMFastest 24GB inference + diffusionTesting before buying
VRAM48GB pooled24GBWhatever you rent
Price (Sep 2026)~$4,600 all-in~$4,300 all-in3090 from $0.07/hr
The catch~700W GPU load, BIOS/PCIe setup work70B needs CPU offload at 8–14 tok/sNothing local, nothing private

Honest take: If the whole point of spending $5,000 is running models a single card can’t hold, build the dual-3090 rig. If you’d mostly run 27B–35B models fast, save $300 and buy the used 4090 instead — and if you’re not sure which person you are, rent both configurations on Vast.ai for a weekend first.

A $5,000 local AI build in September 2026 is a different exercise than it was a year ago. The GPU side has stayed almost rational — used RTX 3090s average $1,264 on ResalePrices’ tracker — while everything made of memory chips has gone vertical: a 64GB DDR5-6000 kit now costs more than a PlayStation 5, and 2TB NVMe drives that were $130 in May are $300+ now. So this parts list makes different trade-offs than the $2,000 build did in the spring: maximum dollars into VRAM, minimum dollars into DRAM and NAND, and a named alternative at every slot so you can rebalance when prices move again.

The parts list

Prices verified September 18, 2026:

PartPickPrice
GPUs2× used RTX 3090 24GB~$2,528 ($1,264 avg each)
CPUAMD Ryzen 9 9950X (16C/32T)~$485
MotherboardASUS ProArt X870E-Creator WiFi~$480
RAM32GB DDR5-6000 CL30 (2×16GB)~$399–$479
StorageSamsung 990 Pro 2TB Gen4 NVMe~$338
PSULian Li Edge 1200W ATX 3.1 Platinum~$179
CaseLian Li Lancool 216~$103
CPU coolerThermalright Peerless Assassin 120 SE~$50
Total~$4,562–$4,642

That leaves roughly $350–$450 of headroom against the $5,000 ceiling, and the last section covers the three competing ways to spend it: a used NVLink bridge ($100), a RAM bump to 64GB ($430 more), or nothing — because the used 3090 market has been drifting upward and your two cards may simply cost more than the tracker average by the time you buy.

GPUs: 2× used RTX 3090 — the whole build exists for this line

Half the budget goes here, on purpose. 48GB of pooled VRAM is what separates this machine from every single-consumer-card build: Llama 3.3 70B at Q4_K_M is a 42.5GB file, and it loads fully resident with each 3090 holding about 21GB (InsiderLLM’s quant ladder has the full table). Expect 7–10 tok/s on that 70B with layer split — above reading speed, fine for chat and batch work, not agent-fast. Models that fit one card run at full single-card speed: ~95 tok/s on a 7B, 161 tok/s on gpt-oss-20b (llama.cpp benchmark thread), ~40 tok/s on a 27B at Q4_K_M.

The per-card math is why it’s this card and not something newer. ResalePrices puts the used 3090 at a $1,264 average ($1,201–$1,299 fair range); September tracker spreads run $1,050–$1,343 depending on cooler model and seller type. That’s 936 GB/s of memory bandwidth per card — more than any other GPU you can buy at the price — and 3090s are the last consumer cards with NVLink support if you later want faster cross-card transfers.

Buy used like it’s a used car: sellers with return windows, and inspect the power connector for browning before first boot. The 48GB VRAM tier guide covers what this configuration runs model-by-model.

Alternatives. A single used RTX 4090 (~$2,527, 30-day average) is the same money for half the VRAM and much faster single-card decode — 225 tok/s on gpt-oss-20b — plus meaningfully better Stable Diffusion and fine-tuning performance. It’s the better pick if 70B models are an occasional curiosity rather than the goal; our used 4090 vs 5090 breakdown has the full price picture. The sleeper option is a used RTX A6000: 48GB on one 300W card, $2,600–$3,800 used, ~18% slower per token than a 3090 but with zero multi-GPU configuration work. If your used-market search turns one up near the $2,600 floor, it simplifies everything downstream of this line — smaller PSU, any motherboard, no split-mode flags.

CPU: Ryzen 9 9950X — bought for its PCIe lanes, not its cores

The 9950X (~$485, verified in our Threadripper comparison this week) is here for one reason: AM5’s 24 usable PCIe 5.0 lanes run two GPUs at x8/x8, and for inference, x8 per card costs you almost nothing — decode is bound by each card’s own memory bandwidth, not the slot. The 16 cores are a bonus for the moments the CPU actually matters: prompt preprocessing, tokenization, and partial offload of MoE experts if you ever run something bigger than 48GB.

The Threadripper conversation starts at GPU number three. If you think you’ll genuinely expand past two cards, that’s a different (and much more expensive) platform — the TRX50 math is here — but don’t pay the Threadripper platform tax for a two-card build.

Alternative. Any AM5 Ryzen with the same lane layout works; an 8-core like the 7700X saves real money if you’ll never CPU-offload, at the cost of slower prompt processing on CPU-bound steps. We’d still take the 16 cores at this budget — the delta is small against a $4,600 total.

Motherboard: ASUS ProArt X870E-Creator WiFi — the x8/x8 requirement

This is the part most $5,000 builds get wrong. Most AM5 boards wire the second full-length slot to the chipset at x4 — electrically fine, but your second 3090’s prompt processing and model loading crawl through a shared uplink. The ProArt X870E-Creator (ASUS spec sheet) splits the CPU’s x16 into true x8/x8 across two PCIe 5.0 slots, spaced far enough apart to physically fit two 3-slot cards. Street price runs ~$480 ($380 open-box appears at Newegg periodically).

Alternative. ASUS’s ProArt B850-Creator does the same x8/x8 split on the cheaper B850 chipset — TechSpot’s 21-board roundup flags it as the value x8/x8 pick — if you can live with fewer USB4 ports and 5GbE instead of 10GbE. Whatever you choose, verify “x8/x8 bifurcation” appears in the manual’s expansion-slot table before ordering. This is the one component where a $150 savings can silently cost you PCIe bandwidth.

RAM: 32GB DDR5-6000 — the memory supercycle decides this one

In a sane market this build would carry 64GB. In September 2026, a name-brand 64GB DDR5-6000 kit is $869.99 — more than a PS5, and roughly 4× its 2024 price (Tom’s Hardware’s price index tracks the whole ugly chart). Mainstream 32GB DDR5-6000 kits cluster at $399–$479 per Newegg’s own DDR5 crisis guide.

Here’s why 32GB is enough for this specific build: the dual-3090 design keeps models fully in VRAM. System RAM only becomes the bottleneck when you offload — and offloading is exactly what this machine was built to avoid. 32GB covers the OS, llama.cpp’s mmap window, and a working set of tooling comfortably.

Alternative. Pay the $430 upgrade to 64GB only if you have a concrete plan to run 100GB+ MoE models with --n-cpu-moe offload. The four-DIMM-slot escape hatch from our $2,000 build still applies: add a second identical 2×16GB kit later, accepting that four sticks may need to run below 6000MHz.

Storage: Samsung 990 Pro 2TB — the floor, not the ceiling

2TB is the practical minimum when a single 70B GGUF is 42.5GB and you’ll want 4–6 models plus checkpoints on disk. The 990 Pro 2TB runs ~$338 now (BestDiskPrices’ September check) — NAND has the same AI-demand problem as DRAM, and drives that cost $130–$155 on spring deals are $300+ today. A Gen4 drive loads a 42GB model in roughly 6 seconds; that’s the number you feel every time you swap models.

Alternative. WD Black SN850X 2TB trades at similar money — buy whichever is cheaper the week you order. Don’t pay a Gen5 premium; model loading doesn’t saturate Gen4, and Gen5 drives run hotter for nothing you’ll notice in Ollama.

PSU: 1200W ATX 3.1 — two 350W cards need real headroom

Two 3090s pull ~700W of GPU load before the CPU says anything, and 3090s are famous for millisecond transient spikes well above TDP. 1200W is the right rating — the full math is in our PSU sizing guide. The good news: 3090s use classic 8-pin PCIe connectors, so you need a unit with four+ 8-pins, not 12V-2x6 cables. The Lian Li Edge 1200W (Platinum, ATX 3.1) at ~$179 covers it; ASRock’s 1300W Platinum at ~$189 is the step-up if it’s in stock (PCTechKits’ 2026 roundup covers the field).

One operating cost note: at the EIA’s 18.83¢/kWh residential average (EIA electricity monthly), this rig fully loaded costs about $0.16/hour to run. Overnight batch jobs are cheap; 24/7 idle is not free — the power bill math is here.

Case and cooling: the two-card heat problem is real

The Lancool 216 (~$103, Tom’s Hardware review) is here for its two 160mm front intakes, because the actual thermal problem in this build isn’t the CPU — it’s the top 3090 inhaling the bottom 3090’s exhaust. With x8/x8 slots three slots apart and two ~3-slot cards, the gap between cards is thin. Undervolt or power-limit both cards (nvidia-smi -pl 280 costs ~3% performance for 20% less heat) and keep front intake unobstructed. The Peerless Assassin 120 SE ($50, GamersNexus’ review made it the default answer) handles the 9950X fine because sustained inference load lives on the GPUs.

First boot: the mistake almost everyone makes once

Here’s the failure you’re most likely to hit, because we’ve seen it repeatedly: both cards show up in nvidia-smi, the model loads, and 70B decode comes in at 3 tok/s instead of 8 — with prompt processing taking minutes. The cause is almost always the second card sitting in the chipset-attached x4 slot instead of the second CPU slot. Check the negotiated link before blaming the software:

$ nvidia-smi --query-gpu=index,name,pcie.link.gen.current,pcie.link.width.current --format=csv
index, name, pcie.link.gen.current, pcie.link.width.current
0, NVIDIA GeForce RTX 3090, 4, 8
1, NVIDIA GeForce RTX 3090, 4, 8

Both cards should report width 8 (Gen4 — the 3090 is a Gen4 card in a Gen5 slot, which is fine). If card 1 reports width 4 or 1, physically move it to the board’s second CPU-wired slot and enable x8/x8 bifurcation in BIOS. Once the lanes are right, a fully-resident 70B looks like this in a current llama.cpp build:

$ ./llama-cli -m Llama-3.3-70B-Instruct-Q4_K_M.gguf -ngl 99 --split-mode layer -c 8192
...
llama_perf_context_print: eval time ... ( 8.4 tokens per second)

Deeper multi-GPU gotchas — IOMMU groups, ACS, NCCL P2P failures — have their own troubleshooting guide, and the NVLink-vs-PCIe question is settled with numbers in our interconnect breakdown. If you’re pairing this rig with a coding agent, the same box serves as a BYOK backend for Cline or Cursor — aicoderscope.com covers that side, and aifoss.dev has the Ollama multi-GPU server configuration.

Where the last $400 goes

Three honest options, in the order we’d pick them:

  1. Keep it as buffer. Used 3090 asks have drifted up all summer; the tracker average is $1,264 but clean, returnable cards from reputable sellers routinely close near $1,343. Two of those and your buffer is gone. This is the most likely outcome.
  2. A used NVLink bridge (~$100). 3090-only, improves cross-card transfers ~50%, helps tensor-parallel workloads and training more than llama.cpp layer split. Check your slot spacing matches the bridge width before buying.
  3. RAM to 64GB (+~$430). Only with a concrete MoE-offload plan, per the RAM section.

What to actually buy

Prices as of September 2026, all taken from the comparison above:

Your situationThe machinePriceWhere
You want 70B models fully in VRAM — this build2× used RTX 3090 + parts list above~$4,600Check price
You run 27B–35B fast, 70B rarelyUsed RTX 4090, same platform~$4,300 all-inCheck price
You want 48GB with zero multi-GPU setupUsed RTX A6000, any board~$2,600–$3,800 + platformCheck price
Undecided — test the 48GB workload firstRented GPU, 3090 from $0.07/hrpay per hourVast.ai

Before committing, run your exact model-plus-context numbers through the VRAM calculator — the difference between “fits in 48GB” and “doesn’t” is often the context window, not the weights.

FAQ

Why not one RTX 5090 instead of two 3090s? Price and capacity. A 5090 is ~$4,299 in-store when Micro Center has stock and $5,199+ online — the card alone eats the entire budget and its 32GB still can’t hold a 70B at Q4. The used 4090 vs 5090 article has September’s full pricing.

Will two 3090s make my 13B model twice as fast? No. Two cards buy capacity, not speed — a model that fits one card should stay on one card. Splitting a small model across two cards is usually slower than one card; a tuned dual-3090 setup running a 27B hit 22.8 tok/s split vs ~40 tok/s on a single card (sanj.dev’s tuning writeup).

Can I start with one 3090 and add the second later? Yes, and it’s a reasonable hedge — buy the x8/x8 board and 1200W PSU up front (the extra cost over single-GPU parts is ~$300) and you can add card two whenever the market or your workload demands it.

Is buying used GPUs at these prices actually safe? It’s manageable risk, not zero risk: sellers with return windows, connector inspection on arrival, and a stress test inside the return period. The market has priced 3090s up 30%+ since spring precisely because they keep proving durable at sustained AI loads.

Sources

Last updated September 18, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.