Mac Studio M5 Ultra vs M3 Ultra for Local AI in 2026: The Used Market Just Made This Decision Easy
TL;DR: A used M3 Ultra now costs the same as a new M5 Ultra — eBay asks for the 96GB config run $4,899–$6,100 against Apple’s $5,499, and 256GB asks sit within $300 of Apple’s $9,499. The M5 Ultra has 46% more bandwidth and ships September 22. Pay used prices for old silicon only at a real discount, which currently doesn’t exist.
| M5 Ultra 96GB (new, $5,499) | M5 Ultra 256GB (new, ~$9,499) | Used M3 Ultra 96GB ($4,899–$6,100 asks) | |
|---|---|---|---|
| Best for | 100B-class MoE + 70B dense, shipping Sep 22 | 200GB-class frontier MoE quants | Needing the machine this week |
| Bandwidth | 1.2 TB/s | 1.2 TB/s | 819 GB/s |
| The catch | No independent benchmarks until after Sep 22 | $4,000 for the memory tier alone | Same money as new, ~32% less bandwidth, clock starts on a 2025 chip |
Honest take: Buy the new M5 Ultra. The used market hasn’t repriced since the August 25 refresh — until an M3 Ultra 96GB sells for $4,000 or less, you’d be paying 2026 prices for 2025 bandwidth.
Apple replaced the Mac Studio lineup on August 25, 2026: the M3 Ultra — the machine that made “DeepSeek R1 at home under 200W” a real sentence — is gone from the configurator, and the Mac Studio M5 Ultra takes its slot at $5,499. Deliveries start September 22. That leaves anyone shopping for a big unified-memory box with exactly two paths: pre-order the new one, or hunt the used market for the old one.
Normally that’s a genuine dilemma, because used Apple hardware depreciates and old flagships become bargains. Not this time. We priced both paths in mid-September 2026, and the used market is asking new-machine money for the old machine. This article lays out the measured numbers for the M3 Ultra, the honest math for the still-unbenchmarked M5 Ultra, and where the line sits at each memory tier. Before committing to any tier, run your target model and context through our VRAM calculator — the fits are less generous than the headline memory numbers suggest.
The used M3 Ultra market hasn’t gotten the memo
The Mac Studio M3 Ultra launched March 5, 2025 at $3,999 for the base 96GB/1TB config (28-core CPU, 60-core GPU), with the 512GB memory tier a $4,000 upgrade on the 32-core/80-core chip. Then the DRAM crunch rewrote its history: Apple pulled the 512GB config around March 2026 and the 256GB tier by May, sold the remaining 96GB-only lineup through the summer, and retired the whole machine on August 25.
Scarcity did what scarcity does. Here’s what the two Ultras actually cost in mid-September 2026:
| Config | Used/leftover M3 Ultra (eBay asks, Sep 2026) | New M5 Ultra (Apple, Sep 2026) |
|---|---|---|
| 96GB | $4,899 pre-owned; $5,690 refurbished; $6,099 sealed; $6,195 new-old-stock 2TB | $5,499 (1TB, orderable now) |
| 256GB | $9,199–$9,399 (2TB) | ~$9,499 (+$4,000 memory tier, orderable now) |
| 512GB | $18,250–$27,500 asks | Price TBA, orders open late October |
Read that middle column again. A pre-owned M3 Ultra 96GB at $4,899 is $600 below a factory-new M5 Ultra with 46% more memory bandwidth and a full warranty. The refurbished and sealed units ask more than the new machine. At 256GB, the gap is under $300. Only the 512GB tier still favors the old chip — because you literally cannot order the new one yet, and sellers know it. We covered that seller’s market in the M3 Ultra vs dual RTX 5090 comparison: asks of $18,250–$27,500 against a $9,499 original price for the 512GB config.
This is the DRAM shortage working exactly as it has all year — the same force that pushed used RTX 3090s up 72% and Strix Halo mini PCs up 60%. Used prices track replacement cost, and when memory is the scarce ingredient, a 512GB machine nobody can buy new stops depreciating entirely.
Spec sheet: what two generations actually changed
There was never an M4 Ultra — Apple skipped the generation, so the M3 Ultra (March 2025) hands off directly to the M5 Ultra (August 2026). The changes that matter for inference:
| Spec | M3 Ultra | M5 Ultra |
|---|---|---|
| Memory bandwidth | 819 GB/s | 1.2 TB/s (+46%) |
| Memory tiers | 96 / 256 / 512GB (as sold) | 96 / 256GB now, 512GB late Oct |
| CPU | 28-core (base) / 32-core (top) | 30-core (base) / 36-core (top) |
| GPU | 60-core / 80-core | 64-core / 80-core, Neural Accelerators |
| Peak AI compute | baseline | 4.3× (Apple’s claim) |
| Base price at launch | $3,999 | $5,499 |
Two of those rows do almost all the work. Decode speed — the tokens-per-second you feel while a model writes — is memory-bandwidth-bound, so the 819 → 1,200 GB/s jump is worth roughly 1.46× on every model that fits in memory, before any software wins. Prompt processing — the wait before the first token — is compute-bound, and that’s where the 4.3× AI-compute claim and the per-core Neural Accelerators aim. Prefill has always been the Mac’s weak leg against NVIDIA hardware, so if Apple’s number survives independent testing at even half strength, it fixes the right problem.
The honest caveat: as of September 16, nobody outside Apple has benchmarked an M5 Ultra, because the machine hasn’t shipped. Apple’s performance figures come from its own July 2026 testing. Every M5 Ultra tok/s figure below is bandwidth arithmetic, and we’ve flagged it as such. The M3 Ultra numbers, by contrast, are measured.
What the M3 Ultra measurably does
Eighteen months in the field means the M3 Ultra’s local-AI behavior is documented to a level the M5 Ultra won’t reach until October. The anchors, all from independent testing on the 512GB/80-core config:
- DeepSeek R1 671B, 4-bit quant: 17–18 tok/s, under 200W. The full 404GB model resident in unified memory with 448GB manually wired to the GPU. This remains the single most power-efficient way to run a frontier-class model at home — a quad-GPU rig with the same capacity pulls 1,400W+.
- gpt-oss-120b (MXFP4): 79.7 tok/s generation, 1,244 tok/s prompt processing in llama.cpp. For comparison, an RTX PRO 6000 does 196 tok/s on the same model and a DGX Spark does 38.5 — the Mac sits squarely between the $14,000 card and the $4,699 appliance.
- DeepSeek V3 671B Q4: ~6.2 tok/s in llama.cpp — consistent with a ~37B-active MoE on 819 GB/s, and a reminder that “it fits” and “it’s pleasant” are different claims.
- The prefill wall is real. Hardware Corner measured a 14-minute wait to first token feeding DeepSeek 671B a long prompt through llama.cpp, and found context sizes in the 16K–32K range introduce serious slowdowns. Generation speed was fine once it started; the compute-bound prompt pass is what crawled. Long-document RAG and agentic workloads that re-read big contexts feel this constantly. Shorter chat contexts don’t.
That last point is the M3 Ultra’s defining trade-off, and it’s exactly the axis the M5 Ultra claims to attack. A 46% decode bump is nice; a multiple on prefill would change what the machine is for. We just can’t verify the second part yet.
One operational note that applies to both machines: macOS caps GPU-wired memory at roughly 75% of unified memory by default, so a 96GB Studio offers about 72GB to the GPU — which is why gpt-oss-120b (68.5GB with full context in llama.cpp) is a tight squeeze and bigger quants refuse to load with memory to spare. The fix is one command:
sudo sysctl iogpu.wired_limit_mb=81920
# expected output:
# iogpu.wired_limit_mb: 0 -> 81920
That raises the wired ceiling to 80GB on a 96GB machine (the 448GB allocation in the R1 test above is the same trick at 512GB scale). Leave the OS 12–16GB and it stays stable; the setting reverts on reboot, so add it to a LaunchDaemon if you serve models 24/7. Details on running the big quants day-to-day are in our 100B models on Mac Studio guide.
What the M5 Ultra should do (arithmetic, not benchmarks)
Decode scales with bandwidth on Apple Silicon reliably enough that the estimate is worth making, as long as it’s labeled. At 1.2 TB/s against the M3 Ultra’s measured numbers:
| Model | M3 Ultra (measured) | M5 Ultra (bandwidth estimate) |
|---|---|---|
| gpt-oss-120b MXFP4 | 79.7 tok/s | ~115 tok/s |
| DeepSeek R1 671B Q4 (needs 512GB) | 17–18 tok/s | ~25 tok/s |
| DeepSeek V3 671B Q4 (needs 512GB) | ~6.2 tok/s | ~9 tok/s |
Those are estimates — treat them as ceilings, not promises. Prefill is the number to actually watch when third-party benchmarks land after September 22. Apple’s 4.3× peak-AI-compute claim, if it translates to even 2–3× in llama.cpp/MLX prompt processing, turns the 14-minute wait into something livable and makes the Mac a credible agentic-workload box rather than a chat box. If it doesn’t translate, the M5 Ultra is “the M3 Ultra, 46% faster, at 2024 prices” — still the better buy at parity, but not a category change. Our M5 Ultra deep dive covers the tier-by-tier pre-order logic, and the DGX Spark vs M5 Max comparison covers why prefill is where Apple loses to CUDA today.
The memory-tier decision
96GB runs 70B dense at Q4 comfortably and gpt-oss-120b with the wired-limit bump, but so does a $2,000-cheaper 128GB Strix Halo box (slower, at 256 GB/s) — the 96GB Ultra’s case is speed on those mid-size models, and at $5,499 it’s a bandwidth purchase, not a capacity purchase. If everything you run fits in 24GB, stop here: a used RTX 3090 decodes those models faster than either Mac for around $1,050.
256GB is where the Ultra does something no GPU short of a $14,000 RTX PRO 6000 can: hold 200GB-class MoE quants (GLM-5 and Qwen3.8 2–3-bit, gpt-oss-120b at full precision alongside a draft model) in one quiet box. The $4,000 tier jump is steep, but the used 256GB M3 Ultra asking $9,199 confirms the market agrees it’s worth roughly that.
512GB — the R1/V3-at-home tier — is a waiting game. New orders open late October at an unannounced price; used M3 Ultra 512s ask $18K+. Unless someone is paying you to run 671B locally starting this week, wait the six weeks. If the itch is unbearable, rent the workload first: a 96GB card on RunPod runs about $2/hour and tells you whether your use case actually needs 400GB of weights resident.
What to actually buy
Prices as of September 2026, all taken from the comparison above:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| Buying a big-memory Mac and can wait until Sep 22 | Mac Studio M5 Ultra 96GB | $5,499 | Check price |
| Running 200GB-class MoE quants | Mac Studio M5 Ultra 256GB | ~$9,499 | Check price |
| Found an M3 Ultra 96GB at $4,000 or less | Used Mac Studio M3 Ultra | ≤$4,000 (rare) | Check price |
| Everything you run fits in 24GB | Used RTX 3090 | ~$1,050 | Check price |
| Want 671B-class local AI | Wait for M5 Ultra 512GB | TBA, late Oct | Apple |
| Undecided — test the workload first | Rented GPU, ~$2/hr | pay per hour | RunPod |
The used-M3-Ultra row deserves its own sentence: $4,000 is the number because that was the machine’s actual new price in 2025, and paying more than that today means paying a scarcity premium for the slower chip while the faster one sits in stock. Current asks are $900–$2,100 above that line. They’ll come down once M5 Ultras start arriving in quantity — if you’re patient and specifically want the M3 Ultra as a budget capacity box, check back in November.
If the Mac is destined to be a coding-model server, our sister site covers wiring local backends into Cursor and Cline, and aifoss.dev covers the Ollama/MLX serving stack that makes a headless Studio pleasant.
FAQ
Was there ever an M4 Ultra Mac Studio? No. Apple skipped the M4 generation for the Ultra tier — the March 2025 Studio paired the M4 Max with the M3 Ultra, and the August 2026 refresh went straight to M5 Max/M5 Ultra. If a listing says “M4 Ultra,” it’s mislabeled.
Is the M5 Ultra’s 1.2 TB/s available on every config? Yes — 96GB, 256GB, and the upcoming 512GB all get the full 1.2 TB/s. Bandwidth doesn’t scale down with the cheaper memory tier, which is unusual for Apple and makes the base config more interesting than the M3 Ultra’s base ever was.
How much faster will the M5 Ultra actually feel? On decode, expect roughly 1.4–1.5× the M3 Ultra — bandwidth arithmetic that history says Apple Silicon hits. Prompt processing is the unknown: Apple claims 4.3× peak AI compute, and no independent test exists yet. If long-context prefill improves by anything close to that, the M5 Ultra is a different class of machine.
Should I buy a used M3 Ultra 512GB to run DeepSeek R1 now? At $18,250–$27,500 asks, no. That’s double the config’s original price for 819 GB/s. The M5 Ultra 512GB opens for orders in late October; even if Apple prices it at $13,000+, it will likely undercut the used asks while adding 46% bandwidth and a warranty.
Does the M3 Ultra still make sense at any price? Yes — at or below its $3,999 original price for the 96GB, it’s a fine capacity box with measured, known behavior. The problem is finding one there. Every dollar above $4,000 narrows the gap to a new M5 Ultra that’s faster in every dimension.
Recommended Gear
- Mac Studio M5 Ultra — the default pick at $5,499 (96GB) or ~$9,499 (256GB); ships from September 22
- Mac Studio M3 Ultra — only at $4,000 or below for the 96GB config; measured 17–18 tok/s on R1 671B (512GB config) at under 200W
- RTX 3090 24GB — the ~$1,050 answer if your models fit in 24GB and you want maximum tok/s per dollar
Sources
- Apple introduces new Mac Studio with M5 Max and M5 Ultra — Apple Newsroom
- New Mac Studio M5 Max and M5 Ultra: Everything you need to know — Macworld
- Mac Studio gets update to M5 Max and M5 Ultra — AppleInsider
- Apple announces the M5 Ultra Mac Studio with up to 512GB of RAM — Macworld
- A Maxed Out M3 Ultra Mac Studio Will Cost You $14,099 — MacRumors
- Mac Studio With M3 Ultra Runs Massive DeepSeek R1 AI Model Locally — MacRumors
- Mac Studio M3 Ultra runs DeepSeek R1 671B entirely in memory using less than 200W — TechRadar
- Running gpt-oss with llama.cpp (M3 Ultra 512GB results) — ggml-org/llama.cpp Discussion #15396
- 14-Minute Wait?! $10K Mac Studio Crawls with DeepSeek 671B + llama.cpp — Hardware Corner
- How Fast is Mac Studio M3 Ultra Running the New DeepSeek V3 LLM? — Hardware Corner
- Apple Mac Studio M3 Ultra 96GB used/refurbished listings — eBay
- Mac Studio roundup (M5 Max / M5 Ultra) — MacRumors
Last updated September 16, 2026. Prices and specs change; verify current rates before purchasing. Used-market figures are seller asking prices, not sold prices.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.