Apple M5 Ultra Mac Studio for Local AI in 2026: 1.2 TB/s, a 256GB Ceiling Until October, and Whether to Pre-Order

mac-studioapple-siliconlocal-llmhardwareunified-memory

TL;DR: Apple announced the M5 Ultra Mac Studio on August 25 — 1.2 TB/s of unified memory bandwidth at every config, $5,499 for 96GB, shipping September 22. The 512GB headline config isn’t orderable until late October and has no announced price. Pre-order the 256GB now only if you’re serving 200GB-class MoE quants; everyone else should wait for real benchmarks.

M5 Ultra 96GB ($5,499)M5 Ultra 256GB (~$9,499)4× used RTX 3090 (~$5,144)
Best for70B dense + 100B-class MoE, one quiet boxHy4/GLM-5.3 1–2-bit frontier quantsMax tok/s on models that fit 24GB shards
Price / Cost$5,499, ~200W under load~$9,499 (+$4,000 memory tier)~$5,144 in cards + ~$1,500 platform, 1,400W+
The catch96GB buys no new model class over a 128GB Strix Halo boxBig MoE quants decode in single digits anywayPower, heat, PCIe plumbing, and no warranty

Honest take: The 1.2 TB/s bandwidth is the real news, not the 512GB number — but until community benchmarks land after September 22, every tok/s figure for this machine (including ours below) is an estimate. Don’t spend $9,499 on math you can’t verify yet.

Apple opened pre-orders for the new Mac Studio M5 Ultra on August 25, 2026, with deliveries starting September 22. The spec sheet reads like it was written for the r/LocalLLaMA front page: up to 512GB of unified memory, 1.2 TB/s of memory bandwidth, and a claimed 4.3× jump in peak AI compute over the M3 Ultra it replaces.

If you’ve been holding off on a big unified-memory box since our M7 wait-or-buy guide, this is the machine you were told to wait for. Whether you should actually pre-order it is a different question, and the answer depends on which of the three memory tiers you’re looking at — one of which you can’t buy yet. Before anything else, check your target model against your candidate config with our VRAM calculator; the fits at each tier are less generous than the headline numbers suggest.

What Apple actually announced (and what ships when)

The August 25 refresh replaced the M4 Max / M3 Ultra Mac Studio lineup with M5 Max and M5 Ultra. The numbers that matter for local AI, verified across Apple’s newsroom release and launch coverage:

SpecM5 Max StudioM5 Ultra Studio (base)M5 Ultra Studio (top chip)
CPU / GPU coresup to 16 / 4030 / 6436 / 80
Unified memoryup to 128GB96GB256GB (512GB late Oct)
Memory bandwidth614 GB/s1.2 TB/s1.2 TB/s
Starting price$2,499$5,499~$9,499 w/ 256GB
ShipsSep 22Sep 22Sep 22 (512GB: late Oct)

Three details in the fine print change the buying decision:

The 512GB tier is not orderable today. Apple’s own launch materials and MacRumors both confirm the 512GB option arrives in late October, with no announced price. Coverage speculates around $15,000, but that’s expectation, not a price list — Apple charged $4,000 to go from 96GB to 256GB, and the 256GB-to-512GB step has no public number. If the 512GB config is the reason you’re excited, you are not making a September decision at all.

256GB and 512GB require the top chip. The bigger memory tiers are tied to the 36-core CPU / 80-core GPU version of the M5 Ultra, so the ~$9,499 street math for a 256GB machine already includes the GPU bump. A maxed 256GB / 16TB configuration rings up at $18,299 today, per AppleInsider.

Bandwidth is flat across configs. Unlike the memory tiers, the 1.2 TB/s figure applies to every M5 Ultra, including the $5,499 base. That’s a 46% step over the M3 Ultra’s 819 GB/s, and it’s the single number that most directly predicts decode speed on large dense models.

What 1.2 TB/s means for tokens per second

No M5 Ultra has reached reviewers yet, so there are no measured LLM numbers — anyone quoting real tok/s before September 22 is guessing. What we can do is bandwidth-scale from machines this site has verified numbers for, because token generation on large dense models is memory-bandwidth-bound:

  • A 546 GB/s M4 Max sustains 28.4 tok/s on a 70B Q4 (measured, from our M5 Max MacBook guide canon).
  • An 819 GB/s M3 Ultra pushes roughly 40 tok/s on the same class of model.

Scaling both anchors to 1.2 TB/s lands in the same window:

Model classM3 Ultra (819 GB/s)M5 Ultra (1.2 TB/s), scaled estimate
8B Q4 dense~95 tok/s120+ tok/s (compute limits kick in)
70B Q4 dense (~40GB)~40 tok/s~55–62 tok/s
gpt-oss-120b MoE (~61GB)~35–45 tok/s~50–65 tok/s
200GB-class 1–2-bit MoEn/a (doesn’t fit 192GB)single digits to low teens

Treat every number in the right-hand column as an estimate until community benchmarks land — that’s how the arithmetic works, not a promise the silicon keeps. Two honest caveats:

Prompt processing is the historic Apple Silicon weakness, and it’s the unverified variable. Decode scales with bandwidth; prefill scales with compute. Apple’s claimed 4.3× peak AI compute over M3 Ultra comes from the M5 generation’s neural accelerators in each GPU core — if that translates to llama.cpp and MLX prefill, it addresses the exact complaint that made long-context work painful on the M3 Ultra. Nobody outside Apple has tested it. Wait for numbers before assuming your 90-second coding-agent prefills disappear.

Big MoE quants are slow everywhere. GLM-5.3’s IQ1_S measured 3–6 tok/s at ~230GB resident on DDR5 server hardware; at 1.2 TB/s the same math suggests high single digits to low teens. Capacity gets you in the door with frontier quants; it doesn’t make them pleasant daily drivers.

What each memory tier actually runs

96GB ($5,499). The awkward tier. macOS reserves a slice of unified memory for the system, so the practical GPU budget is below the sticker (you can raise the wired limit with sudo sysctl iogpu.wired_limit_mb, but leave the OS at least 8–12GB). That’s comfortable for 70B Q4 dense (~40GB), gpt-oss-120b (~61GB), and Tencent Hy3’s 91.8GB Q4 becomes a squeeze that only works with an aggressive wired-limit push. Here’s the problem: a $2,000-class 128GB Strix Halo box or a 128GB M5 Max already covers this model class — see our 128GB unified memory guide — at half to a third the price, just slower. You’re paying $5,499 for speed on models the cheaper boxes already run, not for new capability.

256GB (~$9,499). The tier with a genuine story. This is the first single consumer-orderable box since the M3 Ultra 512GB that holds the new wave of official 1–2-bit frontier quants: Hy4’s STQ1_0 at ~214GiB resident fits with room for context, and GLM-5.3’s ~230GB IQ1_S fits with almost none. Kimi K3’s 1.56TB remains out of reach of everything short of a cluster. When we wrote that 256GB machines are the entire market for Hy4’s STQ1_0 quant, this machine is what that market now looks like — $9,499, one power cable, under 200W while inferencing (the M3 Ultra ran DeepSeek R1 671B in-memory under 200W, per TechRadar, and the M5 Ultra’s envelope is similar).

512GB (late October, price TBD). This tier would hold Hy4’s Q4_K_M at ~435GiB — a materially better quant than STQ1_0, not just more headroom — and it’s the config that would make a single Studio a genuine alternative to the 4-box Thunderbolt cluster approach. But you can’t order it, you don’t know its price, and by late October you’ll also have real M5 Ultra benchmarks. There is no scenario where pre-ordering something else now to “hold your place” for this tier makes sense.

The $5,144 elephant: four used RTX 3090s

The traditional answer to “96GB for local AI” is four used RTX 3090s. The DRAM crisis has mangled that math: used 3090s now average $1,286 each (ResalePrices, August 2026), so four cards run ~$5,144 before you’ve bought the EPYC/Threadripper platform to hang them on — call it $6,500–7,000 all-in, versus $5,499 for the 96GB Studio. Our used 3090 guide covers the per-card decision; here’s the honest split at rig scale:

  • The 3090 rig still wins raw tok/s on anything that fits in one or two cards — 936 GB/s per card beats 1.2 TB/s shared when the model shards cleanly, and CUDA prefill remains far ahead of Apple Silicon’s measured prefill to date.
  • The Studio wins everything else: ~200W vs 1,400W+ under load, silence, no PCIe lane Tetris, no mining-history roulette, a warranty, and one power cable. At $0.12/kWh, the rig’s extra ~1.2kW is roughly $0.14/hour — about $500/year at 10 hours/day — and it heats the room it lives in.
  • Capacity is no longer close. The 96GB rig ceiling is fixed; the Studio line goes to 256GB today and 512GB in October. The 200GB-quant era we’ve entered this August is simply not addressable with consumer NVIDIA cards at any sane price — the 96GB RTX PRO 6000 costs more than the 256GB Studio.

And if you need frontier-quant capacity only occasionally, renting stays cheaper than either: a RunPod multi-GPU pod runs big MoE inference by the hour without a five-figure box in your office. Our take from the Hy4 guide stands — for models this size, rent-don’t-build is still the default until your utilization is high.

Pre-order, wait, or skip

Pre-order the 256GB if you are the specific person the STQ1_0/IQ1_S era created: you want GLM-5.3 or Hy4 class models resident locally, you accept single-digit-to-low-teens decode as the price of frontier weights at home, and $9,499 is justified by privacy or workload. Nothing else orderable today does this in one box.

Wait until late October if you want the 512GB tier (obviously), or if your decision hinges on tok/s rather than capacity — real MLX and llama.cpp benchmarks will exist by then, the prefill question will be answered, and you’ll know the 512GB price. The M5 Ultra is not a limited drop; it will still be for sale in November.

Skip if your models fit in 24–32GB. A used 3090 — even at today’s inflated $1,286 — or an RTX 5090 delivers more tok/s per dollar for the 70B-and-under world, and the Studio’s capacity advantage buys you nothing you’d use. That’s the same verdict tier logic as our M3 Ultra vs dual 4090 comparison, and the M5 generation doesn’t change it.

For the software side once you have one — Ollama, MLX, and Open WebUI setup on Apple Silicon — our friends at aifoss.dev maintain the self-hosting guides we defer to.

FAQ

Is the M5 Ultra Mac Studio’s 512GB config available now? No. Pre-orders opened August 25 for 96GB and 256GB configs (shipping September 22), but Apple says the 512GB option arrives in late October with pricing not yet announced.

How fast will the M5 Ultra run a 70B model? Nobody has measured it yet. Bandwidth-scaling from the M3 Ultra’s ~40 tok/s at 819 GB/s suggests roughly 55–62 tok/s on a 70B Q4 at 1.2 TB/s — treat that as an estimate until community benchmarks land after September 22.

Does the base $5,499 M5 Ultra get the full 1.2 TB/s bandwidth? Yes. Bandwidth is 1.2 TB/s on every M5 Ultra config, including the 30-core CPU / 64-core GPU base model with 96GB. Only the memory capacity tiers are gated to the 36-core/80-core chip.

Can the 256GB M5 Ultra run Kimi K3 or full GLM-5.3? Kimi K3, no — its smallest serving footprint is ~1.56TB. GLM-5.3 only at the ~230GB-resident IQ1_S 1-bit quant, which fits with almost no headroom and decodes slowly; the 320B GLM-5.3-Flash at 93GB is the more realistic Studio target.

Should I buy a discounted M3 Ultra instead? If you find a 512GB M3 Ultra meaningfully under its original $9,499, it remains the cheapest 512GB box in existence and runs the same quants ~46% slower. At small discounts, the M5 Ultra’s bandwidth and (claimed) prefill improvements are worth the wait.

  • Mac Studio M5 Ultra — 96GB/$5,499 for speed on 70B-class models; 256GB for the 200GB-quant era.
  • RTX 3090 — still the tok/s-per-dollar pick for models that fit 24GB, even at 2026 prices.

Sources

Last updated September 1, 2026. Prices and specs change; verify current rates before purchasing. M5 Ultra performance figures are bandwidth-scaled estimates, not measurements — community benchmarks arrive after September 22.

Was this article helpful?