Mistral Large 4 (Le Chonk) for Local AI: What Hardware Runs a 1.05T Open-Weight MoE When the Weights Drop

mistralmistral-large-4local-llmmoehardwareapple-siliconvrambuying-guide

TL;DR: Mistral announced Large 4 (“Le Chonk”) on October 6 — a 1.05-trillion-parameter multimodal MoE with 49B active parameters, with open weights promised by the end of October. A ~3-bit quant should land around 440–470GB, which exactly one home machine can hold: the Mac Studio M5 Ultra 512GB, itself arriving late October at a still-unannounced price. Everyone else uses the API.

Mac Studio M5 Ultra 512GBEPYC DDR5 RAM build (768GB)Mistral API (preview)
Best forOnly single-box local path at usable speedAlready-built multi-channel rigsEveryone else
PriceTBA (256GB is $9,499; 512GB will be more)~$10,000+ in RAM alone post-crisis$0.68/M in, $2.09/M out (preview)
Expected speed~20–30 tok/s (estimate, see ceiling math)~5–8 tok/s (estimate)Instant
The catchPrice unknown, orders open late OctoberDDR5 RDIMM prices tripled in 2026Your data leaves the building

Honest take: Don’t buy anything for this model yet — the weights, the license, and the only machine that fits it all land in the same late-October window. If Large 4’s quality claims survive independent testing, the M5 Ultra 512GB becomes the most interesting local AI purchase of the year. Until then, run Mistral Small 4 on hardware you already own, and judge Le Chonk through the API.

Mistral announced Large 4 on October 6 with a nickname that does half of this article’s work for it: “Le Chonk” (Mistral AI, The Next Web). It is a 1.05-trillion-parameter granular mixture-of-experts model with 49B active parameters per token and a 1.6B vision encoder, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacenters (Unite.AI). The preview is API-only today. The open weights are promised for the end of the month — Mistral’s Hugging Face upcoming-release page currently says October 31.

That makes this the third 1T-class open-weight model a home-labber could theoretically download, after Moonshot’s Kimi K2 family and its successors. We’ve run the hardware math on those twice before (Kimi K2.6, Kimi K2.7), and the honest answer has always been “almost nobody should.” Large 4 changes the question in one specific way: its release lines up, almost to the week, with the first consumer machine that can actually hold a model this size at usable speed. More on that below — first, what the model is.

What Mistral actually announced

The spec sheet, as confirmed across the launch coverage (Unite.AI, Artificial Analysis):

SpecMistral Large 4Mistral Large 3 (Dec 2025)Kimi K2.7 (site canon)
Total parameters1.05T675B~1T
Active per token49B41B32B
ModalitiesText + image (1.6B vision encoder)Text + imageText
Context window1M advertised (preview endpoint lists 524,288)256K256K
LicenseNot yet announcedApache 2.0Modified MIT
Weights”End of October” (HF page: Oct 31)OpenOpen

Two caveats worth flagging before anyone budgets around this. First, the license is unannounced. Mistral Large 3 shipped under Apache 2.0, and Mistral’s messaging frames Large 4 as its return to open frontier weights, but until a LICENSE file exists on Hugging Face, “open” is a press-release word. Second, the expert count, routing top-k, and layer layout haven’t been published — they arrive with the weights. The 49B-active figure is Mistral’s own, and every speed number below hangs on it.

The pricing: the preview API runs $0.68 per million input tokens and $2.09 per million output, half of the $1.36/$4.18 list price, with cached input at $0.07/M (Artificial Analysis, OpenRouter).

The benchmark story: real wins, and one honest asterisk

Mistral’s claim is that Le Chonk is the strongest open-weight model from any Western lab, benchmarked directly against DeepSeek, Qwen, Kimi, and GLM (OfficeChai). The numbers that survive cross-checking (Stork.ai, Yahoo Finance):

  • DeepSWE v1.1 (agentic coding): 61.7% — ahead of GLM-5.3 (61%), DeepSeek V4 Pro (57%), and Qwen3.8 Max (51%)
  • AutomationBench (agentic workflows): 59.9%, ahead of Kimi K3 and DeepSeek V4 Pro
  • FinWorkBench: 67%, tied with DeepSeek V4 Pro
  • Terminal-Bench 4.0: 28.3%
  • Surge AI human coding eval: 3.74/5 — second overall, behind Claude Opus 5 (4.22) but ahead of every open-weight competitor

The asterisk: on the aggregate Vals Index v2.1, Large 4 scores 48.05% — behind GLM-5.3 (53.51%), DeepSeek V4.1 Flash (51.32%), and Kimi K3 (50.30%). So the “best Western open model” framing is defensible, but “best open model, period” is not — the Chinese labs still lead on aggregate indices, and Large 4’s wins are concentrated in agentic coding, finance, and legal tasks. All of these are vendor-era numbers from launch week; independent evals will move them.

For what it’s worth to the coding crowd: at 61.7% DeepSWE with a 1M context, this is squarely aimed at agentic coding harnesses — the use case our sister site covers at aicoderscope.com, where the BYOK local-backend question is going to get interesting if the weights really are permissively licensed.

The GGUF math: what 1.05T parameters weighs on disk

No GGUFs exist yet — the weights aren’t out. But we have a nearly perfect size reference: Kimi K2.7 is also a ~1T-parameter MoE, and its community quants are site canon from our K2.7 guide: Unsloth’s 2-bit UD-Q2_K_XL is ~325GB, and the Q4-class build runs ~605GB on disk. Scaling by Large 4’s 5% larger parameter count (and treating the 1.6B vision encoder as a rounding error):

Quant (projected)Est. file sizeFits in
~2-bit (UD-Q2_K_XL class)~340GB384GB (4× RTX PRO 6000), 512GB Mac
~3-bit (Q3_K_XL class)~440–470GB512GB Mac, 768GB RAM build
~4-bit (Q4_K_M class)~630GB768GB RAM build only
Original weights (~BF16)~2.1TBNobody’s house

These are extrapolations, not measurements — treat them as ±10% until Unsloth or Bartowski publish real files. But the shape of the answer doesn’t move: there is no 128GB path. Even a 1-bit-class quant of 1.05T parameters lands near 180GB, so the entire Strix Halo / DGX Spark / 128GB Mac class — every machine in our 128GB unified-memory guide — is out before the conversation starts. That’s the single biggest practical difference from the 284B DeepSeek V4 Flash, which DS4 squeezed onto 128GB boxes last week.

The speed math: bandwidth ceilings, then reality

Decode speed is memory-bandwidth-bound. The ceiling is bandwidth divided by bytes read per token, and for a MoE you read only the active parameters: 49B, or roughly 22GB per token at 3-bit, ~29GB at 4-bit. Measured systems typically land at 45–95% of ceiling depending on stack maturity — 1T-class MoEs on Apple Silicon have historically sat near the bottom of that band.

The best real-world anchor we have: DeepSeek R1 671B (37B active, Q4 ≈ 22GB read per token — almost identical per-token traffic to Large 4 at 3-bit) runs a measured 17–18 tok/s on the M3 Ultra’s 819 GB/s (site canon). From that anchor:

HardwareBandwidthQuant that fitsCeilingRealistic estimate
Mac Studio M5 Ultra 512GB1,200 GB/s~3-bit (~460GB)~54 tok/s~20–30 tok/s
Used Mac Studio M3 Ultra 512GB819 GB/s~3-bit~37 tok/s~15–18 tok/s
4× RTX PRO 6000 (384GB)1,800 GB/s/card~2-bit onlyhigh, multi-GPU-limitedfastest, at $56,000+
EPYC 12-ch DDR5 + 768GB~460 GB/s~4-bit~16 tok/s~5–8 tok/s
Anything with 128GB—nothing fits——

Every estimate column is labeled an estimate because it is one: nobody outside Mistral has run this model. The K2.7 anchors (8–12 tok/s on 4× RTX 3090 + 256GB RAM, 8–11 tok/s on a 384GB EPYC build, both at 32B active) scale down to roughly 5–8 tok/s at Large 4’s 49B active — background-batch speed, not interactive speed.

Which leaves one genuinely new story. Apple opened orders for the Mac Studio M5 Ultra in late August, but the 512GB configuration ships separately in late October — pricing still unannounced, with the 256GB tier at $9,499 (Macworld, iThinkDiff). The first consumer machine with 512GB of unified memory at 1.2 TB/s and the first Western 1T open-weight model are arriving in the same two-week window. If the M3 Ultra 512GB precedent ($9,499 at launch) holds even loosely, this pairing is the first time “run a trillion-parameter frontier model at home, interactively” has a five-figure rather than six-figure answer. We compared the M5 Ultra against its predecessor in detail here — the short version is 1.2 TB/s vs 819 GB/s, which is the whole game for decode.

The RAM-build alternative died this year, incidentally. A 768GB EPYC build was the budget path to 600GB-class quants in 2025; with 64GB DDR5 RDIMMs now at $900+ against a ~$255 Q3 2025 contract price (our Threadripper platform guide has the full breakdown), 768GB is roughly $10,000 in memory before you’ve bought a CPU. The DRAM crisis didn’t just reprice gaming builds — it quietly deleted the cheap big-RAM path to 1T models.

The one config problem you’ll hit on a 512GB Mac

If you do end up loading a ~460GB quant on a 512GB Mac Studio, you will hit macOS’s GPU wired-memory limit before you hit the physical RAM limit. By default macOS caps GPU-wired allocations at roughly 75% of unified memory — about 384GB on a 512GB machine — so a 460GB model fails to load with Metal buffer allocation errors even though the RAM is physically there. The fix is the same sysctl we documented for DS4 on 128GB Macs, scaled up:

$ sudo sysctl iogpu.wired_limit_mb=491520
iogpu.wired_limit_mb: 393216 -> 491520

That raises the wired ceiling to 480GB, leaving ~32GB for the OS and the KV cache spillover. It resets on reboot; persist it via a LaunchDaemon if the box is a dedicated inference server. Leave at least 24–32GB unwired — starving macOS of pageable memory hard-locks the machine under sustained load.

What to actually buy

Prices as of October 2026, all verified in the comparison above or in the linked guides:

Your situationThe machinePriceWhere
You want Le Chonk local, interactive, single boxMac Studio M5 Ultra 512GB (orders late Oct)TBA ($9,499 for 256GB tier)Check availability
You want Mistral-quality coding on one card todayUsed RTX 3090 24GB + Devstral/Small-4 family$1,190–$1,400Check price
You already own a 4×3090 or EPYC rigWait for the ~2-bit Unsloth quant, re-check license first$0 today—
Undecided — want to test big-model workflows firstRented GPU, 3090 from $0.07/hrpay per hourVast.ai
You just want the model’s outputMistral API preview$0.68/M in, $2.09/M outmistral.ai

Run your own model-plus-context numbers in the VRAM calculator before committing to anything here.

FAQ

Can I run Mistral Large 4 on an RTX 4090 or 5090? No — not at any quantization, not with offloading you’d want to use. A ~2-bit quant is ~340GB; a 24–32GB card holds under 10% of that, and MoE expert routing makes NVMe/RAM offload of the rest brutally slow. For single-card Mistral, the ceiling is Mistral Small 4 (119B MoE, ~74GB at Q4) on multi-GPU, or Devstral Small 2 on a 24GB card.

Is it actually open weights? The weights are promised by end of October (Hugging Face lists October 31), but the license is unannounced. Large 3 was Apache 2.0; assume nothing until the LICENSE file exists. If it ships research-only, the entire local story above collapses to “use the API.”

Will it fit a 128GB machine like Strix Halo, DGX Spark, or a MacBook Pro M5 Max? No. Even a ~1-bit-class quant of 1.05T parameters is ~180GB. The smallest realistic home target is 384GB of pooled VRAM or a 512GB unified-memory Mac. The 128GB class tops out around the 284B/custom-quant tier — see what DS4 did with DeepSeek V4 Flash.

How does it compare to Kimi K2.7 for local use? Same weight class, worse per-token economics: 49B active vs 32B means roughly 1.5× the memory traffic per token, so every rig that runs K2.7 at 8–12 tok/s runs Large 4 proportionally slower. What Large 4 adds is native vision, a 1M context, and (pending license) Western-lab provenance — which matters to some compliance departments and not at all to your GPU.

Should I wait for the M5 Ultra 512GB price before deciding? Yes — that’s the whole recommendation. Weights, license, 512GB pricing, and the first independent benchmarks all land within about three weeks. Nothing about this model rewards buying early.

Products linked in this article:

  • Mac Studio M5 Ultra — the only single-box path to 1T-class local inference; 512GB tier ships late October
  • Used RTX 3090 — still the right answer for the Mistral models you can actually run today

Sources

Last updated October 7, 2026. Prices and specs change; verify current rates before purchasing. Mistral Large 4’s weights, license, and the M5 Ultra 512GB price were all unannounced at publication time — re-verify before spending money on this model’s account.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.