Reflection AI's Beam 501B: What Hardware Runs America's First Frontier Open-Weight Model When the Weights Drop

reflection-aibeamlocal-llmmoeopen-weightshardwareapple-siliconbuying-guide

TL;DR: Reflection AI announced Beam on October 5 — a 501B-parameter sparse MoE with 23B active, weights promised under Apache 2.0 later this month. Nothing is downloadable yet. When it lands, a ~163GB 2-bit quant fits a $9,499 Mac Studio M5 Ultra 256GB, and the low active count means it should decode unusually fast for its size.

Wait for the weightsMac Studio M5 Ultra 256GBHosted API beta
Best forAlmost everyoneAlready planning a 200GB-class quant boxTesting Beam this week
Price / Cost$0from $9,499Unpriced — no rate card yet
The catchWeights could slip past OctoberEvery tok/s figure is still an estimateWaitlist, 262K beta context cap

Honest take: Don’t buy anything for Beam yet — but if you already own or planned a 256GB Mac, Beam is the most interesting thing that will land on it this year: GLM 5.2-class coding scores at roughly half the memory footprint and half the per-token traffic.

Reflection AI — the startup founded in March 2024 by Misha Laskin (reward modeling lead on DeepMind’s Gemini) and Ioannis Antonoglou (AlphaGo co-developer), backed by a $2 billion round that NVIDIA led with $800 million — announced its first model on October 5, 2026. Beam is pitched, bluntly, as the American answer to DeepSeek and Qwen: an open-weight frontier model from a US lab, Apache 2.0, built for coding and agentic work.

The announcement went to the top of Hacker News (321 points), and most of the coverage framed it as “the largest open-weight model ever released.” Before you plan hardware around that, three corrections — because the real spec sheet changes the math in your favor.

What was actually announced (and what wasn’t)

Beam is a sparse MoE, not a 501B dense model. 501B total parameters, 23B active per token. Reflection hasn’t published the expert count or routing details yet — those are promised with the technical report. The distinction matters in opposite directions: total parameters decide whether the model fits in your memory, active parameters decide how fast it decodes once it does. We walked through that split in Why Local LLMs Got Good in 2026, and Beam is the most extreme example yet from a Western lab — only 4.6% of the model computes per token.

It is not the largest open-weight model. Kimi K3’s weights have been live since July at 2.8T parameters and a 1.56TB download — 5.6× Beam’s size. DeepSeek R1 is 671B. What Beam actually is: the largest open-weight release from a US lab to date, and the first US open-weight model that trades blows with the Chinese frontier on coding benchmarks. If provenance matters to your compliance department (it increasingly does for government and defense-adjacent work, which is explicitly Reflection’s market), that’s the story. Your GPU doesn’t care about provenance.

The weights are not out. As of October 8, access is a hosted API beta behind a waitlist. Reflection says the weights, model card, technical report, and — unusually — its internal safety evaluations all ship “later in October” under Apache 2.0. The company also says Beam is still in final red-teaming, which is the kind of sentence that precedes schedule slips. Treat late October as a plan, not a date. We said the same about Mistral Large 4 two days ago; October is turning into open-weight release season, with both promised inside the same two-week window.

There is no API pricing. Reflection’s developer docs had no per-token rate card as of October 6, and Artificial Analysis has no entry for it. One outlet circulated $0.15/$0.60 per million tokens — those are DeepSeek V4.1 Flash’s prices, pasted onto the wrong model. Ignore them.

One real limitation worth knowing before you join the beta: Beam claims a 1M-token context window, but the hosted beta currently caps combined input/output at 262,144 tokens, with generation capped at 131,072 inside that. If you plan to throw a whole repository at it this week, you’ll hit the 262K wall, not the advertised 1M. The fix for now is the boring one — chunk your context or wait for the weights, where the full window is yours to fund with your own KV cache memory.

The vendor benchmarks, flagged as vendor benchmarks

Every number below is Reflection’s own table. No independent lab has reproduced any of them yet, and GLM 5.2’s row is blank (“not reported”) on several benchmarks, so the head-to-head covers only where both sides posted scores.

BenchmarkBeamGLM 5.2Current open-weight leaders
SWE-bench Verified80.9NRInkling 77.6, Nemotron 3 Ultra 70.7
SWE-bench Pro v165.562.1—
DeepSWE v1.144.444.0—
Terminal-Bench v2.180.181.0DeepSeek V4.1 Flash 90.6, Kimi K3 88.3, GLM-5.3 88.2

Two honest readings. First: Beam lands at GLM 5.2 parity on coding — level or slightly ahead on the SWE-bench family, a point behind on Terminal-Bench. Reflection’s actual pitch is efficiency, not peak scores: comparable reasoning results at a claimed 3–4× less inference compute, which is what a 23B-active design buys versus GLM 5.2’s 40B active. Second: GLM 5.2 is no longer the frontier. GLM-5.3, Kimi K3, and DeepSeek V4.1 Flash all clear Beam by 8+ points on Terminal-Bench. Beam matches the best open model of mid-2026, at the end of 2026, with less compute. That’s genuinely useful — it’s just not the frontier-topping debut the headlines implied.

The training recipe explains the shape: 23.8 trillion curated tokens, then a multi-teacher on-policy distillation step combining a large RL teacher with a safety/alignment teacher. Beam is a distilled model by design. Keep that word in mind — it comes back when we talk about aggressive quantization.

The GGUF math: what 501B/23B actually needs

No GGUFs exist, because no weights exist. But this is the fourth 500B-plus MoE we’ve sized this year, and the extrapolation method has held within ±10% every time: scale from Kimi K2.7’s community quants (1T parameters → ~325GB at 2-bit UD-Q2_K_XL, ~605GB at Q4-class), with DeepSeek R1’s IQ1_S (671B → 131GiB) as the second anchor for 1-bit-class builds.

At 0.501× K2.7’s parameter count:

Quant classEst. size on diskWhat holds it (weights + context)
~1-bit dynamic (IQ1/TQ class)~85–105GB128GB unified-memory boxes — barely, maybe
~2-bit (UD-Q2_K_XL class)~163GBMac Studio M5 Ultra 256GB; 192GB Framework (tight to no)
~4-bit (Q4_K_M class)~303GBM5 Ultra 512GB (late Oct), 384GB+ EPYC RAM build
BF16~1,000GBNobody’s house

Treat these as ±10% until Unsloth or Bartowski publish real files. The interesting rows are the first two.

The 2-bit build fits a machine you can order today. 163GB of weights plus KV cache sits comfortably inside the Mac Studio M5 Ultra 256GB (from $9,499, the tier we covered when pre-orders opened). That makes Beam materially more accessible than Mistral Large 4, whose smallest useful quant (~340GB) needs the 512GB tier that still isn’t orderable, and than GLM 5.2, whose 241GB 2-bit quant squeezes a 256GB Mac with almost no context headroom. Half the total parameters is a practical difference, not a spec-sheet one.

The 1-bit row is the wildcard. ~85–105GB would be the first time a 500B-class model even theoretically lands on the 128GB class — Strix Halo boxes with their ~120GB Linux GTT pool, DGX Spark, the 192GB Framework Desktop with room to spare. Two loud caveats before anyone orders a 128GB machine on this basis: 1-bit dynamic quants are the most quality-destructive tier there is, and Beam is a distilled 23B-active model — the usual rule of thumb is that distilled models tolerate aggressive quantization worse than their over-parameterized teachers, because the redundancy that quantization eats has already been distilled away. Nobody can test that claim until the weights drop. Possible, unproven, do not pre-spend on it.

Speed: why 23B active is the headline

Decode speed is memory-bandwidth-bound: the ceiling is bandwidth divided by bytes read per token, and a MoE reads only its active experts. At Q4-class quantization, 23B active ≈ 13.7GB per token (same bytes-per-parameter as R1’s measured 22GB/token at 37B active); at 2-bit, roughly 7.5GB.

$ python3 - <<'EOF'
bw = {"M5 Ultra (1,200 GB/s)": 1200, "M3 Ultra (819 GB/s)": 819,
      "EPYC 12ch DDR5 (~460 GB/s)": 460, "Strix Halo (256 GB/s)": 256}
for label, gb_per_tok in [("Q4-class, 13.7 GB/token", 13.7), ("2-bit, ~7.5 GB/token", 7.5)]:
    print(label)
    for hw, b in bw.items():
        print(f"  {hw:28s} ceiling {b/gb_per_tok:5.1f} tok/s")
EOF
Q4-class, 13.7 GB/token
  M5 Ultra (1,200 GB/s)        ceiling  87.6 tok/s
  M3 Ultra (819 GB/s)          ceiling  59.8 tok/s
  EPYC 12ch DDR5 (~460 GB/s)   ceiling  33.6 tok/s
  Strix Halo (256 GB/s)        ceiling  18.7 tok/s
2-bit, ~7.5 GB/token
  M5 Ultra (1,200 GB/s)        ceiling 160.0 tok/s
  M3 Ultra (819 GB/s)          ceiling 109.2 tok/s
  EPYC 12ch DDR5 (~460 GB/s)   ceiling  61.3 tok/s
  Strix Halo (256 GB/s)        ceiling  34.1 tok/s

Real decode lands at 40–70% of ceiling for big MoE models. The best measured anchor we have: DeepSeek R1 671B (37B active, Q4) runs 17–18 tok/s on the M3 Ultra’s 819 GB/s — about 45% of its 39 tok/s ceiling (site canon). Applying that fraction:

MachineQuantEst. decodeStatus
M5 Ultra 256GB2-bit (~163GB)~55–75 tok/sEstimate — no one has run this
M5 Ultra 512GBQ4 (~303GB)~35–50 tok/sEstimate; machine ships late Oct
Used M3 Ultra 256GB2-bit~40–55 tok/sEstimate
EPYC 384GB+ DDR5Q4~11–15 tok/sScaled from K2.7’s measured 8–11 tok/s at 32B active
Strix Halo 128GB1-bit (if it fits)~15–25 tok/sDoubly speculative: size and quality both unproven

Every row says estimate because every row is one. But notice what the 23B active count does: Mistral Large 4 (49B active) pencils out at 20–30 tok/s on the 512GB M5 Ultra; Beam at 2-bit on the 256GB tier pencils out at twice that, on a machine that costs less and is orderable today. Per dollar of hardware, Beam is the most runnable frontier-class model announced this year — pending one giant asterisk over how a distilled model survives 2-bit quantization.

What there is no path to: consumer GPU stacks. Four used RTX 3090s give you 96GB — under even the optimistic 1-bit estimate once KV cache joins, and MoE expert offload to system RAM drops you into the 3–9 tok/s crawl we measured on GLM 5.2’s offload path. The DRAM crisis killed the cheap big-RAM alternative too: 64GB DDR5 RDIMMs at $900+ put a 384GB EPYC build north of $5,000 in memory alone (the full breakdown hasn’t gotten cheaper since September).

What to actually buy

Prices as of October 2026, all verified in the sections above:

Your situationThe machinePriceWhere
Want Beam-class coding on hardware you own todayUsed RTX 3090 24GB running GLM-5.3 Flash / Qwen3.6$1,150–$1,350Check price
Already building a 200GB-class quant box this quarterMac Studio M5 Ultra 256GBfrom $9,499Check price
Want the Q4 build when weights landMac Studio M5 Ultra 512GBlate Oct, price TBDNot orderable — wait
Want to try Beam before spending anythingRented GPU pod, by the hourpay per hourRunPod

The first row is the one most readers should sit in. A 501B model at 2-bit is not what you buy a first machine for; it’s what people who already own 256GB machines get to play with. If Beam’s scores hold up under independent testing, the practical local-AI win for a 24GB card owner won’t be running Beam — it’ll be the distills and the pressure it puts on GLM and Qwen pricing. For self-hosting the weights when they land, the serving stack will almost certainly be vLLM day one (our sister site’s vLLM review covers the setup), with llama.cpp GGUF support following at community speed — and if you want it behind an editor, any OpenAI-compatible endpoint slots into Cline or Continue.dev as a BYOK backend.

FAQ

Can I run Beam on a 24GB GPU? No. The smallest plausible quant is ~85–105GB, and that’s the quality-destroying 1-bit tier. A 24GB card holds under 15% of the 2-bit build, and MoE routing makes RAM offload of the rest brutally slow (single-digit tok/s). On 24GB, run GLM-5.3 Flash or Qwen3.6 35B-A3B instead — both post scores in Beam’s neighborhood on the benchmarks that fit in your VRAM budget.

Is Beam the biggest open-weight model ever? No. Kimi K3 (2.8T) and DeepSeek R1 (671B) are both larger and both downloadable today. Beam is the largest open-weight model from a US lab, which matters for compliance-sensitive deployments and not at all for your hardware.

When exactly do the weights drop? Reflection says “later in October 2026,” Apache 2.0, alongside the technical report and its internal safety evals. The company also says final red-teaming is still underway, so the date can slip. We’ll update this guide when files actually appear — the same watch we’re keeping on Mistral Large 4’s October 31 promise.

Is Beam better than GLM 5.2? On Reflection’s own table: even on coding, a hair behind on Terminal-Bench, at a claimed 3–4× less inference compute. No independent verification exists yet, and GLM-5.3 — already downloadable — beats both. Efficiency is the real pitch, and for local use, efficiency (23B active) is exactly the property that makes it runnable at all.

  • Mac Studio M5 Ultra — the 256GB tier is the only orderable machine today that should hold Beam’s 2-bit quant with context headroom.
  • Used RTX 3090 24GB — still the right answer for everyone not buying a $9,499 machine; runs the 24GB-class models that match Beam’s benchmark neighborhood.

Sources

Last updated October 8, 2026. Prices and specs change; verify current rates before purchasing. All Beam performance figures are vendor-reported or bandwidth-ceiling estimates — no independent benchmarks exist until the weights ship.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.