Reflection AI's Beam 501B: What Hardware Runs America's First Frontier Open-Weight Model When the Weights Drop
TL;DR: Reflection AI announced Beam on October 5 — a 501B-parameter sparse MoE with 23B active, weights promised under Apache 2.0 later this month. Nothing is downloadable yet. When it lands, a ~163GB 2-bit quant fits a $9,499 Mac Studio M5 Ultra 256GB, and the low active count means it should decode unusually fast for its size.
| Wait for the weights | Mac Studio M5 Ultra 256GB | Hosted API beta | |
|---|---|---|---|
| Best for | Almost everyone | Already planning a 200GB-class quant box | Testing Beam this week |
| Price / Cost | $0 | from $9,499 | Unpriced — no rate card yet |
| The catch | Weights could slip past October | Every tok/s figure is still an estimate | Waitlist, 262K beta context cap |
Honest take: Don’t buy anything for Beam yet — but if you already own or planned a 256GB Mac, Beam is the most interesting thing that will land on it this year: GLM 5.2-class coding scores at roughly half the memory footprint and half the per-token traffic.
Reflection AI — the startup founded in March 2024 by Misha Laskin (reward modeling lead on DeepMind’s Gemini) and Ioannis Antonoglou (AlphaGo co-developer), backed by a $2 billion round that NVIDIA led with $800 million — announced its first model on October 5, 2026. Beam is pitched, bluntly, as the American answer to DeepSeek and Qwen: an open-weight frontier model from a US lab, Apache 2.0, built for coding and agentic work.
The announcement went to the top of Hacker News (321 points), and most of the coverage framed it as “the largest open-weight model ever released.” Before you plan hardware around that, three corrections — because the real spec sheet changes the math in your favor.
What was actually announced (and what wasn’t)
Beam is a sparse MoE, not a 501B dense model. 501B total parameters, 23B active per token. Reflection hasn’t published the expert count or routing details yet — those are promised with the technical report. The distinction matters in opposite directions: total parameters decide whether the model fits in your memory, active parameters decide how fast it decodes once it does. We walked through that split in Why Local LLMs Got Good in 2026, and Beam is the most extreme example yet from a Western lab — only 4.6% of the model computes per token.
It is not the largest open-weight model. Kimi K3’s weights have been live since July at 2.8T parameters and a 1.56TB download — 5.6× Beam’s size. DeepSeek R1 is 671B. What Beam actually is: the largest open-weight release from a US lab to date, and the first US open-weight model that trades blows with the Chinese frontier on coding benchmarks. If provenance matters to your compliance department (it increasingly does for government and defense-adjacent work, which is explicitly Reflection’s market), that’s the story. Your GPU doesn’t care about provenance.
The weights are not out. As of October 8, access is a hosted API beta behind a waitlist. Reflection says the weights, model card, technical report, and — unusually — its internal safety evaluations all ship “later in October” under Apache 2.0. The company also says Beam is still in final red-teaming, which is the kind of sentence that precedes schedule slips. Treat late October as a plan, not a date. We said the same about Mistral Large 4 two days ago; October is turning into open-weight release season, with both promised inside the same two-week window.
There is no API pricing. Reflection’s developer docs had no per-token rate card as of October 6, and Artificial Analysis has no entry for it. One outlet circulated $0.15/$0.60 per million tokens — those are DeepSeek V4.1 Flash’s prices, pasted onto the wrong model. Ignore them.
One real limitation worth knowing before you join the beta: Beam claims a 1M-token context window, but the hosted beta currently caps combined input/output at 262,144 tokens, with generation capped at 131,072 inside that. If you plan to throw a whole repository at it this week, you’ll hit the 262K wall, not the advertised 1M. The fix for now is the boring one — chunk your context or wait for the weights, where the full window is yours to fund with your own KV cache memory.
The vendor benchmarks, flagged as vendor benchmarks
Every number below is Reflection’s own table. No independent lab has reproduced any of them yet, and GLM 5.2’s row is blank (“not reported”) on several benchmarks, so the head-to-head covers only where both sides posted scores.
| Benchmark | Beam | GLM 5.2 | Current open-weight leaders |
|---|---|---|---|
| SWE-bench Verified | 80.9 | NR | Inkling 77.6, Nemotron 3 Ultra 70.7 |
| SWE-bench Pro v1 | 65.5 | 62.1 | — |
| DeepSWE v1.1 | 44.4 | 44.0 | — |
| Terminal-Bench v2.1 | 80.1 | 81.0 | DeepSeek V4.1 Flash 90.6, Kimi K3 88.3, GLM-5.3 88.2 |
Two honest readings. First: Beam lands at GLM 5.2 parity on coding — level or slightly ahead on the SWE-bench family, a point behind on Terminal-Bench. Reflection’s actual pitch is efficiency, not peak scores: comparable reasoning results at a claimed 3–4× less inference compute, which is what a 23B-active design buys versus GLM 5.2’s 40B active. Second: GLM 5.2 is no longer the frontier. GLM-5.3, Kimi K3, and DeepSeek V4.1 Flash all clear Beam by 8+ points on Terminal-Bench. Beam matches the best open model of mid-2026, at the end of 2026, with less compute. That’s genuinely useful — it’s just not the frontier-topping debut the headlines implied.
The training recipe explains the shape: 23.8 trillion curated tokens, then a multi-teacher on-policy distillation step combining a large RL teacher with a safety/alignment teacher. Beam is a distilled model by design. Keep that word in mind — it comes back when we talk about aggressive quantization.
The GGUF math: what 501B/23B actually needs
No GGUFs exist, because no weights exist. But this is the fourth 500B-plus MoE we’ve sized this year, and the extrapolation method has held within ±10% every time: scale from Kimi K2.7’s community quants (1T parameters → ~325GB at 2-bit UD-Q2_K_XL, ~605GB at Q4-class), with DeepSeek R1’s IQ1_S (671B → 131GiB) as the second anchor for 1-bit-class builds.
At 0.501× K2.7’s parameter count:
| Quant class | Est. size on disk | What holds it (weights + context) |
|---|---|---|
| ~1-bit dynamic (IQ1/TQ class) | ~85–105GB | 128GB unified-memory boxes — barely, maybe |
| ~2-bit (UD-Q2_K_XL class) | ~163GB | Mac Studio M5 Ultra 256GB; 192GB Framework (tight to no) |
| ~4-bit (Q4_K_M class) | ~303GB | M5 Ultra 512GB (late Oct), 384GB+ EPYC RAM build |
| BF16 | ~1,000GB | Nobody’s house |
Treat these as ±10% until Unsloth or Bartowski publish real files. The interesting rows are the first two.
The 2-bit build fits a machine you can order today. 163GB of weights plus KV cache sits comfortably inside the Mac Studio M5 Ultra 256GB (from $9,499, the tier we covered when pre-orders opened). That makes Beam materially more accessible than Mistral Large 4, whose smallest useful quant (~340GB) needs the 512GB tier that still isn’t orderable, and than GLM 5.2, whose 241GB 2-bit quant squeezes a 256GB Mac with almost no context headroom. Half the total parameters is a practical difference, not a spec-sheet one.
The 1-bit row is the wildcard. ~85–105GB would be the first time a 500B-class model even theoretically lands on the 128GB class — Strix Halo boxes with their ~120GB Linux GTT pool, DGX Spark, the 192GB Framework Desktop with room to spare. Two loud caveats before anyone orders a 128GB machine on this basis: 1-bit dynamic quants are the most quality-destructive tier there is, and Beam is a distilled 23B-active model — the usual rule of thumb is that distilled models tolerate aggressive quantization worse than their over-parameterized teachers, because the redundancy that quantization eats has already been distilled away. Nobody can test that claim until the weights drop. Possible, unproven, do not pre-spend on it.
Speed: why 23B active is the headline
Decode speed is memory-bandwidth-bound: the ceiling is bandwidth divided by bytes read per token, and a MoE reads only its active experts. At Q4-class quantization, 23B active ≈ 13.7GB per token (same bytes-per-parameter as R1’s measured 22GB/token at 37B active); at 2-bit, roughly 7.5GB.
$ python3 - <<'EOF'
bw = {"M5 Ultra (1,200 GB/s)": 1200, "M3 Ultra (819 GB/s)": 819,
"EPYC 12ch DDR5 (~460 GB/s)": 460, "Strix Halo (256 GB/s)": 256}
for label, gb_per_tok in [("Q4-class, 13.7 GB/token", 13.7), ("2-bit, ~7.5 GB/token", 7.5)]:
print(label)
for hw, b in bw.items():
print(f" {hw:28s} ceiling {b/gb_per_tok:5.1f} tok/s")
EOF
Q4-class, 13.7 GB/token
M5 Ultra (1,200 GB/s) ceiling 87.6 tok/s
M3 Ultra (819 GB/s) ceiling 59.8 tok/s
EPYC 12ch DDR5 (~460 GB/s) ceiling 33.6 tok/s
Strix Halo (256 GB/s) ceiling 18.7 tok/s
2-bit, ~7.5 GB/token
M5 Ultra (1,200 GB/s) ceiling 160.0 tok/s
M3 Ultra (819 GB/s) ceiling 109.2 tok/s
EPYC 12ch DDR5 (~460 GB/s) ceiling 61.3 tok/s
Strix Halo (256 GB/s) ceiling 34.1 tok/s
Real decode lands at 40–70% of ceiling for big MoE models. The best measured anchor we have: DeepSeek R1 671B (37B active, Q4) runs 17–18 tok/s on the M3 Ultra’s 819 GB/s — about 45% of its 39 tok/s ceiling (site canon). Applying that fraction:
| Machine | Quant | Est. decode | Status |
|---|---|---|---|
| M5 Ultra 256GB | 2-bit (~163GB) | ~55–75 tok/s | Estimate — no one has run this |
| M5 Ultra 512GB | Q4 (~303GB) | ~35–50 tok/s | Estimate; machine ships late Oct |
| Used M3 Ultra 256GB | 2-bit | ~40–55 tok/s | Estimate |
| EPYC 384GB+ DDR5 | Q4 | ~11–15 tok/s | Scaled from K2.7’s measured 8–11 tok/s at 32B active |
| Strix Halo 128GB | 1-bit (if it fits) | ~15–25 tok/s | Doubly speculative: size and quality both unproven |
Every row says estimate because every row is one. But notice what the 23B active count does: Mistral Large 4 (49B active) pencils out at 20–30 tok/s on the 512GB M5 Ultra; Beam at 2-bit on the 256GB tier pencils out at twice that, on a machine that costs less and is orderable today. Per dollar of hardware, Beam is the most runnable frontier-class model announced this year — pending one giant asterisk over how a distilled model survives 2-bit quantization.
What there is no path to: consumer GPU stacks. Four used RTX 3090s give you 96GB — under even the optimistic 1-bit estimate once KV cache joins, and MoE expert offload to system RAM drops you into the 3–9 tok/s crawl we measured on GLM 5.2’s offload path. The DRAM crisis killed the cheap big-RAM alternative too: 64GB DDR5 RDIMMs at $900+ put a 384GB EPYC build north of $5,000 in memory alone (the full breakdown hasn’t gotten cheaper since September).
What to actually buy
Prices as of October 2026, all verified in the sections above:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| Want Beam-class coding on hardware you own today | Used RTX 3090 24GB running GLM-5.3 Flash / Qwen3.6 | $1,150–$1,350 | Check price |
| Already building a 200GB-class quant box this quarter | Mac Studio M5 Ultra 256GB | from $9,499 | Check price |
| Want the Q4 build when weights land | Mac Studio M5 Ultra 512GB | late Oct, price TBD | Not orderable — wait |
| Want to try Beam before spending anything | Rented GPU pod, by the hour | pay per hour | RunPod |
The first row is the one most readers should sit in. A 501B model at 2-bit is not what you buy a first machine for; it’s what people who already own 256GB machines get to play with. If Beam’s scores hold up under independent testing, the practical local-AI win for a 24GB card owner won’t be running Beam — it’ll be the distills and the pressure it puts on GLM and Qwen pricing. For self-hosting the weights when they land, the serving stack will almost certainly be vLLM day one (our sister site’s vLLM review covers the setup), with llama.cpp GGUF support following at community speed — and if you want it behind an editor, any OpenAI-compatible endpoint slots into Cline or Continue.dev as a BYOK backend.
FAQ
Can I run Beam on a 24GB GPU? No. The smallest plausible quant is ~85–105GB, and that’s the quality-destroying 1-bit tier. A 24GB card holds under 15% of the 2-bit build, and MoE routing makes RAM offload of the rest brutally slow (single-digit tok/s). On 24GB, run GLM-5.3 Flash or Qwen3.6 35B-A3B instead — both post scores in Beam’s neighborhood on the benchmarks that fit in your VRAM budget.
Is Beam the biggest open-weight model ever? No. Kimi K3 (2.8T) and DeepSeek R1 (671B) are both larger and both downloadable today. Beam is the largest open-weight model from a US lab, which matters for compliance-sensitive deployments and not at all for your hardware.
When exactly do the weights drop? Reflection says “later in October 2026,” Apache 2.0, alongside the technical report and its internal safety evals. The company also says final red-teaming is still underway, so the date can slip. We’ll update this guide when files actually appear — the same watch we’re keeping on Mistral Large 4’s October 31 promise.
Is Beam better than GLM 5.2? On Reflection’s own table: even on coding, a hair behind on Terminal-Bench, at a claimed 3–4× less inference compute. No independent verification exists yet, and GLM-5.3 — already downloadable — beats both. Efficiency is the real pitch, and for local use, efficiency (23B active) is exactly the property that makes it runnable at all.
Recommended Gear
- Mac Studio M5 Ultra — the 256GB tier is the only orderable machine today that should hold Beam’s 2-bit quant with context headroom.
- Used RTX 3090 24GB — still the right answer for everyone not buying a $9,499 machine; runs the 24GB-class models that match Beam’s benchmark neighborhood.
Sources
- Introducing Beam: Reflection’s 501B open-weight model — Reflection AI
- Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute cost — TechCrunch
- Reflection AI unveils Beam — Fortune
- Reflection AI releases first open model to rival China — The Hill
- Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters — MarkTechPost
- Beam: Reflection AI’s 501B Open-Weight Model — DataCamp
- Reflection Beam: Specs, Benchmarks & API Limits — Kingy AI
- Reflection Beam: 501B Total, 23B Active, Weights Not Out — OrcaRouter
- Reflection Beam API beta: the weights are two weeks out — OrcaRouter
- Reflection AI Raises $2B, Nvidia Leads Open Source Push — AI Business
- Reflection Beam: 501B Open-Weight Model, Benchmarks — CellCog
Last updated October 8, 2026. Prices and specs change; verify current rates before purchasing. All Beam performance figures are vendor-reported or bandwidth-ceiling estimates — no independent benchmarks exist until the weights ship.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.