llama.cpp 'Streaming Expert Loading': The Viral 96K-on-24GB Claim vs. the Flags That Already Do It

llama-cppmoegpulocal-llmvramlong-context

TL;DR: A widely recirculated digest claims llama.cpp’s “streaming expert loading” now delivers 96K+ context on 24GB GPUs. We traced it: the feature is an experimental pull request tested on a 48GB unified-memory laptop, whose own author told maintainers not to merge it. Meanwhile, flags that merged back in August 2025 already run a 35B MoE at 131K–262K context on a used RTX 3090.

Used RTX 3090 24GB3× used RTX 3090 (72GB)GMKtec EVO-X2 128GB
Best for30–35B MoE at 131K+ contextgpt-oss-120b at 93K context120B-class MoE, low power, one box
Long-context reality~148–158 tok/s on Qwen3.6-35B-A3B41–73 tok/s decode, 833 tok/s prefill at 94K~31 tok/s on gpt-oss-120b
Price (Sep 2026)~$1,150–$1,350~$3,450–$4,050 in cards alone~$2,199
The catch120B-class models crawl at 1.6 tok/sPlatform cost, 1,000W+, Linux tuningSoldered RAM, no upgrade path

Honest take: Don’t wait for streaming expert loading — nothing has merged, and the PR behind the headline doesn’t even target discrete GPUs. A used RTX 3090 plus three flags that already ship (--fa, KV quantization, --n-cpu-moe) is the real 96K-context-on-24GB story, as long as you match the model size to the card.

The claim has been bouncing around aggregators since September 21: “llama.cpp’s streaming expert loading enables >96K context on 24GB GPUs — a new benchmark for long-context inference on consumer hardware.” One version we found, an auto-generated “AI Infrastructure Digest” published as a GitHub issue in a bot-run tracker repo, goes further: it tells readers to “use --load-mode streaming --gpu-pill for Qwen3.8 MoE models,” claims a ~50% peak-VRAM reduction, and says an 85GB model runs “on a 24GB GPU with NVMe swap.”

Every load-bearing part of that is wrong or garbled. We read the pull request it cites, the two other expert-streaming efforts in the llama.cpp repo, and the community benchmarks for what a 24GB card actually does with long context in September 2026. The gap between the viral version and the real one changes what you should run tonight — and what hardware is worth buying. This is the same pattern we found last month when we chased the “28x AMD Q2_K speedup” back to a PR that never landed; AI-generated digests keep reading benchmark tables out of open PRs and reporting them as shipped features.

Where the 96K claim came from

The digest cites llama.cpp PR #29191, “streams experts from qwen 3.8 moe with gpu prefill on vulkan and hip,” opened September 20, 2026 by contributor Atomic-Germ. The PR is real, and it’s genuinely interesting work: it adds --load-mode streaming and --gpu-pill options that stream Mixture-of-Experts expert weights instead of resident-loading them, and the author demonstrates an ~85GB Q6_K Qwen 3.8 MoE running with a roughly 96,000-token context window.

Now the parts the digest dropped:

  • The test machine is not a 24GB GPU. It’s a Framework 13 laptop with an AMD Ryzen AI 340 and 48GB of unified memory, plus a 24GB NVMe swap file. The “24GB” in the viral claim appears to be the swap file.
  • It only works on unified-memory (UMA) systems. The author says so directly: it “only works on a system with UMA. So, practically, the options are a laptop or a very expensive gpu.” A desktop RTX 3090, 4090, or 5090 — the cards the headline made you think of — is exactly what this PR does not cover.
  • It runs on Vulkan; the author hadn’t tested HIP at the time of writing, and Metal and Intel Arc don’t work. CUDA isn’t mentioned at all.
  • The author asked for it not to be merged. Quoting the PR thread: the code is “probably niche and big to the point that it should not actually be pulled.” It sits open as a proof of concept people can cherry-pick.

So the “new benchmark for consumer long-context inference” is one developer’s unmerged experiment on a 48GB laptop — closer to our Strix Halo unified-memory coverage than to anything a discrete-GPU owner can download.

Three expert-streaming efforts, zero merged

PR #29191 isn’t even the first attempt. The llama.cpp repo currently hosts at least three independent takes on the same idea — keep only hot experts in fast memory, stream the rest — and none has merged:

PR #25294 — “llama: stream MoE routed experts from disk” (freedomljc, opened July 4, 2026, still open). This one targets the opposite extreme: models bigger than your system RAM, streamed from SSD through a fixed per-layer expert cache (--moe-stream, --moe-stream-cache, --moe-stream-direct for O_DIRECT reads). The author’s benchmarks on a ~254GB GLM-5.2 Q2_K_XL tell you what disk streaming costs: with a ~55GB cache, decode runs at ~1.8 tok/s with a 73% cache hit rate; a ~79GB cache gets ~2.2 tok/s at 79% hits. That’s 2.4x faster than the mmap fallback it replaces — and still a reading-speed-or-slower experience. It also currently supports a single context at a time, and the thread documents a 158GB Qwen3.8-Flash-Next quant loading on a 123GB machine, which is the actual headline use case: running a model your RAM can’t hold, slowly, instead of not at all.

Discussion #24528 — “RFC: MoE expert cache” (June 2026, implemented in community forks, not upstream). This is the most promising of the three for people who already run big MoE models split across GPU and CPU: it caches the hottest CPU-resident experts in leftover VRAM. Fork benchmarks on 2×–4× RTX 3090 rigs show +25% decode on a 754B GLM-5.1 (13.96 → 17.49 tok/s) and +46% on DeepSeek-V4-Flash, composing to +64.9% with speculative decoding — at the price of 2–2.3x slower prefill in many configurations, and outright regressions on older cards like the GTX 1080 Ti. Real gains, real trade-offs, fork-only.

PR #29191 — the UMA laptop experiment above.

The pattern matters more than any single number: expert streaming is an active research frontier in llama.cpp, the maintainers haven’t accepted any version of it, and every working prototype trades a lot of speed for capacity. Any article telling you it’s a shipped feature that made 24GB cards better is describing software you cannot download.

What’s actually merged: --cpu-moe and friends

The boring truth is that llama.cpp’s merged answer to “MoE bigger than my VRAM” is over a year old. PR #15077 (merged August 2025) added --cpu-moe and --n-cpu-moe N: keep every always-active tensor — attention, dense FFN, shared experts, KV cache — on the GPU, and park routed expert weights in system RAM, either all of them (--cpu-moe) or the first N layers’ worth (--n-cpu-moe 27). No cache-hit gambling, no SSD in the token path; experts are read from RAM every token, so decode speed becomes a function of your memory bandwidth. Doctor-Shotgun’s MoE offload guide is the canonical walkthrough, including the -ot "exps=CPU" regex form for manual placement.

Tooling support, September 2026:

  • llama.cpp / llama-server: full support, all backends.
  • LM Studio: shipped the same thing as a checkbox — “Force Model Expert Weights onto CPU” — in version 0.3.23 back on August 12, 2025.
  • Ollama: still doesn’t expose it. The feature request (ollama #11772, open since August 7, 2025) sits unimplemented; Ollama’s automatic layer split is the old slow path that expert offload exists to beat. If your MoE model overflows VRAM under Ollama, this is the single biggest reason to run llama-server or LM Studio instead — see our Ollama speed guide for what Ollama can and can’t tune.

The real 96K-on-24GB recipe (it already works)

Here’s what the viral claim should have said: a 24GB card already runs a modern 30–35B-class MoE at 96K+ context, using nothing but merged flags. The recipe has three parts — flash attention, KV-cache quantization, and picking a sparse model whose weights leave VRAM room for context.

Qwen3.6-35B-A3B is the demonstration case (35B total parameters, ~3B active — our full guide). Two independent RTX 3090 benchmark writeups converge on the same picture: Amine Raji’s benchmark measured the Unsloth UD-Q4_K_M quant fully resident on a 24GB RTX 3090 at 157.66 tok/s, and a Japanese community guide (zephel01) reports Q4_K_M at 148.64 tok/s with a 131K context configured, and IQ4_NL at 149.81 tok/s with 262K — the model’s full native window — on the same card. The launch command looks like this:

$ ./llama-server -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \
    -ngl 999 -fa on -c 131072 -ctk q8_0 -ctv q8_0
...
llama_context: n_ctx = 131072
llama_context: KV cache size, offloaded to GPU

Two of those flags do the heavy lifting. -fa on (flash attention) is mandatory — Hardware Corner’s large-context testing found big-context loads “simply fail” on 24GB-class hardware without it. -ctk q8_0 -ctv q8_0 quantizes the KV cache to 8-bit, which halves its footprint at negligible quality cost; that halving is exactly the difference between 131K fitting alongside an ~19GB model file and not fitting. When you still come up short, the escalation order is: reduce context, quantize KV, then --n-cpu-moe a few layers — each step trades speed you have for VRAM you don’t. Run your exact model-plus-context arithmetic in our VRAM calculator before downloading 20GB of weights.

And the honest ceiling

What a 24GB card does not do is run a 120B-class MoE quickly, no matter which loading trick you use. Hardware Corner tested gpt-oss-120b (~60GB of MXFP4 weights) on a single RTX 3090 with an EPYC 7343 and 64GB of DDR4-3200: the default layer split fit 13 of 32 layers and generated at 0.90 tok/s; switching to -ngl 99 --n-cpu-moe 27 — the merged, correct approach — nearly doubled it to 1.6 tok/s average. That’s the fix working as designed, and it’s still far below usable. The bottleneck isn’t the flag; it’s DDR4’s ~51GB/s feeding 5.1B active parameters per token. Faster dual-channel DDR5 roughly doubles that arithmetic, but no consumer desktop RAM turns a CPU-offloaded 120B into an interactive experience.

Getting gpt-oss-120b genuinely fast at long context took Hardware Corner three RTX 3090s: 41–73 tok/s decode, over 1,100 tok/s prefill at small contexts and 833 tok/s at ~94K — with 93K tokens the max context that fit in 72GB, flash attention required. That 93K figure is probably the origin of the digest’s “96K on 24GB” — right neighborhood, wrong number of GPUs. If that’s the class of rig you’re considering, our multi-GPU guide and Threadripper platform breakdown cover the PCIe and power realities.

The RAM tax nobody mentions

Every CPU-offload and streaming technique in this article shifts cost from VRAM to system RAM — in the middle of a DRAM price crisis. Offloading gpt-oss-120b’s experts wants 64GB of system memory, and a DDR5-6000 64GB kit runs $680–$1,070 in September 2026 (Tom’s Hardware RAM price index, September 11) — several times its mid-2025 price, and potentially more than the used GPU you’re trying to spare. A 32GB kit at $399–$479 covers the 35B-MoE-fully-in-VRAM recipe above with room to spare, which is one more reason the “sparse 35B on one 24GB card” configuration is the sweet spot of 2026: it’s the setup that doesn’t pay the RAM tax.

Does any of this change what to buy?

No — and that’s the finding. If streaming expert loading had actually shipped for discrete GPUs, it would strengthen the case for a cheap 24GB card over a $4,300+ RTX 5090. It hasn’t, but the merged flags already deliver the useful version of that promise: our 24GB VRAM model guide stands, with the used RTX 3090 at $1,150–$1,350 still the cheapest seat for 96K-plus context on models people actually run daily. The 262K-context Qwen3.6-35B-A3B setup is also a legitimate full-codebase coding backend — aicoderscope.com covers wiring a local endpoint into Cline and Cursor.

For 120B-class ambitions, the choice is unchanged from our $3K/$6K/$10K build breakdown: stack used 3090s for speed, or buy unified memory for capacity-per-dollar-per-watt — a GMKtec EVO-X2 128GB runs gpt-oss-120b at ~31 tok/s in one quiet box at ~$2,199. Ironically, that Strix Halo machine is a UMA system — the one hardware class PR #29191 actually targets, if it ever matures.

What to actually buy

Prices as of September 2026, all verified in the comparison above:

Your situationThe machinePriceWhere
Want 96K–262K context on 30–35B MoE models tonightUsed RTX 3090 24GB~$1,150–$1,350Check price
Want gpt-oss-120b fast at ~93K context3× used RTX 3090 + workstation platform~$3,450–$4,050 in cardsCheck price
Want 120B-class capacity, one quiet low-power boxGMKtec EVO-X2 128GB~$3,649Check price
Undecided — test your model + context firstRented RTX 3090from $0.07/hrVast.ai

FAQ

Is streaming expert loading in llama.cpp right now? No. Three separate efforts exist — PR #29191 (UMA laptops, Vulkan), PR #25294 (SSD streaming for models bigger than RAM), and the expert-cache RFC #24528 (fork-only) — and none has merged as of late September 2026. --cpu-moe/--n-cpu-moe, merged in August 2025, is the shipped tool for MoE models that overflow VRAM.

Can a 24GB GPU really run 96K context? Yes, today, for the right model. Qwen3.6-35B-A3B at Q4 fits fully on an RTX 3090 and community benchmarks run it at ~148–158 tok/s with 131K–262K context using flash attention and q8_0 KV quantization. What 24GB can’t do is run a 120B-class MoE quickly — expert offload gets gpt-oss-120b from 0.9 to only 1.6 tok/s on one 3090 with DDR4.

Do the --load-mode streaming and --gpu-pill flags work in my llama.cpp build? No — they exist only in the unmerged PR #29191 branch, only help on unified-memory systems, and the author has said the PR shouldn’t be pulled. A stock build will reject both flags.

Does Ollama support expert CPU offload? Not as of September 2026 — the request (issue #11772) has been open since August 2025. Use llama-server directly or LM Studio 0.3.23+ (“Force Model Expert Weights onto CPU”) to get --cpu-moe behavior.

Should I wait for expert streaming to merge before buying a GPU? No. Every prototype trades speed for capacity — disk streaming benchmarks at ~2 tok/s on huge models, and the VRAM expert cache gains 25–46% only on rigs that already hold the model in RAM. None of it changes the buying math the way bandwidth and VRAM capacity do. Buy for what merged software does today.

  • RTX 3090 24GB — the cheapest card that runs a 35B-class MoE at 131K+ context fully in VRAM
  • GMKtec EVO-X2 128GB — unified-memory route to 120B-class models, and the hardware class expert streaming actually targets

Sources

Last updated September 25, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.