NVIDIA's IFA 2026 llama.cpp Update: Up to 1.9x Faster Local AI — What 24GB RTX Owners Actually Get, and How to Enable It

llama-cppnvidiaollamalm-studioperformancelocal-llm

TL;DR: At IFA 2026 (September 3), NVIDIA announced a round of llama.cpp and vLLM optimizations — CUDA kernel work, enhanced speculative decoding, and faster prefill — worth up to 1.9x throughput, measured on an RTX 5090. The good news for everyone else: most of it runs on Ampere and Ada too, and it ships in stock Ollama 0.34.2 and LM Studio 0.4.24. Update, add two flags for MTP, and a used RTX 3090 picks up real speed for $0.

What you’ll be able to do after this guide:

  • Get the September llama.cpp optimizations in Ollama, LM Studio, or a raw llama.cpp build, and confirm you’re actually on a version that has them
  • Enable multi-token prediction (MTP) with the right flags and the right GGUF — the step most people miss
  • Read a before/after benchmark on your own card instead of trusting NVIDIA’s 5090 marketing number

Honest take: The 1.9x headline is a best-case RTX 5090 number, but this is not a Blackwell-only party trick — community benchmarks show a used RTX 3090 going from 38 to 47–65 tok/s on Qwen3.6-27B. Free speed on a $1,050 used card beats paying $4,300+ for a 5090.

What NVIDIA actually announced at IFA 2026

On September 3 at IFA Berlin, NVIDIA announced a set of local AI improvements built with the llama.cpp and vLLM open-source communities. Strip away the keynote language and there are three separate things in the bag:

1. llama.cpp performance work. New CUDA kernel optimizations, enhanced speculative decoding, and faster prefill, which NVIDIA says add up to 1.9x higher throughput on a GeForce RTX 5090. The gains are model-dependent: NVIDIA’s own numbers are up to a 50% boost on Qwen3.6-27B and 90% on Qwen3.6-35B-A3B on RTX platforms, and up to 35% faster token generation on mixture-of-experts models in llama.cpp (30% in Ollama) per NVIDIA’s developer blog. The 1.9x is the ceiling where all three improvements stack, not the floor.

2. vLLM optimizations. New XQA attention kernels in FlashInfer plus backend work. This matters if you serve concurrent requests; for single-user chat on a consumer card, llama.cpp (and everything built on it) is where you’ll feel the difference. Our vLLM vs Ollama comparison covers when each engine wins.

3. One-click local agents, gated at 24GB. Three agent apps — Hermes Agent, OpenClaw, and Perplexity’s Portable Computer app — get simplified local model setup on Windows, all running llama.cpp under the hood. This is what the “24GB or more of VRAM” line in the coverage actually gates (Wccftech): the app makers picked 24GB as the floor where an agent-grade local model runs well. The llama.cpp kernel and speculative-decoding work itself is not locked to 24GB cards — a 16GB card running a 14B model benefits from the same code.

NVIDIA also teased compact RTX Spark Windows PCs for October; we covered the platform’s 128GB/300 GB/s memory story in our RTX Spark analysis — the short version is that bandwidth, not capacity, still decides tokens per second.

Does your RTX 3090 or 4090 get the speedup?

Yes, with one asterisk. The three ingredients arrived at different times and have different hardware floors:

  • Multi-token prediction (MTP) — the biggest single ingredient — was merged into llama.cpp back on May 16, 2026 (PR #22673) and runs on any CUDA-capable card. It’s a speculative-decoding flavor where the model’s own built-in draft heads propose 2–4 tokens per forward pass and the main model verifies them in one go. Decode is memory-bandwidth-bound, so verifying 3 tokens per weight-read is nearly free speed. Because acceptance-checking scales with bandwidth, faster cards see bigger multipliers — but a 936 GB/s RTX 3090 is exactly the kind of card that benefits.
  • The prefill work includes a chunked CUDA kernel for Gated DeltaNet layers (PR #26001) that needs Ampere or newer with BF16 tensor cores — so RTX 30-series and up qualify, but Turing (RTX 20-series) and older Pascal cards don’t.
  • The September kernel optimizations land as regular llama.cpp CUDA improvements; the peak numbers were tuned and measured on Blackwell (RTX 5090), which is why NVIDIA quotes that card.

Real numbers from the community, not the keynote:

CardBandwidthQwen3.6-27B Q4_K_M, beforeAfter (MTP path)Source
Used RTX 3090 24GB936 GB/s38 tok/s47–65 tok/s (+24% to +71%)Cloudmagazin, Neoteric
RTX 5090 32GB1,792 GB/s63 tok/s84+ tok/scommunity benchmarks via Cloudmagazin
RTX 5090 32GB1,792 GB/s—up to 1.9x combined (NVIDIA)NVIDIA blog

The spread on the 3090 row is honest: the +71% figure (38→65 tok/s) is Cloudmagazin’s measured best case, while Neoteric’s Qwen3 runs landed at +24–33%. Your result depends on the model, quantization, prompt length, and whether the optimized code path applies to your workload at all. Treat anything between +25% and +70% as normal, and the 1.9x as what a 5090 does when kernel work, MTP, and fast prefill all stack on a favorable model.

Step 1: Update to a build that has the optimizations

Ollama — the September llama.cpp update landed in Ollama 0.34.2, published September 15, 2026 (release notes). The desktop app on Windows and macOS auto-updates on restart; on Linux, re-run the install script. Verify with:

ollama -v
# ollama version is 0.34.2

If that prints 0.33.x or older, you’re benchmarking last quarter’s engine. A plain ollama pull of a model does not update the engine — the inference code ships with the Ollama binary itself, so the app update is the step that matters.

LM Studio — 0.4.24 shipped September 9, 2026, and since 0.4.19 the “LM Studio Engine Protocol” delivers llama.cpp engine updates independently of app releases (LM Studio blog). Open the runtimes manager (Ctrl+Shift+R / Cmd+Shift+R), and update the CUDA runtime — the engine version is listed separately from the app version, and an old CUDA engine on a new app silently skips the new kernels.

Raw llama.cpp — grab a release newer than mid-September 2026 from the releases page, or build from source:

cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j

MTP support specifically needs any build after May 16, 2026, but the prefill and kernel work is newer — when in doubt, current release.

Step 2: Enable MTP (the part that isn’t automatic)

This is where most of the “I updated and nothing changed” complaints come from. Two requirements, and both trip people up:

You need the flags. MTP doesn’t switch on by itself in llama.cpp — you ask for it explicitly (DataCamp tutorial):

llama-server -hf ggml-org/Qwen3.6-27B-MTP-GGUF \
  --spec-type draft-mtp \
  --spec-draft-n-max 2 \
  -ngl 99 -fa

--spec-draft-n-max 2 is the conservative setting; 3 can be faster on high-bandwidth cards but wastes work when acceptance drops. Our speculative decoding guide covers the flag family in depth, including the draft-model variant for models without MTP heads.

You need an MTP GGUF. This is the failure mode all over the community threads: running the standard Qwen3.6-27B-Q4_K_M.gguf with --spec-type draft-mtp produces no speedup — same ~38 tok/s as before — because the standard Qwen3.6 quants don’t contain the MTP draft heads at all. The GGUF has to be converted with the heads included. The fix: pull an MTP-converted GGUF — ggml-org/Qwen3.6-27B-MTP-GGUF or the havenoammo MTP quants — after which the same command lands in the 47–65 tok/s range that the benchmark table above shows for a 3090-class card. Per Unsloth’s MTP docs, Qwen3.6 is the family that still needs a separate MTP GGUF; newer model families increasingly ship the heads in the standard file — check the model card for MTP in the metadata before re-downloading 17GB.

In Ollama and LM Studio there are no flags to set: when the bundled engine supports the technique and the model provides the heads, the September builds use the fast path automatically. That’s the trade — convenience, but less visibility into which path you’re on, which is why you should measure.

Step 3: Measure it on your card

Don’t trust any of these numbers — including ours. Run the same prompt through llama-bench before and after updating:

llama-bench -m qwen3.6-27b-mtp-Q4_K_M.gguf -ngl 99 -fa 1

The output table’s t/s column is your truth. In Ollama, ollama run <model> --verbose prints an eval rate: line after each response — same idea, one flag. If your tokens per second didn’t move after the update, work through our Ollama speed checklist — the usual suspects are an old engine, a model too big for VRAM spilling to system RAM, or flash attention off.

What this changes about buying hardware (and what it doesn’t)

The uncomfortable question under every “free speedup” announcement: does this reshuffle which GPU to buy?

Mostly it strengthens the used RTX 3090 case. The site’s standing recommendation for local LLM work is a used 24GB card — around $1,050–$1,343 on the used market as of September 2026 — and this update makes the same card 25–70% faster on supported models without spending a dollar. If the optimizations had been Blackwell-only, the gap to the RTX 5090 (street $4,329+ when you can find one) would have widened; instead, the 5090’s relative advantage is roughly where it was, at a price that’s still ~4x the 3090’s. Our 24GB VRAM model guide and used RTX 3090 breakdown cover what that tier runs; if you’re buying, a used RTX 3090 remains the entry point we’d pick, and if you’d rather test the workload before buying any card, a rented 3090 on Vast.ai runs from $0.07/hr.

What it doesn’t change: the VRAM wall. Speculative decoding makes tokens arrive faster; it doesn’t make a 70B model fit in 24GB. Capacity still decides what you can run — speed optimizations only decide how pleasant it is.

One more angle worth naming: if you use a local model as a coding backend (Cline, Continue.dev — see our sister site aicoderscope.com for those setups), MTP-style gains matter more than in chat, because agentic tools burn thousands of output tokens per task. A 40% decode speedup is the difference between a refactor that takes four minutes and one that takes six. The open-source serving stack side of this — Ollama, vLLM configs — lives at aifoss.dev.

FAQ

Is the 1.9x speedup only for the RTX 5090? The 1.9x figure was measured on an RTX 5090 and is the best case with every optimization stacked. The underlying work — MTP speculative decoding, CUDA kernel improvements — runs on RTX 30-series (Ampere) and newer. Community numbers on a used RTX 3090 show +24% to +71% on Qwen3.6-27B depending on workload.

Do I need to change anything in Ollama or LM Studio? Just update: Ollama 0.34.2+ (September 15, 2026) and LM Studio 0.4.24+ (September 9, 2026) ship engines with the new code. In LM Studio, also update the CUDA runtime in the runtimes manager — the engine updates separately from the app.

Why am I not seeing any speedup after updating? The most common cause: your GGUF doesn’t have MTP draft heads. Standard Qwen3.6 quants don’t include them — you need an MTP-converted GGUF plus --spec-type draft-mtp in raw llama.cpp. Second most common: your model doesn’t hit the optimized code path at all; gains are model- and workload-dependent.

Does this work on a 16GB card like the RTX 5060 Ti or 4060 Ti? The llama.cpp optimizations aren’t gated at 24GB — that floor applies to NVIDIA’s one-click agent apps (Hermes Agent, OpenClaw, Perplexity’s Portable Computer). A 16GB card running a 14B-class model gets the same kernel and MTP benefits.

Should I wait for the RTX Spark PCs instead of buying a GPU? The Spark machines (October) are a capacity play — 128GB of unified memory at ~300 GB/s. For raw tokens per second on models that fit in 24GB, a used RTX 3090 at 936 GB/s stays faster. Different tools for different model sizes.

Sources

Last updated September 21, 2026. Prices and software versions change; verify current rates and releases before purchasing or benchmarking.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.