Ollama 0.34 Turns On Speculative Decoding by Default: Free 2× tok/s on Qwen and Gemma — and One Case Where It Backfires
TL;DR: Ollama’s current Qwen 3.6, Qwen 3.8, and Gemma 4 26B builds ship with speculative decoding switched on in their default parameters — no flag, no separate draft model download. On an RTX 4090 running Ollama 0.34.4, qwen3.6:27b jumps from 48 to 115 tok/s and gemma4:26b from 157 to 294. It’s the largest free speed bump your existing GPU has gotten all year, but one known regression (0.34.2 + qwen3.8:27b) makes generation 4× slower with speculation on, so spend 60 seconds verifying before you trust it.
What you’ll be able to do after this article:
- Confirm whether your models are actually speculating — and measure the real gain on your card with one A/B test using
draft_num_predict. - Recognize the two failure modes (acceptance collapse and bolt-on draft models) that turn “2× faster” into “4× slower,” and fix each one.
- Decide whether this changes what GPU is worth buying — short answer: it makes the card you already own better.
Honest take: This is the rare update where doing nothing gets you most of the benefit — update Ollama, re-pull your models, and dense 27B-class models roughly double. But run the 60-second A/B check anyway: if your tok/s went down after 0.34, you’ve hit the known Qwen 3.8 regression, and the fix is one parameter.
Speculative decoding has been the “free speed if you’re willing to fiddle” trick all year. We covered the llama.cpp setup in July and the DFlash2 drafter for Qwen 3.8 in September — both required picking flags, matching vocabularies, and budgeting VRAM by hand. As of the Ollama 0.34 series (v0.34.0 shipped September 5, 2026; v0.34.4 is current as of September 23), the fiddling is gone for the models most home labs actually run: the defaults do it for you.
What actually changed
Multi-token prediction (MTP) landed in Ollama on Apple Silicon first — the official announcement measured Gemma 4 up to ~90% faster in coding-agent use on MLX, with Ollama auto-tuning the draft length so speculation never becomes a net loss. We covered that phase in our Ollama 0.32 breakdown, when the speedup was Mac-only and CUDA users watched from the sidelines.
The 0.34-era change is that speculation is now in the default parameters of the model builds themselves. Pull the current qwen3.6, qwen3.8, or gemma4:26b from the Ollama registry and the Modelfile carries draft_num_predict 3 (plus repeat_penalty 1, which the release notes say speeds up speculative decoding). That parameter tells Ollama to propose three cheap draft tokens per step and have the full model verify them in a single pass — on any GPU, CUDA included (ComputingForGeeks’ September cheat sheet documents the defaults and the benchmark deltas).
Two details matter:
- There is no separate draft model to download. These families use their own bundled MTP head — small extra tensors inside the model you already pulled.
ollama pull qwen3.6:27bgets everything. This is the same mechanism we explained in why local LLMs got good in mid-2026. - The parameter is inert without the tensors.
draft_num_predictonly does work when the loaded GGUF actually contains MTP tensors (Open WebUI discussion #27479). If you pulled your models months ago, re-pull — an old local copy may predate the MTP head and the new defaults, and it will silently run at the old speed.
The numbers: what “on by default” is worth
The cleanest published comparison comes from ComputingForGeeks, run on Ollama 0.34.4 with an RTX 4090 (24GB) under Ubuntu 24.04 — mean eval rate across five prompts, 256-token outputs, temperature 0, seed 42:
| Model | Speculation off | Speculation on (default) | Gain |
|---|---|---|---|
| qwen3.6:27b | 48 tok/s | 115 tok/s | 2.4× |
| gemma4:26b | 157 tok/s | 294 tok/s | 1.9× |
Setting draft_num_predict to 0 in the request options drops the same card right back to baseline (122.6 → 48.9 tok/s on qwen3.6:27b in their follow-up run) — which is also your verification method, below.
On an RTX 3090, nobody has published the exact same A/B yet, so treat 3090 figures as follows. Measured: an independent MTP test on a 3090 with Qwen 3.6-27B Q4_K_M found coding and stepwise-math workloads at 1.59–1.62× (70–76% draft acceptance) and translation/creative prose at 1.30–1.35× (54–58% acceptance). A lever-by-lever 3090 write-up took Qwen 3.6-27B from 35.7 tok/s (stock Ollama, pre-default era) to roughly 75–80 tok/s with llama.cpp plus MTP. Extrapolating from the site’s usual 3090 ≈ 60–65% of 4090 scaling, expect qwen3.6:27b around 65–75 tok/s on a 3090 with the 0.34 defaults — an estimate until a controlled benchmark lands, but both independent measurements point at the same neighborhood.
The reason the gain varies by workload is acceptance rate. Speculation only pays when the big model agrees with the drafts: at 70%+ acceptance the speedup compounds; below roughly 30%, verification overhead eats the savings and you’re paying to go slower. Code, JSON, and boilerplate accept well. High-entropy creative prose accepts worst.
Verify it in 60 seconds
Don’t assume — measure. Run your model with --verbose and note the eval rate, then kill speculation for one request and compare:
$ ollama run qwen3.6:27b --verbose "Write a Python function that parses ISO 8601 timestamps."
...
eval count: 256 token(s)
eval duration: 2.24s
eval rate: 114.3 tokens/s
Now the same request with speculation off, via the API:
curl -s http://localhost:11434/api/generate -d '{
"model": "qwen3.6:27b",
"prompt": "Write a Python function that parses ISO 8601 timestamps.",
"stream": false,
"options": {"num_predict": 256, "temperature": 0, "draft_num_predict": 0}
}' | python3 -c "import sys,json; d=json.load(sys.stdin); print(round(d['eval_count']/d['eval_duration']*1e9,1),'tok/s')"
Expected result on a healthy setup: the second number is roughly half the first (that’s the speculation you just turned off). In Open WebUI, the same knob is a custom advanced parameter named draft_num_predict — a runtime value there overrides the Modelfile default.
Three outcomes and what they mean:
- First number ≈ 2× the second — speculation is working. Done.
- Both numbers identical — your local model copy has no MTP tensors or predates the new defaults.
ollama pullthe model again and re-test. - First number lower than the second — speculation is actively hurting you. That’s the failure mode below.
When it backfires: the Qwen 3.8 regression and the bolt-on-draft trap
The regression. Ollama issue #18541 documents a severe one: on 0.34.2, qwen3.8:27b Q4_K_M with MTP enabled collapses to ~17.4 tok/s on an RTX 5090 — versus 75.1 tok/s with MTP off, and ~193 tok/s that the same card got on 0.33.2 with MTP working. Draft acceptance falls to about one token per step, meaning essentially every proposed draft gets rejected and you pay full verification cost for nothing. The issue was still open when we checked on September 28, 2026. If your Qwen 3.8 got slower after updating: set draft_num_predict to 0 for that model (/set parameter draft_num_predict 0 in an ollama run session, then /save), or roll back to 0.33.2 until a fix ships. Note this is model-specific — qwen3.6 and gemma4:26b are unaffected in the same reports.
The bolt-on-draft trap. Default-on speculation uses the model’s own MTP head. That is not the same thing as adding a separate small “draft model” yourself, and the difference is brutal. A reproducible RTX 3090 study on Qwen3.6-35B-A3B (UD-Q4_K_XL, llama.cpp) measured it directly:
| Configuration | Throughput | vs baseline | Acceptance |
|---|---|---|---|
| No speculation | 115.7 tok/s | — | — |
| DFlash draft head, n=2 | 146.2 tok/s | +26% | 72% |
| MTP head, n=2 | 141.9 tok/s | +23% | 78% |
| External 0.8B draft model, n=8 | 30.9 tok/s | −73% | 30% |
Purpose-built heads win; a mismatched external drafter can cost you three-quarters of your throughput. If you experimented with manual draft models earlier this year, remove them and let the defaults work. (Also worth knowing: that study’s own audit retracted its viral “MoE makes speculation useless” headline — corrected measurement showed standard acceptance-rate economics, r = +0.998 between acceptance and speedup.)
Mac note. If you run long generations on Apple Silicon, update to at least v0.34.2 — it fixed unbounded memory growth during MLX speculative decoding by releasing freed buffers every 256 tokens. The v0.40 pre-release (September 25) goes further and makes MLX the default engine on Apple Silicon; our MLX upgrade guide covers that stack.
Does this change what GPU you should buy?
It strengthens the card you already own. A used RTX 3090 at $1,150–$1,350 was already our 24GB value pick; at an estimated 65–75 tok/s on a dense 27B it now clears comfortable reading speed with a 7× margin, and the case for spending $3,822–$5,000 on an RTX 5090 for speed alone gets weaker — the 5090’s real arguments remain 32GB of capacity and NVFP4, as we found in the IFA llama.cpp speedup analysis. This is software repricing hardware, and it’s repricing it in favor of the used market.
Prices as of September 2026, all verified this month:
| Your situation | The move | Price | Where |
|---|---|---|---|
| Already own a 24GB card | Update Ollama, re-pull models, run the A/B | $0 | — |
| Buying into local AI now | Used RTX 3090 24GB | $1,150–$1,350 | Check price |
| Want 27B at 100+ tok/s today | Used RTX 4090 24GB | $2,150–$2,350 | Check price |
| Want to test the workload before buying | Rented 3090, from $0.07/hr | pay per hour | Vast.ai |
Check your exact model-plus-context fit in the VRAM calculator first — speculation changes speed, not capacity.
FAQ
Does speculative decoding change the model’s output?
In principle the verification step keeps output equivalent to non-speculative sampling, and tool-call/JSON schema integrity holds identically in testing. One careful 3090 Ti test found output is not always bit-for-bit identical at temperature 0 on free-form prose, so if you diff outputs across runs for evaluation work, pin draft_num_predict 0 for reproducibility.
Which models get the default?
The current registry builds of Qwen 3.6, Qwen 3.8, and Gemma 4 26B ship draft_num_predict 3 by default. On Apple Silicon, the MLX engine additionally applies MTP speculation automatically for supported models (Gemma 4, Qwen 3.5+). Other models are unchanged — the parameter does nothing without MTP tensors in the weights.
Do I need to re-download my models?
If you pulled them before September 2026, yes — ollama pull <model> refreshes both the Modelfile defaults and the weights. An old copy without the MTP head runs at the old speed with no error.
Does it cost VRAM? The MTP head adds a small amount of weight data compared to a whole separate draft model — an RTX 5080 with 16GB ran a 9B model with MTP at +26% in community testing, so mid-size models on 16GB cards still fit. On a card that’s already at 99% VRAM with your target model, watch for offload after re-pulling; our Ollama speed guide shows how to catch silent CPU spill.
My tokens/sec got worse after updating. Why?
Either you hit the qwen3.8:27b acceptance-collapse regression on 0.34.2 (set draft_num_predict 0 for that model or roll back), or you’re running a manually configured external draft model that no longer matches — remove it. Measure with the A/B above rather than guessing.
If you drive these models from a coding agent, the speedup compounds: agents burn thousands of output tokens per task, and code is exactly the high-acceptance workload where speculation shines — see our Continue.dev + Ollama stack and the self-hosted Ollama review at aifoss.dev for the serving side.
Recommended Gear
- RTX 3090 24GB (used) — $1,150–$1,350; the default-on speedup makes the value king faster, not obsolete.
- RTX 4090 24GB (used) — $2,150–$2,350; measured 115 tok/s on qwen3.6:27b with the 0.34 defaults.
Sources
- Release v0.34.2 — ollama/ollama, GitHub
- Ollama releases index (v0.34.0–v0.34.4, v0.40.0 pre-release) — GitHub
- Faster Gemma 4 on MLX with multi-token prediction — Ollama Blog
- Ollama Models Cheat Sheet 2026 (0.34.4 RTX 4090 speculation benchmarks) — ComputingForGeeks
- Ollama 0.34.2 MTP speculative decoding regression with Qwen3.8 27B — Issue #18541, ollama/ollama
- llama.cpp speculative decoding measured on one RTX 3090 (Qwen3.6-35B-A3B) — thc1006, GitHub
- DFlash vs MTP on RTX 3090: I Tested Both Locally — InsiderLLM
- Doubling Qwen3.6-27B on One RTX 3090, Lever by Lever — DEV Community
- Enable Ollama MTP in Open WebUI with draft_num_predict — open-webui Discussion #27479
- Three Months of Speed-Up Experiments on a 3090 Ti (DFlash/MTP, Qwen3.6-27B) — Ian Paterson, DEV Community
- mlx: Gemma4 MTP speculative decoding — PR #15980, ollama/ollama
Last updated September 28, 2026. Prices and software versions change; verify current rates and release notes before purchasing or updating.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.