The Viral 28x AMD Q2_K Speedup: What Actually Shipped in llama.cpp ROCm, and What It Changes for GPU Buyers

amdrocmllama-cppgpuquantizationlocal-llm

TL;DR: A widely shared article claims llama.cpp gave AMD ROCm a 28x Q2_K speedup and 15% faster prompt processing. We traced every ROCm performance PR in the repo: the merged RDNA 4 kernel work is real (2.6-8x faster prefill, batch-dependent), but the PR that matches the Q2_K headline was closed without merging. AMD buying math improves, but not 28x worth.

Used RX 7900 XTXUsed RTX 3090RX 9070 XT (new)
Best forCheapest 24GB, Linux-only stack24GB that works everywhereRDNA 4 kernels + Windows ROCm
VRAM / bandwidth24GB / 960 GB/s24GB / 936 GB/s16GB / 640 GB/s
Price (Sep 2026)~$815–$949~$1,150–$1,350~$740–$900 (MSRP $599)
The catchNo Windows ROCm, no CUDA~$350 premium, 3+ years old16GB ceiling blocks 27B at Q4

Honest take: Don’t buy hardware because of a headline multiplier. The used RTX 3090 is still the default 24GB pick, but at $815–$949 a used RX 7900 XTX is now roughly $350–$450 cheaper for the same VRAM — a real value play if you run Linux and stay inside llama.cpp.

The claim spread the way GPU claims usually do: a blog post — “llama.cpp Breakthrough for AMD ROCm: 15% Prompt Processing Boost and 28x Speedup for Q2_K Quantization,” published on baguaai.com on July 21, 2026 — got picked up, resummarized, and recirculated through r/LocalLLaMA well into September. The framing is irresistible: AMD cards were “barely functional” for a whole quantization class, one pull request fixed it, and suddenly the cheap 24GB Radeon is a giant-killer.

Numbers like that change buying decisions, so before repeating them we did what the original article didn’t show: went through llama.cpp’s actual pull requests and benchmark threads to find where the 28x lives. Short version — we couldn’t find it in anything that merged. Here’s what’s really in the tree, what almost made it, and what the honest AMD-vs-NVIDIA math looks like in September 2026.

Where the Q2_K problem came from

The underlying complaint is real and old. Issue #11931, opened February 17, 2025, documented that ROCm underperformed across most quantization formats relative to the hardware’s theoretical capability, with Q2 quants called out as “very slow” — at the time, on Vega-class cards like the MI60 and Radeon VII. K-quants in general lagged because the matrix-multiply kernels (MMQ) that llama.cpp uses during prompt processing were tuned for NVIDIA’s architectures and fell back to slow paths on AMD.

That matters most for prefill — the phase where your prompt gets ingested. Token generation (decode) is memory-bandwidth-bound, so a 960 GB/s card generates tokens at a competitive rate even with mediocre kernels. Prefill is compute-bound, and that’s where AMD was leaving multiples on the table. If you feed a 27B model a 4,000-token prompt and wait 40 seconds before the first token appears, that’s the Q2_K/MMQ problem in practice.

What actually merged

Three changes account for nearly all of the real ROCm progress in the current llama.cpp tree:

PR #17156 — “HIP: WMMA-MMQ kernels for RDNA 4” (merged November 24, 2025). This is the substantial one. It rewrote the quantized matrix-multiply path to use RDNA 4’s WMMA instructions, covering Q2_K through Q8_0 plus the IQ formats. The author’s benchmarks on an RX 9060 XT running an 8B model show the gains are batch-dependent, which is exactly what you’d expect from a prefill kernel:

  • Batch 1–8 (interactive decode): ~1.00x — no change
  • Batch 16–32: 1.17–1.25x
  • Batch 64–2048 (prompt processing): 6.5–8.2x on Q4_0, up to 7.1x on Q8_0, 3.5–5.2x on Q6_K
  • Q2_K specifically: 2.58–2.73x at high batch

So the biggest merged Q2_K improvement on record is about 2.7x, on RDNA 4 only, during prefill only.

PR #24127 — “CUDA: refactor MMQ kernel configuration” (merged July 13, 2026). A structural refactor replacing macro spaghetti with a parameter table. Performance-neutral on NVIDIA (0.99–1.01x across the board), mixed on AMD — a few configurations like Q4_0 on the Ryzen AI 8060S iGPU gained ~3x, while MI100 lost 8–14% on some legacy quants. This PR matters mostly because it’s the foundation later tuning work has to build on.

PR #26046 — “HIP: remove rocWMMA FlashAttention” (merged August 3, 2026). Housekeeping that removed the old rocWMMA FlashAttention path. If you’re following a 2025-era build guide that passes -DGGML_HIP_ROCWMMA_FATTN=ON, current builds no longer want it.

The PR that matches the headline never landed

The closest thing to the viral claim is PR #25587 — “HIP: Tune RDNA4 MMQ tiles for Q6_K / Q2_K (gfx1201),” opened July 12, 2026, nine days before the baguaai article. Its numbers, measured on a Radeon AI PRO R9700:

  • Q2_K prefill throughput: 1.93 → 6.50 TFLOPS (+237%)
  • Qwen3.6-27B at Q2_K: pp512 87.21 → 263.99 tok/s (+203%)
  • Qwen3.6-27B at Q6_K: pp512 705.83 → 888.79 tok/s (+26%)

Note two things. First, the shape of the claim matches the headline — a big Q2_K multiplier plus a double-digit prefill gain on other quants — but the magnitudes are ~3x and ~26%, not 28x and 15%. We could not trace either headline number to any merged change, this PR, or any benchmark table in the repo. Second, and more important for anyone deciding what to buy: the PR was closed pending a rebase onto #24127 and was never resubmitted. As of llama.cpp’s tree in late September 2026, this tuning is not something you can download.

That’s the pattern worth internalizing: the 87 tok/s prefill figure for a 27B at Q2_K on a $1,300 RDNA 4 workstation card is the current shipping reality. The 264 tok/s figure is a proposal that stalled. An article written from the PR’s benchmark table describes hardware behavior nobody can reproduce with a stock build.

Q2_K was the wrong quant for most people anyway

Even if the tuning had merged, the framing — “run ultra-large models on cheap AMD VRAM” — skips the reason Q2_K is a niche format. llama.cpp’s own quantization quality table puts Q2_K at +0.87 perplexity over the F16 baseline on a 7B model, flagged as extreme quality loss and explicitly not recommended. Q4_K_M’s penalty on the same scale is +0.05 — roughly 16x less degradation. Our quantization quality guide covers why even the Q4-to-Q5 gap shows up in multi-step reasoning; Q2 is a different category of compromise, and it hits chain-of-thought tasks hardest.

Where Q2_K genuinely earns its place is the capacity edge case. Unsloth’s Qwen3.8-27B GGUFs make the math concrete: the UD-Q2_K_XL file is 10.7GB, versus roughly 17GB for Q4_K_M. That’s the difference between a 27B model fitting on a 12GB or 16GB card at all versus not loading. If you own an RX 7800 XT 16GB or an RX 9070 XT, a Q2 quant of a 27B may still beat a Q4 quant of an 8B for some tasks — a real trade, but a “make it fit” tool, not a performance strategy. Run your exact model-plus-context math in our VRAM calculator before assuming either way.

What AMD actually delivers today

The canonical numbers in llama.cpp’s ROCm performance thread for the RX 7900 XTX on the standard 7B Q4_0 llama-bench workload, current builds:

$ ./llama-bench -m llama-7b-q4_0.gguf -fa 1
| model          |     backend |  test |            t/s |
| llama 7B Q4_0  |  ROCm (HIP) | pp512 | 3874.25 ± 11.92|
| llama 7B Q4_0  |  ROCm (HIP) | tg128 |  170.12 ± 0.56 |

Without flash attention those drop to about 3,552 and 167 tok/s. For calibration: 170 tok/s decode on a 7B is entirely usable — the same thread and our own ROCm 7.2 testing put RDNA 3 at roughly 1.5x behind NVIDIA per dollar on raw inference, driven mostly by prefill and software friction, not decode.

The software situation, September 2026:

  • ROCm 7.2 (released January 21, 2026) officially supports RDNA 4 (RX 9070/9070 XT, gfx1200/gfx1201) on both Linux and Windows. RDNA 3 — the RX 7900 XTX included — remains Linux-only for ROCm. On Windows a 7900 XTX runs llama.cpp through Vulkan instead; our RDNA 4 Vulkan vs ROCm benchmark covers how those backends compare.
  • Building from source is two lines on Ubuntu with ROCm installed — note the flag is -DGGML_HIP=ON now, not the old -DGGML_HIPBLAS=ON that most 2024–2025 tutorials still show:
HIPCXX="$(hipconfig -l)/clang" HIP_PATH="$(hipconfig -R)" \
cmake -S . -B build -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1100 -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j$(nproc)

(Use gfx1201/gfx1200 for RDNA 4, gfx1100 for the 7900 series.) Our ROCm Ubuntu setup guide walks the full stack.

  • If you don’t want to compile: AMD’s lemonade-sdk publishes nightly llama.cpp builds with ROCm 7 bundled for Windows and Ubuntu, covering gfx110X (7900 XTX/XT/GRE, 7800 XT, 7700, 7600) and gfx120X (9070 series) — no separate ROCm install. That’s the fastest way to get the current merged kernels, including the #17156 RDNA 4 work. The broader Lemonade server builds on the same binaries.
  • Ollama and LM Studio bundle their own llama.cpp snapshots and trail upstream by weeks to months. If you’re benchmarking kernel changes, benchmark llama.cpp directly; a wrapper’s version lag will otherwise masquerade as “the speedup didn’t work."

"I updated and nothing got faster” — reading your own results

The most common failure mode after news like this: someone with an RX 7900 XTX pulls the latest build, runs their usual chat workload, and sees identical tokens per second. Three separate reasons, all by design:

  1. The merged gains are RDNA 4 kernels. PR #17156 targets gfx120x WMMA instructions. A 7900 XTX (RDNA 3, gfx1100) doesn’t execute those paths. There is no merged equivalent uplift for RDNA 3.
  2. The gains are prefill, not decode. At batch sizes 1–8 — which is what an interactive chat session mostly is — even the RDNA 4 tables show 1.00x. You’ll feel the improvement in time-to-first-token on long prompts and in batch/server workloads, not in the streaming speed you watch.
  3. The Q2_K-specific tuning never merged. If your test case was literally “Q2_K got 28x faster,” there is nothing to observe.

To verify what your build actually does, run llama-bench at both ends: -p 2048 -n 128 and compare pp against tg across two builds. If pp512 moves and tg128 doesn’t, the kernels are working exactly as shipped.

Does this change which GPU to buy?

The queue premise we started from asked whether a used RX 7900 XTX at “$550” versus a used RTX 3090 at “$1,050” becomes the value pick. Both of those prices are stale, which changes the shape of the answer.

September 2026 street prices: used RX 7900 XTX listings on eBay run around $815, with BestValueGPU’s index at $949 used — the card has held value unusually well against its $999 launch MSRP because 24GB of anything is scarce. Used RTX 3090 cards run $1,150–$1,350 in the same trackers, pushed up all year by the DRAM squeeze. A new RX 9070 XT sits at $740–$900 against its $599 MSRP.

So the AMD discount for equal VRAM is real — about $350–$450 — and it did not come from a kernel breakthrough. What the actual merged work changes is confidence, not headline speed: ROCm 7.2 made the Linux stack officially supported and boring, prefill on RDNA 4 improved multiples at batch, and AMD engineers are visibly contributing upstream. The remaining honest gaps: RDNA 3 has no Windows ROCm, no CUDA means friction with everything outside the llama.cpp/vLLM path (fine-tuning, ComfyUI custom nodes, whisper variants), and the 1.5x-per-dollar inference gap on RDNA 3 hasn’t closed.

If your workload is “llama.cpp or Ollama on Linux, chat and coding models up to 32B” — the 7900 XTX at $815–$949 is the cheapest working 24GB seat in the house, and the money saved covers a lot of electricity. If you touch anything CUDA-shaped even occasionally, the 3090 premium keeps buying its usual insurance. Our GPU buying guide and used RTX 3090 deep dive cover both cards’ full profiles; for coding-assistant backends specifically (Cline, Continue.dev with a local endpoint), the cheaper AMD card is a legitimate BYOK play — see aicoderscope.com for the tool-side setup.

What to actually buy

Prices as of September 2026, all verified in the comparison above:

Your situationThe cardPriceWhere
Linux home lab, llama.cpp/Ollama only — cheapest 24GBUsed RX 7900 XTX~$815–$949Check price
Need CUDA compatibility or Windows — the safe 24GBUsed RTX 3090~$1,150–$1,350Check price
Want the RDNA 4 kernels + official Windows ROCmRX 9070 XT 16GB~$740–$900Check price
Undecided — test your model on rented 24GB firstRented RTX 3090from $0.07/hrVast.ai

FAQ

Did llama.cpp really make Q2_K 28x faster on AMD? Not in anything you can download. The merged RDNA 4 kernel work (PR #17156, November 2025) improved Q2_K prefill about 2.6–2.7x at high batch. A July 2026 PR (#25587) showed ~3x Q2_K prefill gains on the R9700 but was closed without merging. We could not trace the 28x figure to any merged change or benchmark table in the repository.

Does the RX 7900 XTX benefit from the new kernels? Mostly no. The WMMA-MMQ work targets RDNA 4 (RX 9070 series, R9700). The 7900 XTX still runs well — ~170 tok/s decode and ~3,874 tok/s prefill on a 7B Q4_0 with flash attention — but its numbers haven’t jumped in 2026.

Should I run 27B models at Q2_K on a 12–16GB card? Only when nothing else fits. Qwen3.8-27B’s Q2_K_XL is 10.7GB versus ~17GB at Q4_K_M, but Q2_K carries a +0.87 perplexity penalty on the llama.cpp quality table — the “extreme quality loss” tier. Try the 27B at Q2 against a strong 8–14B at Q4 on your actual tasks before committing.

Is prompt processing or generation speed what I’d actually notice? Prefill improvements show up as shorter waits before the first token on long prompts (RAG, long documents, agent context). Streaming speed while the model writes is decode, which none of these changes touched at interactive batch sizes.

What’s the fastest way to get current ROCm kernels without compiling? The lemonade-sdk llamacpp-rocm nightly builds bundle ROCm 7 for RDNA 3 and RDNA 4 on Windows and Ubuntu. Ollama and LM Studio lag upstream llama.cpp by weeks to months.

Sources

Last updated September 24, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.