Best Open-Source LLMs for Your Home Lab, September 2026: The 24GB Crown Changes Hands
TL;DR: Seven open-weight releases landed since our July leaderboard, and exactly one changed what you should run: Qwen3.8-27B is the new default on 24GB cards, and a single llama.cpp flag most people aren’t using lifts it from ~41 to ~66 tok/s on an RTX 3090. The frontier keeps growing (GLM-5.3, Kimi K3, Qwen3.8-Max) and keeps drifting away from clean open licenses. Below 24GB, nothing moved.
| Qwen3.8-27B | GLM-5.3-Flash | GLM-5.3 | |
|---|---|---|---|
| Best for | Daily driver on one 24GB card | Biggest model a 128GB box can hold | Frontier coding via API |
| License | Apache 2.0 | MIT | Custom (revenue-triggered review clause) |
| Smallest usable quant | 16.8GB Q4_K_M | 93GB 1-bit | 239GB 2-bit |
| Runs on your hardware? | Yes — 41–66 tok/s on a used RTX 3090 | Only via Unsloth’s llama.cpp branch (mainline still unmerged) | No — 256GB+ box at 3–9 tok/s |
| The catch | Dense, so slower than the 35B MoE in agent loops | Ollama/LM Studio support still stuck in review | $10B revenue clause, and you can’t fit it anyway |
Honest take: If you own a 24GB card, pull Qwen3.8-27B, add
--spec-type draft-mtpto your llama.cpp flags, and you’re done — that combination is the real news of the last two months. Everything bigger is either an API product or a 96GB+ hardware project, and the newest release of all (Qwen3.8-Omni-Flash) shipped no weights at all.
Two months ago, the July edition of this leaderboard ended with a prediction: two pending releases — Inkling-Small and the Qwen3.8 generation — could actually move the consumer tiers, which nothing in July had managed. Both landed. One of them did move a tier. This is the September accounting: what shipped, what it runs on, which licenses grew teeth, and the one config change that’s worth more than any download this month.
Eight weeks, seven releases, one theme: less open
Here’s everything that hit Hugging Face between the July edition and today, with the numbers that decide whether you can run it. BenchLM’s September board and Thunder Compute’s monthly roundup rank most of these near the top of the open-weight field; the columns that matter for a home lab are the last three.
| Model (weights date) | Total / active | License | Smallest usable quant | What actually runs it |
|---|---|---|---|---|
| Kimi K3 (Jul 27) | 2.8T / 104B | Custom, MIT-style | 1.56TB (BF16 shards) | 8× GB300 node, minimum |
| Inkling-Small (Jul 30) | 276B / 12B | Apache 2.0 | 87.9GB 2-bit | RTX PRO 6000 96GB, 128GB unified box |
| Qwen3.8-Max (Aug 12) | 2.4T / 95B | Custom ($50M revenue trigger) | 397GB 1-bit | 24× H200 — nothing you own |
| Qwen3.8-27B (Aug 14) | 27.78B dense | Apache 2.0 | 16.8GB Q4_K_M | One 24GB card |
| GLM-5.3-Flash (Aug 26) | 320B / 18B | MIT | 93GB 1-bit | 128GB unified memory |
| Qwen3.8-Flash-Next (Aug 26) | 180B / 6B | Qwen Community License 1.0 | 75GB combined RAM+VRAM | 128GB tower or Strix Halo box |
| GLM-5.3 (Aug 28) | 744B / 40B | Custom ($10B security-review clause) | 239GB 2-bit | 256GB+ box, 3–9 tok/s |
Read the license column top to bottom and the trend is hard to miss. Of seven releases, only two ship Apache 2.0 and one MIT; the other four carry custom licenses with revenue triggers, display requirements, or security-review clauses. None of those clauses touch a home lab — we read the GLM-5.3 and Kimi K3 texts in full when they dropped, and the thresholds start at $10 billion and “hyperscaler” respectively — but if your employer runs an Apache-or-MIT-only allowlist, the list of frontier models that clear legal is shrinking, not growing.
And the newest release of the batch took the trend to its endpoint. Qwen3.8-Omni-Flash, announced September 18 with a 1M-token context and audio-video benchmarks Alibaba claims rival Gemini 3.8 Flash, is API-only — no published weights at all, a break from Qwen2.5-Omni, which was Apache 2.0 on Hugging Face. It topped Hacker News on Friday anyway. If you were waiting for our hardware guide on it: there’s no hardware to guide. It’s a hosted product wearing an open-family name.
The 24GB tier flipped — and most owners are leaving a third of the speed off
For over a year, the answer at 24GB was some flavor of Qwen3.6. That’s over. Qwen3.8-27B — Apache 2.0, natively multimodal, 262K context — fits the same 16.8GB at Q4_K_M as its predecessor and has become the default pull for a used RTX 3090 or RTX 4090. Stock llama.cpp decode on a 3090 is about 41 tok/s — respectable for a dense 27B, and not the interesting number.
The interesting number comes from the multi-token-prediction head Alibaba trained into the release checkpoint. It ships inside every Unsloth GGUF you’ve already downloaded, llama.cpp loads it — and then ignores it unless you pass a flag. Two community benchmark projects measured what turning it on is worth:
- At 32K context with f16 KV cache, a five-rep RTX 3090 bench measured 41.55 → 66.39 tok/s (+59.8%) on Unsloth’s UD-Q4_K_XL.
- At full 131K context with q4_0 KV cache, the qwen38-mtp benchmark repo measured 31.0 → 41.3 tok/s (+33%) on the same card across 70+ configurations — and 47.7 → 76.3 tok/s (+60%) on an RTX 4090, 74.3 → 179.7 tok/s on an RTX 5090 at draft depth 4.
The flag pair, from the repo’s reference command:
llama-server -m Qwen3.8-27B-Q4_K_M.gguf \
-c 131072 -ngl 999 -fa 1 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--spec-type draft-mtp --spec-draft-n-max 2 --parallel 1
The problem you’ll actually hit: you add the flags, and tokens per second doesn’t move. Run the same prompt with and without --spec-type draft-mtp and compare the reported speed — if the two numbers match within noise, the drafter isn’t engaged. The usual causes, in order: a llama.cpp build older than the MTP path (the benchmark repo’s runs span builds b10335–b10990; update first), a GGUF conversion that stripped the MTP tensors (use Unsloth’s), or running through Ollama — which doesn’t expose speculative decoding flags at all, so Ollama users are capped at the baseline speed until that lands. Output is bit-for-bit identical to normal decode; speculative execution only changes who proposes the tokens. On Apple Silicon, skip it — the same repo measured no gain on an M4.
Two honest footnotes. First, if you already run the DFlash2 draft-model setup we covered in early September, MTP buys you nothing extra — vLLM benchmarks put them within a few percent of each other; MTP’s advantage is needing zero extra VRAM and zero extra downloads. Second, speed is still the one argument for the old guard: Qwen3.6-35B-A3B, with only ~3B active parameters, does 107 tok/s on a 3090 in Ollama and remains the pick for high-volume agent loops. Smartest model versus fastest model is now a real choice at 24GB; the full 24GB tier guide walks through it.
Tier by tier: what to run in September 2026
| Your memory | Run this | Expect | Tier guide |
|---|---|---|---|
| 8–12GB | Gemma 4 12B QAT (~7GB) | Best quality that fits; smaller Qwen tags for speed | 12GB guide |
| 16GB | Gemma 4 26B-A4B QAT (~15GB); Codestral 2 (13.3GB) for code | Fits with KV room; Qwen3.8-27B does NOT fit here | 16GB guide |
| 24GB | Qwen3.8-27B Q4_K_M + MTP flag (16.8GB); Qwen3.6-35B-A3B for speed | 41–66 tok/s / 107 tok/s on a used 3090 | 24GB guide |
| 32GB (RTX 5090) | Qwen3.8-27B at Q5/Q8 with big context | 74 tok/s stock, 179.7 with MTP at depth 4 | 32GB guide |
| 48GB (2× 3090) | Llama 3.3 70B Q4_K_M, fully resident | 7–10 tok/s; the cheapest real 70B | 48GB guide |
| 96GB | Inkling-Small 2-bit (87.9GB) — new since July; gpt-oss-120b for speed | First single-card frontier-class fit | 96GB guide |
| 128GB unified | Qwen3.8-Flash-Next (75GB floor) today; GLM-5.3-Flash 1-bit (93GB) via Unsloth’s branch | 22–82 tok/s measured on Strix Halo | 128GB guide |
The July edition’s headline — “five weeks of frontier releases and zero movement in the consumer tiers” — died in three places: the 24GB flip, Inkling-Small giving the 96GB class its first frontier-adjacent model, and the two 128GB-class MoEs. Below 16GB, though, the July advice stands untouched: Gemma 4’s QAT checkpoints are still the quality ceiling, and nothing released since June targets small VRAM at all. Run your own card and context through the VRAM calculator before downloading anything measured in tens of gigabytes.
What didn’t happen (September’s quiet corrections)
GLM-5.3-Flash still isn’t in mainline llama.cpp. When we covered the weights on August 29, competing merge PRs had just been filed and we wrote that the gap “should resolve within days.” Wrong call — as of September 20, the lead PR is still open, a third competing implementation has joined the queue, and the tracking issue sits unresolved. Practical consequence: Ollama and LM Studio — which inherit llama.cpp’s model support — still can’t load the 93GB GGUF; Ollama’s only glm-5.3-flash tag calls their hosted cloud service. Unsloth’s llama.cpp branch remains the one working consumer path, three and a half weeks after release. If a 128GB box is your daily driver and you want zero branch-wrangling, Qwen3.8-Flash-Next is the one with working mainline support today.
The Omni line went closed, as covered above — worth restating as a pattern, because it’s the second Qwen3.8-family model (after the API-only multimodal Max) where the open weights and the flagship capability quietly parted ways.
Frontier rankings churned without consequence. Kimi K3, DeepSeek V4 Pro, Qwen3.8-Max, and GLM-5.3 traded places on the September boards — and every one of them needs a datacenter node or a 256GB+ machine. The July conclusion survives its third straight month: for hard problems, use the frontier through an API and let your own silicon run the 17–22GB class it’s actually good at. Our sister site aifoss.dev covers the self-hosting stacks for the open-weight giants, and aicoderscope.com covers wiring local models into Cursor and Cline as coding backends.
What to actually buy
Prices as of September 2026, all discussed above or in the linked tier guides:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| Want the September 24GB stack (Qwen3.8-27B + MTP) tonight | Used RTX 3090 24GB | ~$1,343 avg ($1,287–$1,411 fair range) | Check price |
| Want 179 tok/s MTP decode and Q8 headroom | RTX 5090 32GB | from ~$4,600 street, first-party retail out of stock | Check price |
| Want the 93GB-class MoEs (Flash-Next, GLM-5.3-Flash) at home | GMKtec EVO-X2 128GB | ~$3,649 | Check price |
| Undecided — test a model before buying anything | Rented 3090 | from $0.07/hr | Vast.ai |
The first column is doing the honest work: the used 3090’s price rose ~11% over the last 90 days and BestValueGPU’s September tracker shows the same climb, so “wait for a dip” has been a losing strategy all year — the full used-3090 case is here. The EVO-X2 row buys capacity, not speed — it holds models a 24GB card never will, at 22–82 tok/s rather than three digits; our EVO-X2 head-to-head has the full machine review.
FAQ
Should I replace Qwen3.6-35B-A3B with Qwen3.8-27B on my 24GB card? Keep both. The 27B is smarter and multimodal; the 35B-A3B is 1.6–2.6× faster in decode. Pull the 27B for quality work, keep the MoE loaded for agent loops and bulk tasks.
Does GLM-5.3’s custom license affect my home lab? No. The added clause requires a security review only for companies clearing $10 billion in trailing-12-month revenue. It’s not OSI open source anymore, which matters for corporate allowlists — not for you.
Can I run any of the new frontier models on a 24GB GPU? Not meaningfully. The smallest of the new frontier class is Inkling-Small at 87.9GB for the 2-bit quant — 96GB-territory. On 24GB, offloading a 239GB GLM-5.3 quant is a 3–9 tok/s batch tool at best.
Where’s the Qwen3.8-Omni-Flash local guide? There won’t be one unless weights appear. It’s a hosted API product — no Hugging Face repo, no GGUFs, nothing to size a GPU for.
Is the MTP flag safe to use for production output? Yes. Speculative decoding is lossless — the main model verifies every drafted token, so output is identical to standard decode. The only cost is a small VRAM overhead for the draft state.
Sources
- qwen38-mtp benchmark repo (flags, 70+ configs) — GitHub/sudoingX
- Qwen3.8-27B MTP speculative decoding bench, RTX 3090 — HackMD/thc1006
- llama-bench: Qwen 27B-class on RTX 3090 — ahelpme.com
- Qwen/Qwen3.8-27B — Hugging Face
- Alibaba releases Qwen3.8-Omni-Flash — MarkTechPost
- Qwen3.8-Omni-Flash vs Qwen 3.8: weights status — OrcaRouter
- Qwen3.8-Omni-Flash model page — LLM Reference
- GLM-5.3-Flash (glm5next) PR — ggml-org/llama.cpp #27754
- GLM-5.3 Flash support tracking issue — ggml-org/llama.cpp #27922
- GLM-5.3-Flash GGUF sizes and requirements — Unsloth docs
- Inkling-Small model card — Thinking Machines Lab
- Best Open Source LLMs, September 2026 — Thunder Compute
- Open-source LLM leaderboard — BenchLM.ai
- RTX 3090 used price and fair range — ResalePrices
- RTX 3090 price history — BestValueGPU
- RTX 5090 price tracker, September 2026 — videocardprices.com
- RTX 5090 listings touch $6,000 as official retailers remain out of stock — PCGamesN
Last updated September 20, 2026. Prices and specs change; verify current rates before purchasing.
Recommended Gear
- RTX 3090 24GB (used) — the card the September 24GB stack is built on
- RTX 5090 32GB — 179 tok/s MTP decode and Q8 headroom
- GMKtec EVO-X2 128GB — the cheapest door into the 93GB MoE class
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.