Best Open-Source LLMs for Your Home Lab, September 2026: The 24GB Crown Changes Hands

local-llmleaderboardqwen3-8glm-5-3open-sourcegpuhome-lab

TL;DR: Seven open-weight releases landed since our July leaderboard, and exactly one changed what you should run: Qwen3.8-27B is the new default on 24GB cards, and a single llama.cpp flag most people aren’t using lifts it from ~41 to ~66 tok/s on an RTX 3090. The frontier keeps growing (GLM-5.3, Kimi K3, Qwen3.8-Max) and keeps drifting away from clean open licenses. Below 24GB, nothing moved.

Qwen3.8-27BGLM-5.3-FlashGLM-5.3
Best forDaily driver on one 24GB cardBiggest model a 128GB box can holdFrontier coding via API
LicenseApache 2.0MITCustom (revenue-triggered review clause)
Smallest usable quant16.8GB Q4_K_M93GB 1-bit239GB 2-bit
Runs on your hardware?Yes — 41–66 tok/s on a used RTX 3090Only via Unsloth’s llama.cpp branch (mainline still unmerged)No — 256GB+ box at 3–9 tok/s
The catchDense, so slower than the 35B MoE in agent loopsOllama/LM Studio support still stuck in review$10B revenue clause, and you can’t fit it anyway

Honest take: If you own a 24GB card, pull Qwen3.8-27B, add --spec-type draft-mtp to your llama.cpp flags, and you’re done — that combination is the real news of the last two months. Everything bigger is either an API product or a 96GB+ hardware project, and the newest release of all (Qwen3.8-Omni-Flash) shipped no weights at all.

Two months ago, the July edition of this leaderboard ended with a prediction: two pending releases — Inkling-Small and the Qwen3.8 generation — could actually move the consumer tiers, which nothing in July had managed. Both landed. One of them did move a tier. This is the September accounting: what shipped, what it runs on, which licenses grew teeth, and the one config change that’s worth more than any download this month.

Eight weeks, seven releases, one theme: less open

Here’s everything that hit Hugging Face between the July edition and today, with the numbers that decide whether you can run it. BenchLM’s September board and Thunder Compute’s monthly roundup rank most of these near the top of the open-weight field; the columns that matter for a home lab are the last three.

Model (weights date)Total / activeLicenseSmallest usable quantWhat actually runs it
Kimi K3 (Jul 27)2.8T / 104BCustom, MIT-style1.56TB (BF16 shards)8× GB300 node, minimum
Inkling-Small (Jul 30)276B / 12BApache 2.087.9GB 2-bitRTX PRO 6000 96GB, 128GB unified box
Qwen3.8-Max (Aug 12)2.4T / 95BCustom ($50M revenue trigger)397GB 1-bit24× H200 — nothing you own
Qwen3.8-27B (Aug 14)27.78B denseApache 2.016.8GB Q4_K_MOne 24GB card
GLM-5.3-Flash (Aug 26)320B / 18BMIT93GB 1-bit128GB unified memory
Qwen3.8-Flash-Next (Aug 26)180B / 6BQwen Community License 1.075GB combined RAM+VRAM128GB tower or Strix Halo box
GLM-5.3 (Aug 28)744B / 40BCustom ($10B security-review clause)239GB 2-bit256GB+ box, 3–9 tok/s

Read the license column top to bottom and the trend is hard to miss. Of seven releases, only two ship Apache 2.0 and one MIT; the other four carry custom licenses with revenue triggers, display requirements, or security-review clauses. None of those clauses touch a home lab — we read the GLM-5.3 and Kimi K3 texts in full when they dropped, and the thresholds start at $10 billion and “hyperscaler” respectively — but if your employer runs an Apache-or-MIT-only allowlist, the list of frontier models that clear legal is shrinking, not growing.

And the newest release of the batch took the trend to its endpoint. Qwen3.8-Omni-Flash, announced September 18 with a 1M-token context and audio-video benchmarks Alibaba claims rival Gemini 3.8 Flash, is API-only — no published weights at all, a break from Qwen2.5-Omni, which was Apache 2.0 on Hugging Face. It topped Hacker News on Friday anyway. If you were waiting for our hardware guide on it: there’s no hardware to guide. It’s a hosted product wearing an open-family name.

The 24GB tier flipped — and most owners are leaving a third of the speed off

For over a year, the answer at 24GB was some flavor of Qwen3.6. That’s over. Qwen3.8-27B — Apache 2.0, natively multimodal, 262K context — fits the same 16.8GB at Q4_K_M as its predecessor and has become the default pull for a used RTX 3090 or RTX 4090. Stock llama.cpp decode on a 3090 is about 41 tok/s — respectable for a dense 27B, and not the interesting number.

The interesting number comes from the multi-token-prediction head Alibaba trained into the release checkpoint. It ships inside every Unsloth GGUF you’ve already downloaded, llama.cpp loads it — and then ignores it unless you pass a flag. Two community benchmark projects measured what turning it on is worth:

  • At 32K context with f16 KV cache, a five-rep RTX 3090 bench measured 41.55 → 66.39 tok/s (+59.8%) on Unsloth’s UD-Q4_K_XL.
  • At full 131K context with q4_0 KV cache, the qwen38-mtp benchmark repo measured 31.0 → 41.3 tok/s (+33%) on the same card across 70+ configurations — and 47.7 → 76.3 tok/s (+60%) on an RTX 4090, 74.3 → 179.7 tok/s on an RTX 5090 at draft depth 4.

The flag pair, from the repo’s reference command:

llama-server -m Qwen3.8-27B-Q4_K_M.gguf \
  -c 131072 -ngl 999 -fa 1 \
  --cache-type-k q4_0 --cache-type-v q4_0 \
  --spec-type draft-mtp --spec-draft-n-max 2 --parallel 1

The problem you’ll actually hit: you add the flags, and tokens per second doesn’t move. Run the same prompt with and without --spec-type draft-mtp and compare the reported speed — if the two numbers match within noise, the drafter isn’t engaged. The usual causes, in order: a llama.cpp build older than the MTP path (the benchmark repo’s runs span builds b10335–b10990; update first), a GGUF conversion that stripped the MTP tensors (use Unsloth’s), or running through Ollama — which doesn’t expose speculative decoding flags at all, so Ollama users are capped at the baseline speed until that lands. Output is bit-for-bit identical to normal decode; speculative execution only changes who proposes the tokens. On Apple Silicon, skip it — the same repo measured no gain on an M4.

Two honest footnotes. First, if you already run the DFlash2 draft-model setup we covered in early September, MTP buys you nothing extra — vLLM benchmarks put them within a few percent of each other; MTP’s advantage is needing zero extra VRAM and zero extra downloads. Second, speed is still the one argument for the old guard: Qwen3.6-35B-A3B, with only ~3B active parameters, does 107 tok/s on a 3090 in Ollama and remains the pick for high-volume agent loops. Smartest model versus fastest model is now a real choice at 24GB; the full 24GB tier guide walks through it.

Tier by tier: what to run in September 2026

Your memoryRun thisExpectTier guide
8–12GBGemma 4 12B QAT (~7GB)Best quality that fits; smaller Qwen tags for speed12GB guide
16GBGemma 4 26B-A4B QAT (~15GB); Codestral 2 (13.3GB) for codeFits with KV room; Qwen3.8-27B does NOT fit here16GB guide
24GBQwen3.8-27B Q4_K_M + MTP flag (16.8GB); Qwen3.6-35B-A3B for speed41–66 tok/s / 107 tok/s on a used 309024GB guide
32GB (RTX 5090)Qwen3.8-27B at Q5/Q8 with big context74 tok/s stock, 179.7 with MTP at depth 432GB guide
48GB (2× 3090)Llama 3.3 70B Q4_K_M, fully resident7–10 tok/s; the cheapest real 70B48GB guide
96GBInkling-Small 2-bit (87.9GB) — new since July; gpt-oss-120b for speedFirst single-card frontier-class fit96GB guide
128GB unifiedQwen3.8-Flash-Next (75GB floor) today; GLM-5.3-Flash 1-bit (93GB) via Unsloth’s branch22–82 tok/s measured on Strix Halo128GB guide

The July edition’s headline — “five weeks of frontier releases and zero movement in the consumer tiers” — died in three places: the 24GB flip, Inkling-Small giving the 96GB class its first frontier-adjacent model, and the two 128GB-class MoEs. Below 16GB, though, the July advice stands untouched: Gemma 4’s QAT checkpoints are still the quality ceiling, and nothing released since June targets small VRAM at all. Run your own card and context through the VRAM calculator before downloading anything measured in tens of gigabytes.

What didn’t happen (September’s quiet corrections)

GLM-5.3-Flash still isn’t in mainline llama.cpp. When we covered the weights on August 29, competing merge PRs had just been filed and we wrote that the gap “should resolve within days.” Wrong call — as of September 20, the lead PR is still open, a third competing implementation has joined the queue, and the tracking issue sits unresolved. Practical consequence: Ollama and LM Studio — which inherit llama.cpp’s model support — still can’t load the 93GB GGUF; Ollama’s only glm-5.3-flash tag calls their hosted cloud service. Unsloth’s llama.cpp branch remains the one working consumer path, three and a half weeks after release. If a 128GB box is your daily driver and you want zero branch-wrangling, Qwen3.8-Flash-Next is the one with working mainline support today.

The Omni line went closed, as covered above — worth restating as a pattern, because it’s the second Qwen3.8-family model (after the API-only multimodal Max) where the open weights and the flagship capability quietly parted ways.

Frontier rankings churned without consequence. Kimi K3, DeepSeek V4 Pro, Qwen3.8-Max, and GLM-5.3 traded places on the September boards — and every one of them needs a datacenter node or a 256GB+ machine. The July conclusion survives its third straight month: for hard problems, use the frontier through an API and let your own silicon run the 17–22GB class it’s actually good at. Our sister site aifoss.dev covers the self-hosting stacks for the open-weight giants, and aicoderscope.com covers wiring local models into Cursor and Cline as coding backends.

What to actually buy

Prices as of September 2026, all discussed above or in the linked tier guides:

Your situationThe machinePriceWhere
Want the September 24GB stack (Qwen3.8-27B + MTP) tonightUsed RTX 3090 24GB~$1,343 avg ($1,287–$1,411 fair range)Check price
Want 179 tok/s MTP decode and Q8 headroomRTX 5090 32GBfrom ~$4,600 street, first-party retail out of stockCheck price
Want the 93GB-class MoEs (Flash-Next, GLM-5.3-Flash) at homeGMKtec EVO-X2 128GB~$3,649Check price
Undecided — test a model before buying anythingRented 3090from $0.07/hrVast.ai

The first column is doing the honest work: the used 3090’s price rose ~11% over the last 90 days and BestValueGPU’s September tracker shows the same climb, so “wait for a dip” has been a losing strategy all year — the full used-3090 case is here. The EVO-X2 row buys capacity, not speed — it holds models a 24GB card never will, at 22–82 tok/s rather than three digits; our EVO-X2 head-to-head has the full machine review.

FAQ

Should I replace Qwen3.6-35B-A3B with Qwen3.8-27B on my 24GB card? Keep both. The 27B is smarter and multimodal; the 35B-A3B is 1.6–2.6× faster in decode. Pull the 27B for quality work, keep the MoE loaded for agent loops and bulk tasks.

Does GLM-5.3’s custom license affect my home lab? No. The added clause requires a security review only for companies clearing $10 billion in trailing-12-month revenue. It’s not OSI open source anymore, which matters for corporate allowlists — not for you.

Can I run any of the new frontier models on a 24GB GPU? Not meaningfully. The smallest of the new frontier class is Inkling-Small at 87.9GB for the 2-bit quant — 96GB-territory. On 24GB, offloading a 239GB GLM-5.3 quant is a 3–9 tok/s batch tool at best.

Where’s the Qwen3.8-Omni-Flash local guide? There won’t be one unless weights appear. It’s a hosted API product — no Hugging Face repo, no GGUFs, nothing to size a GPU for.

Is the MTP flag safe to use for production output? Yes. Speculative decoding is lossless — the main model verifies every drafted token, so output is identical to standard decode. The only cost is a small VRAM overhead for the draft state.

Sources

Last updated September 20, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.