Inkling-Small Weights Are Live: 276B Params, an 88GB Quant, and the First Single-Card Path to a Frontier-Class Model
TL;DR: Thinking Machines released Inkling-Small’s full weights on July 30 — 276B total, 12B active, Apache 2.0, multimodal. Unsloth’s 2-bit dynamic GGUF is 87.9GB, which means a single RTX PRO 6000 96GB or a 128GB unified-memory box can hold a model that trades blows with the 975B flagship. That has never been true of a model in this class before.
| Inkling-Small 2-bit, local | Inkling-Small via Tinker API | Qwen3.6-35B-A3B on 24GB | |
|---|---|---|---|
| Best for | Privacy-first labs with 96GB+ of VRAM or unified memory | Trying the model today, full quality | Everyone with a normal GPU |
| Hardware cost | ~$3,500 (128GB mini PC) to ~$12,099 (RTX PRO 6000) | $0 upfront | Used RTX 3090, ~$1,050 |
| The catch | 2-bit quality loss is unmeasured; single-card fit leaves ~8GB for KV cache | Prompts leave your machine; 256K ctx cap | It’s a 35B — smart, but not 276B smart |
Honest take: This is the release we put a HOLD on the queue for since July 20, and it delivered better than our estimates — the 2-bit quant came in 10GB smaller than we projected. If you already own 96GB of anything, download it this weekend. If you don’t, a used RTX 3090 running Qwen3.6-35B-A3B is still the sane play; don’t buy $12K of hardware for a model whose low-bit quality nobody has measured yet.
Thinking Machines said “once that work is complete,” and on July 30, 2026, it was: Inkling-Small shipped with full open weights on Hugging Face, fifteen days after the 975B flagship. We’ve been re-checking this release on nearly every run since the July 15 announcement, because the pitch was obvious from the spec sheet: 276B total parameters, only 12B active per token, and benchmark parity with a model 3.5× its size. The flagship Inkling 975B needs a ~$40K Mac cluster or a 512GB server to run at usable quality. Inkling-Small was always going to be the one that matters for home labs.
Now the weights are real, the GGUFs are up, and the question has a concrete answer: which hardware actually runs it, and how fast?
What shipped on July 30
The release, per Thinking Machines’ announcement and the official X post:
| Spec | Value |
|---|---|
| Total parameters | 276B (Mixture-of-Experts) |
| Active per token | 12B |
| License | Apache 2.0 — no revenue caps, no badge clauses |
| Modalities | Text, image, and audio in → text out |
| Context | 1M tokens on open weights (256K on the Tinker API) |
| Thinking | Controllable effort levels, like the flagship |
| Training hardware | NVIDIA GB300 NVL72 systems |
| Checkpoints | BF16 on Hugging Face, plus an NVFP4 checkpoint for Blackwell |
Two things stand out against the current open-weight field. First, the license: Apache 2.0 with zero strings, cleaner than Kimi K3’s revenue-clause MIT variant (which we read in full last week) and far cleaner than anything from Meta. Second, the day-one tooling: NVIDIA published an NVFP4 checkpoint alongside the release, and Unsloth had dynamic GGUFs up almost immediately — this is the same landing pattern that made Tencent’s Hy3 instantly runnable in mid-July, and it’s becoming the difference between a model people run and a model people read about.
The benchmark story: small wins some, loses the agentic rows
Thinking Machines’ claim is that Inkling-Small achieves “comparable performance to Inkling at a quarter of its size,” and the company’s own numbers — flagged as vendor-reported, as always — mostly back it up, with instructive exceptions.
Where Small matches or beats the 975B flagship: IFBench 83.4 vs 79.8 (instruction following, a meaningful win) and HLE with tools 46.6 vs 46.0. Where it’s materially behind, per Latent Space’s roundup: Terminal-Bench 2.1, Tau 3 Banking, SimpleQA Verified, and AudioMC. The pattern is coherent — the smaller model holds up on reasoning and instruction-following, and gives ground on long-horizon agentic work and factual recall, exactly the categories where raw parameter count (and the 41B-active compute budget of the flagship) still buys something. Thinking Machines positions Small for “workloads where cost and latency matter”: coding, LLM-as-judge grading, synthetic data generation.
One honest caveat before the hardware math: every score above was measured on the full-precision model. Nobody — not Thinking Machines, not Unsloth, not the community — has published quality numbers for the 1-bit and 2-bit quants yet, and the weights are barely a day old. Unsloth’s dynamic quants of the 975B flagship retained roughly 74% accuracy at 1-bit; the same technique applied here should do better (2-bit is a much gentler cut than 1-bit), but treat everything below as “fits and runs,” not “fits and runs at benchmark quality.”
The real GGUF sizes — smaller than we estimated
When we wrote the 975B hardware guide on July 17, we projected Inkling-Small’s footprint by scaling the flagship’s confirmed sizes: ~82GB at 1-bit, ~98GB at 2-bit, ~155GB at Q4. The actual Unsloth GGUF builds came in below those estimates at the low end:
| Quant | Size on disk | Our July estimate | Fits |
|---|---|---|---|
| UD-IQ1_S | 74.8GB | ~82GB | 96GB card, easily |
| UD-IQ1_M | 78.8GB | — | 96GB card |
| UD-IQ2_XXS | 82.3GB | — | 96GB card |
| UD-IQ2_M | 82.4GB | — | 96GB card, ~13GB headroom |
| UD-Q2_K_XL | 87.9GB | ~98GB | 96GB card, tight; 128GB unified, comfortable |
| UD-Q4_K_M | 163GB | ~155GB | 192GB+ multi-GPU or 256GB RAM |
The number that matters is 87.9GB for UD-Q2_K_XL. Every prior model in this class missed the single-card window: GLM-5.2’s 2-bit is ~239GB, Kimi K2.7’s is ~325GB, and even Hy3 — July’s “fits a $2,000 box” story — needs ~102GB at Q2_K, sailing past 96GB of VRAM. Inkling-Small’s 2-bit lands under the 96GB line with the strongest quant tier Unsloth publishes below 4-bit.
Hardware paths, from cheapest to fastest
Path 1: a 128GB unified-memory box (~$3,500, and rising)
An AMD Strix Halo machine with 128GB of LPDDR5X — the GMKtec EVO-X2 being the canonical example — holds UD-Q2_K_XL with room for context. These boxes can assign up to 96GB to the GPU in BIOS, and the kyuz0 community toolboxes already demonstrated ~93GB Hy3 quants running this way.
Price warning: when we reviewed the EVO-X2 in June it was $1,999–2,199 for the 128GB/2TB config. The memory supercycle caught up with it — VideoCardz reports the new 128GB/1TB configuration debuting at $3,499, positioned $150 below the existing 2TB version, which puts the 2TB config around $3,649 in late July. Deal listings as low as $1,799 circulated earlier in the summer. If you can find one anywhere near $2K, that’s now the anomaly, not the list price.
Speed expectation, clearly flagged as an estimate: no one has published Inkling-Small numbers on Strix Halo yet (the weights are a day old). Scaling from the closest measured anchor — Hy3, a 295B MoE with 21B active parameters, decodes at ~17 tok/s at a similar quant size on the same 256 GB/s platform — Inkling-Small’s 12B active parameters should land somewhere in the ~20–30 tok/s range. That’s above reading speed and genuinely usable.
Path 2: one RTX PRO 6000 96GB (~$12,099)
The RTX PRO 6000 Blackwell is the first single consumer-buyable card that holds a 276B-class model entirely in VRAM: UD-IQ2_M at 82.4GB leaves ~13GB for KV cache and buffers; UD-Q2_K_XL at 87.9GB fits with roughly 8GB of headroom, which in practice means running 8–16K context, not the headline 1M. At 1,792 GB/s of GDDR7 bandwidth — 7× a Strix Halo box — decode speed on a 12B-active MoE should be a multiple of the mini-PC path. We won’t print a number until someone measures one; bandwidth math has flattered MoE decode before, and routing overhead is real.
The catch is the price trajectory: $8,565 at launch, $13,250 on NVIDIA’s own marketplace as of July — a 55% hike in 16 months — with Newegg listing it around $12,099. We covered who this card actually makes sense for in the RTX PRO 6000 guide; Inkling-Small is the first model that makes the 96GB the exact right size rather than an awkward middle.
Path 3: 4× used RTX 3090 (~$4,200 in cards)
Four used 3090s pool 96GB of VRAM for roughly $4,200 at July’s ~$1,050 eBay floor (average listings run higher — Best Value GPU tracks used at ~$1,050–1,254 and new at $1,488). Add a server board, a big PSU, and the PCIe topology homework that quad-GPU builds demand, and you’re at $5K+ all-in. It works — llama.cpp splits MoE layers across cards fine — but you’re buying a project, and 4× 350W of Ampere pulls more wall power than a single 600W Blackwell card doing the same job. This path makes most sense if you already own two 3090s and can stomach buying two more.
Path 4: 24GB GPU + 96GB system RAM (CPU offload)
Unsloth’s standard MoE recipe applies: keep attention and shared layers on the GPU, spill routed experts to RAM.
llama-server -hf unsloth/Inkling-Small-GGUF:UD-Q2_K_XL \
--ctx-size 16384 -ngl 99 \
-ot ".ffn_.*_exps.=CPU"
Watch the startup log for the offloaded N/N layers to GPU line to confirm the split took. Expect single-digit to low-teens tok/s — the 12B active parameters help, but every token still reads expert weights over the PCIe bus. Same honest verdict as our 70B-on-24GB guide: fine for batch and overnight work, frustrating for interactive coding. With 64GB DDR5 kits still elevated ($688–770 on July sales), building 96GB+ of fresh RAM for this costs real money too.
Path 5: rent or API
RunPod’s A100 80GB runs $1.39/hr Secure — two of them (160GB) fit UD-Q4_K_M with headroom to spare at $2.78/hr, and an H100 pair at $2.89/hr each gives you Q4 with speed. That’s the right way to evaluate whether 2-bit quality holds up on your workload before committing to any of the purchases above. Or skip hardware entirely: the model runs on Thinking Machines’ Tinker platform (API + fine-tuning), at 256K context.
The problem you’ll hit first: your runtime is too old
The error that will greet most day-one downloaders isn’t VRAM — it’s unknown model architecture from a stale llama.cpp build. Inkling support went into llama.cpp with the 975B release in mid-July, so any binary from before then (and any Ollama/LM Studio runtime that hasn’t rebased since) will refuse to load the GGUF no matter how much memory you have. The fix is the same as every new-architecture launch: git pull && cmake --build build for source builds, or grab a release binary newer than the arch merge before filing bug reports. Per Unsloth’s docs, both Inkling and Inkling-Small now run in llama.cpp and Unsloth Studio; Ollama support tends to lag these merges by days to weeks.
Second trap, familiar from the flagship: the 1M-token context headline is a weights capability, not a single-card reality. At 87.9GB of weights in 96GB, your KV cache budget is ~8GB. Run 8–16K context on a single card and be happy; long-context work belongs on the multi-GPU and rental paths.
Where this leaves the leaderboard
Our July open-source leaderboard closed with the consumer tiers unchanged since June: nothing frontier-class fit under 100GB, so the 24GB champion stayed Qwen3.6-35B-A3B and everything bigger was API territory. Inkling-Small is the first genuine tier-break of the summer — a top-five-open-weights-class model with a legal-to-ship Apache 2.0 license that fits hardware a determined individual can actually buy.
It doesn’t dethrone the 24GB picks. A used RTX 3090 running Qwen3.6-35B-A3B at 50–65 tok/s remains the best speed-per-dollar in local AI, and for coding agents specifically, Inkling-Small’s weak Terminal-Bench showing suggests the flagship-or-nothing rule still applies to long agentic sessions. What changed is the ceiling: the gap between “what a home lab can run” and “what the leaderboard respects” just narrowed from 239GB (GLM-5.2’s 2-bit) to 87.9GB — and for the first time, one card covers it.
We’ll revisit with measured tok/s once community benchmarks land; if the Hy3 pattern repeats, the Strix Halo numbers will show up within a week. For coding-tool integration — wiring a local Inkling-Small into Cline or Cursor as a BYOK backend — watch aicoderscope.com, and for the Apache 2.0 self-hosting angle, aifoss.dev covers the FOSS side.
FAQ
Is Inkling-Small really Apache 2.0, no strings? Yes. Both Inkling models ship under plain Apache 2.0 — no revenue thresholds (Kimi K3), no attribution badges (Nemotron’s OpenMDW), no geo restrictions. Fine-tunes and quantized redistributions are clearly permitted, which is why Unsloth GGUFs appeared within hours.
Can my RTX 4090 or 3090 run it? Not in VRAM. The smallest quant is 74.8GB — three times a 24GB card. Your options are CPU offload with 96GB+ of system RAM (slow) or pooling four 24GB cards. For a single 24GB card, Qwen3.6-35B-A3B remains the right model.
Is the 2-bit quant actually good? Unknown as of July 31 — no published quality measurements exist yet for Inkling-Small quants. Unsloth’s dynamic 2-bit builds of comparable MoEs retained ~81–82% of full accuracy. Rent a RunPod pod and test your own workload before spending on hardware.
What about a Mac? Weak option right now. Apple pulled the high-RAM Mac Studio configs earlier this year, and a 96GB M3 Ultra’s default ~75% GPU memory allocation (~72GB) doesn’t hold the 82.4GB IQ2_M without risky limit overrides. The 128GB-unified x86 boxes are the cleaner unified-memory path.
Small vs the 975B flagship — which should I run? If you have to ask, Small. The flagship needs ~290GB at 1-bit; Small needs 88GB at 2-bit and beats it on instruction following. The flagship’s edge is agentic long-horizon work and factual recall — categories where you’d honestly be better served by an API anyway.
Recommended Gear
- RTX 3090 — used ~$1,050–1,254; one for Qwen3.6, four for Inkling-Small
- GMKtec EVO-X2 128GB — the unified-memory path; verify current pricing before buying
- NVIDIA RTX PRO 6000 Blackwell — the single-card path, 96GB GDDR7
Sources
- Introducing Inkling-Small — Thinking Machines Lab
- Inkling-Small release announcement — Thinking Machines on X
- unsloth/Inkling-Small-GGUF — Hugging Face
- Inkling: How to Run Locally — Unsloth Documentation
- Thinky’s Inkling: 975B-A41B multimodal, with Inkling-Small 276B-A12B — Latent Space AINews
- Inkling-Small NVFP4 checkpoint — NVIDIA AI on X
- NVIDIA now lists RTX PRO 6000 Blackwell 96GB at $13,250 — VideoCardz
- RTX PRO 6000 Blackwell price hits $13,250, over 50% hike — Wccftech
- RTX 3090 Price Tracker US, July 2026 — Best Value GPU
- GMKtec EVO-X2 with 128GB RAM debuts at $3,500 — VideoCardz
- Runpod H100 Pricing 2026 — Spheron
Last updated July 31, 2026. Prices and specs change; verify current rates before purchasing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →