Breaking the 1.58-Bit Barrier: What Ternary LLM Quantization Means for Your Home Lab VRAM Budget in 2026
TL;DR: Ternary weights ({-1, 0, +1}) cut a 27B model from 16.8GB at Q4_K_M to 5.95GB on disk — and since September 2026 that’s not theory: Ternary Bonsai 2 27B runs on an 8GB card and decodes at 91 tok/s on an RTX 4090. Intel’s new BITCOS paper squeezes storage further, to 1.485 bits per weight. The catch: ternary only works when a lab trains it in, so it covers a handful of models, and today’s flagship needs a llama.cpp fork.
| Qwen3.8-27B Q4_K_M | Ternary Bonsai 2 27B | Rent first | |
|---|---|---|---|
| Best for | Full quality, 32K+ context headroom | 27B-class reasoning on an 8–12GB card | Testing ternary before buying anything |
| VRAM / cost | 16.8GB weights → 24GB card, used ~$1,150–$1,350 | 5.95GB weights → runs on 8GB, comfortable on 12GB | RTX 3090 from $0.07/hr on Vast.ai |
| The catch | The card costs more than some complete PCs | Needs PrismML’s llama.cpp fork; vendor-measured quality | You never own the hardware |
Honest take: Don’t sell your 24GB card — KV cache and every other model you run still don’t shrink. But the old rule that 12GB is a dead end for serious local models died on September 18, 2026, and a used RTX 4070 is suddenly a defensible first GPU.
Every few months a quantization paper makes the rounds claiming your next GPU needs a fraction of the VRAM you thought. Most of them stay papers. This cycle is different, because two things landed in the same week of September 2026: an Intel paper that pushed ternary weight storage below the information-theoretic 1.58-bit floor, and a 27B-class ternary model you can actually download and run on hardware this site covers. Here’s the mechanism, the measured numbers, and the honest buying math.
What the “1.58-bit barrier” actually is
A ternary LLM stores every weight as one of three values: -1, 0, or +1. Three states carry log₂3 ≈ 1.585 bits of information, so 1.58 bits per weight has been treated as the floor for ternary storage — hence “1.58-bit LLMs,” the name Microsoft attached to its BitNet b1.58 line.
In practice nobody even hits 1.585. Packing trits into bytes is awkward: llama.cpp’s two ternary formats, added in PR #8151 for TriLM and BitNet models, land at 1.6875 bits per weight (TQ1_0) and 2.0625 bits per weight (TQ2_0) — TQ2_0 wastes bits on purpose because 2-bit alignment unpacks faster on real CPUs and GPUs.
The paper that hit Hacker News on September 17 (arXiv 2609.16338, 163 points, from Intel’s Evangelos Georganas, Alexander Heinecke, and Pradeep Dubey) starts from a measurement, not a theory. The authors profiled the actual symbol distribution of 29 ternary models across seven families — BitNet b1.58, Bonsai, CAT-Q, ParetoQ, TriLM/Spectra, Maple, BitCPM — and found the three symbols are nowhere near uniformly distributed: zeros make up as much as 51.5% of all weights. Uniform ternary packing pays for information the weights don’t contain.
Their layout, BITCOS, splits storage into a dense presence bitmap (is this weight zero or not?) plus a compacted sign vector for the non-zeros. On the sparsest models that reaches 1.485 bits per weight — under the “barrier,” which was only ever a bound for uniformly distributed ternary symbols. It beats five-trit-per-byte packing in 26 of the 29 models tested, and because the bitmap unpacks with cheap mask operations, their AVX-512, AVX2, and Intel Xe2 kernels run up to 1.28× faster than production ternary matrix-vector kernels.
What BITCOS is not: a new quantization method. It doesn’t make models more accurate or make ternary training easier — it stores and serves already-ternary weights more efficiently. And as of late September 2026 it lives in the paper’s kernels, not in llama.cpp, Ollama, or LM Studio. Nothing you run today gets faster because this paper exists. Its significance is directional: CPU-vendor engineers are now optimizing ternary serving, which is what mainstream support looks like before it happens.
The VRAM math, applied to a model you actually want
Abstract bits-per-weight numbers only matter when you multiply them by a real model. Take Qwen3.8-27B (27.78B dense parameters), the current default recommendation for 24GB cards — our hardware guide measured 16.8GB at Q4_K_M and ~41 tok/s on an RTX 3090:
| Format | Bits/weight | 27B-class weights on disk | Fits |
|---|---|---|---|
| FP16 | 16 | ~55.6GB | 3× RTX 3090, or 64GB+ unified memory |
| Q4_K_M GGUF | ~4.8 | 16.8GB (measured) | 24GB card |
| PQ2_0 ternary (Bonsai 2) | 2.13 | 7.21GB (measured) | 12GB card |
| PTQ1_0 ternary (Bonsai 2) | 1.72 | 5.95GB (measured) | 8GB card, barely |
| log₂3 floor | 1.585 | ~5.5GB | theoretical |
| BITCOS layout | 1.485 (sparsest) | ~5.2GB | research kernels only |
Two things jump out of that table. First, the step from Q4_K_M to ternary is enormous — 16.8GB down to 5.95GB, a 2.8× cut, which moves a 27B model down two whole GPU price tiers. Second, the step from today’s ternary formats to BITCOS is small — roughly 0.8GB on a 27B model. The Intel paper is an efficiency win for whoever serves ternary models at scale; it is not, by itself, a reason to change what card you buy.
Why you can’t just quantize your favorite model to 1.58 bits
Here’s the constraint that separates ternary from every quantization level you’ve used before: Q4 is a post-processing step; ternary is a training decision.
When you quantize a normal FP16 model to Q4_K_M, you keep enough precision that the model still works — our Q4/Q5/Q6/Q8 quality-loss measurements put the typical hit at a few percent. Push the same post-training rounding to three values and the model collapses into noise. This isn’t hypothetical; it’s a documented failure mode. llama.cpp issue #15193 reports exactly what happens when someone runs llama-quantize with TQ1_0 on an ordinary model:
$ ./llama-quantize model-f16.gguf model-tq1_0.gguf tq1_0
# quantization completes without error...
$ ./llama-server -m model-tq1_0.gguf
# → server starts, model loads, output is garbled nonsense
The fix is understanding what the format is for: TQ1_0 and TQ2_0 are storage formats for models whose weights were already ternary when training finished — BitNet-architecture models trained with quantization-aware training, or ternary conversions done with distillation and calibration by the model’s authors. They are not a knob you turn on Llama or Qwen weights yourself. (Post-training ternarization is an active research area — a separate 2026 paper, ScaleQ-1.58, gets reasoning models to ternary by calibrating on the model’s own reasoning traces — but nothing production-grade has shipped for arbitrary models.)
The practical consequence: your ternary options are exactly the models someone trained or converted for you. That list, as of September 2026:
| Model | Params | Ternary size | How you run it |
|---|---|---|---|
| BitNet b1.58 2B4T (Microsoft) | 2B | 0.4GB non-embedding | bitnet.cpp, CPU-first |
| Falcon-E 1B / 3B (TII) | 1B / 3B | ~1GB class | bitnet.cpp / onebitllms |
| TriLM / Spectra suite | 99M–3.9B | sub-1GB | llama.cpp TQ formats |
| Bonsai 27B (PrismML, Jul 2026) | 27.32B | 7.17GB (Q2_0_g128) | PrismML llama.cpp fork |
| Ternary Bonsai 2 27B (PrismML, Sep 18, 2026) | 27B-class | 5.95GB (PTQ1_0) | PrismML llama.cpp fork / MLX fork |
Microsoft’s BitNet b1.58 2B4T is the proof-of-concept everyone cites — 0.4GB of non-embedding memory, 29ms CPU decode latency against 41–124ms for comparable FP16 2B models, and 0.028J per token versus 0.186–0.649J (technical report). Impressive engineering, but a 2B model isn’t why you built a home lab. The bottom row is.
Bonsai 2 27B: the first ternary model worth a home-labber’s time
On September 18, 2026, PrismML released Ternary Bonsai 2 27B: a ternary conversion of Qwen3.8-27B, Apache 2.0, with the vendor claiming 98.2% of the FP16 base model’s benchmark performance — math within half a point of full precision, coding level with the baseline. (Its July predecessor, built from Qwen3.6-27B, retained a claimed 95%, averaging 80.49 across 15 thinking-mode benchmarks.) Those are vendor-run numbers and you should hold them loosely until independent evals accumulate — but even with a few points of slippage, this is 27B-class reasoning in a 5.95GB file.
The measured throughput, from PrismML’s numbers plus community benchmarks:
| GPU | Ternary Bonsai 2 decode | Same card on Qwen3.8-27B Q4_K_M |
|---|---|---|
| RTX 5090 32GB | 129.9 tok/s (PQ2_0) | ~45 tok/s at the heavier Q5 |
| RTX 4090 24GB | 91.1 tok/s (PTQ1_0) | ~70 tok/s |
| RTX 3090 24GB | ~54 tok/s (community-run) | ~41 tok/s |
| RTX 4070 12GB | 37.2 tok/s (ternary Q2_0) | can’t load it at all |
That last row is the story. A 12GB card cannot run Qwen3.8-27B at Q4_K_M in any configuration — the weights alone are 16.8GB. The same card runs the ternary build at reading speed with headroom. Fewer bits per weight also means fewer bytes read from VRAM per token, which is why the ternary build is faster than Q4 on the same silicon: decode is memory-bandwidth-bound, the same physics covered in our NPU vs GPU analysis.
The three catches
1. Stock tooling won’t run it. Mainline llama.cpp — and therefore Ollama and LM Studio — either refuses the Bonsai GGUF outright or loads it and emits nonsense, because the g128 group-scale ternary kernels and Hadamard rotation logic aren’t upstream. You build PrismML’s llama.cpp fork for CUDA or CPU (standard llama.cpp build: cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release), or their MLX fork on Apple Silicon. If you saw gibberish output from a ternary GGUF in stock Ollama, that’s not a broken download — it’s the wrong runtime.
2. The KV cache doesn’t shrink. Ternary compresses weights, not context. Per VRAM-requirement measurements, the 5.95GB PTQ1_0 file needs about 6.8 GiB total at 4K context, 8.6 GiB at 32K, and 22.9 GiB at the full 262,144-token window with an f16 cache. Read that last number again: at max context, the “8GB model” needs a 24GB card anyway. On a 12GB card you’re realistically working with 32K–64K context — plenty for chat and coding sessions, not for whole-repo dumps. Check your own model + context combination in our VRAM calculator.
3. It’s one model. Ternary coverage is whatever labs choose to train. If your workflow depends on a specific fine-tune, a vision model, or next month’s release, you’re back on the Q4 ladder until someone does the ternary conversion work.
Does this change which GPU you should buy?
The question that actually costs money. Three situations, priced at September 2026 street rates (a reminder from our price-crisis coverage: MSRPs are fiction this year).
If you were about to spend $1,150–$1,350 on a used RTX 3090: still do it. The 24GB card runs the full-quality Q4_K_M 27B at 41 tok/s with 32K-context headroom, runs every other model family without waiting for ternary conversions, and — the detail most ternary coverage skips — is the only tier on this page that handles Bonsai 2’s longer contexts, since KV cache eats the savings back. Ternary makes a 24GB card more capable, not less necessary.
If your ceiling is ~$500–$700: a 12GB card is no longer a dead end. A used RTX 4070 ran $485–$585 on the used market in September 2026 (~$703 new, per BestValueGPU), and 37 tok/s on a 27B-class reasoning model is genuinely usable — human reading speed is 7–10 tok/s. A year ago the honest advice at this budget was “save up for 24GB.” That advice is now conditional, not absolute.
If you’re buying because of ternary: don’t, yet. One flagship model, a required fork, and vendor-measured quality is an ecosystem seed, not an ecosystem. And don’t wait for BITCOS-era hardware either — the paper’s gain over today’s formats is ~0.8GB on a 27B model, with kernels demonstrated on Intel silicon. Rent an hour first and see if Bonsai 2’s output quality holds up for your workload.
What to actually buy
Prices as of September 2026, all verified above:
| Your situation | The move | Price | Where |
|---|---|---|---|
| Want full-quality 27B + context headroom today | Used RTX 3090 24GB | $1,150–$1,350 | Check price |
| Budget caps near $600, OK with ternary + a fork build | Used RTX 4070 12GB | $485–$585 used | Check price |
| Want to test Bonsai 2 quality before spending anything | Rented 3090/4090 by the hour | from $0.07/hr | Vast.ai |
FAQ
Can I convert my own model to 1.58-bit with llama-quantize? No. TQ1_0/TQ2_0 quantization of a normal FP16 model completes without errors and then produces garbage output (llama.cpp #15193). Ternary has to be trained in (BitNet-style QAT) or converted with distillation by people with the base weights and a GPU cluster. Use the ternary builds labs publish; quantize your own models to Q4_K_M like before.
Is BITCOS something I can enable? Not today. It’s a storage layout with research kernels for AVX-512, AVX2, and Intel Xe2 GPUs, published September 14, 2026. No llama.cpp, Ollama, or bitnet.cpp integration has shipped as of this writing. If it lands, expect roughly a 10–14% weight-size cut versus TQ1_0-class packing and up to 1.28× faster ternary matvec — nice, not tier-changing.
Does ternary hurt quality as much as Q2_K does? Different mechanism entirely. Q2_K post-training rounds a model that was never trained for it, and quality falls off a cliff. A ternary model was optimized under the ternary constraint from the start (or distilled into it), which is how Bonsai 2 posts a claimed 98.2% retention at 1.72 bits per weight while Q2 post-quants of comparable models lose far more. Vendor numbers, again — but the architecture-vs-afterthought distinction is real and shows up across BitNet, Falcon-E, and TriLM evals.
Will Ollama support Bonsai 2 natively?
Mainline llama.cpp already carries TQ1_0/TQ2_0 for standard BitNet/TriLM models, but Bonsai’s g128 group scales and Hadamard rotation are fork-only as of late September 2026. Community Ollama repacks exist but stock Ollama can’t run the official packs. Watch PrismML’s repos for an upstreaming effort; until then it’s a cmake build.
Should I wait for a ternary 70B? Nothing announced. Note the scaling: a 70B ternary model at 1.72 bpw would be ~15GB of weights — inside a 24GB card with room for cache. If ternary conversion becomes routine at that scale, the used RTX 3090 gets another life extension. One more reason the 24GB recommendation survives this paper.
Sources
- Breaking the 1.58-bit Barrier for Ternary LLMs — arXiv 2609.16338 (Georganas, Heinecke, Dubey / Intel)
- Hacker News discussion of the BITCOS paper — Sep 17, 2026
- ggml-quants: ternary packing for TriLMs and BitNet b1.58 — llama.cpp PR #8151
- TQ1_0-quantized normal models output garbage — llama.cpp issue #15193
- BitNet b1.58 2B4T Technical Report — Microsoft, arXiv 2504.12285
- microsoft/bitnet-b1.58-2B-4T — Hugging Face
- Falcon-Edge: fine-tunable 1.58-bit language models — TII
- PrismML Ternary Bonsai 2 27B GGUF — Hugging Face
- PrismML releases Ternary Bonsai 2 27B, 5.9GB, 98.2% of Qwen3.8-27B — MarkTechPost, Sep 18, 2026
- Bonsai 2 27B VRAM requirements by context length — VRAMCalculator
- Community Bonsai 2 27B benchmarks on RTX 3090 — PrismML-Eng/Bonsai-demo issue #181
- ScaleQ-1.58: post-training ternary quantization for reasoning LLMs — arXiv 2608.01078
- RTX 4070 price history, new and used — BestValueGPU, Sep 2026
- 1.58-bit large language model — Wikipedia
Last updated September 23, 2026. Prices and specs change; verify current rates before purchasing. Running the ternary stack self-hosted? Our sister sites cover the software side: Ollama review on aifoss.dev and local BYOK coding backends on aicoderscope.com.
Recommended Gear
Products linked in this article:
- Used RTX 3090 24GB — $1,150–$1,350 used; the full-quality path: Qwen3.8-27B Q4_K_M at ~41 tok/s, plus every ternary model with max-context headroom.
- Used RTX 4070 12GB — $485–$585 used; the new budget entry to 27B-class reasoning via Ternary Bonsai 2 at 37 tok/s.
- No hardware yet: rent an RTX 3090 from $0.07/hr on Vast.ai and test the ternary stack for the price of a coffee.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.