Aleph Alpha Kolibri-1 for Local AI in 2026: The 78B Apache 2.0 MoE Your Tools Can't Load Yet

kolibri-1aleph-alphamoelocal-llmgpuhardware-guide

TL;DR: Aleph Alpha released Kolibri-1 on October 3, 2026 — a 78.1B-parameter MoE with only 3.46B active, Apache 2.0, German and English. Community GGUFs exist (Q4_K_M is 47.45 GB), but mainline llama.cpp, Ollama, and LM Studio can’t load it. Don’t buy hardware for this model; run it if you already own 64GB+.

Kolibri-1 (patched llama.cpp)Qwen3.6-35B-A3B (mainline)Rent first
Best forGerman-language work, 64GB+ ownersCoding + everything else on 24GBTesting before any purchase
Memory needed47.5GB (Q4_K_M) + context~22GB, fits one 24GB cardnone
Price / Cost$0 if you own the RAMUsed RTX 3090: $1,399–$1,4503090s from $0.07/hr
The catchPatched llama.cpp only; CUDA untestedLower ceiling on German tasksSetup time per session

Honest take: Kolibri-1 is the most interesting European open-weight release of 2026, and almost nobody should buy hardware for it — it trails Qwen3.6-35B-A3B on coding benchmarks while needing twice the memory, and the local toolchain is held together by a community patch. Run it tonight if you already own the RAM; otherwise keep your money.

Aleph Alpha — the Heidelberg lab that has spent years positioning itself as Europe’s sovereign-AI answer — shipped open weights on October 3, 2026: Kolibri-1, a 78.1B-parameter mixture-of-experts model under Apache 2.0. The pitch is unusual for this site’s beat: a bilingual German-English reasoning model, validated out to a 1,048,576-token context, from a lab that explicitly isn’t American or Chinese.

The question that matters here is narrower: can the hardware in your office actually run it, and should you spend money to make that happen? Short answers: only with a patched build, and almost certainly not. The longer answers involve real measured numbers — 11.9–14.9 tok/s on a desktop CPU, 64 tok/s on a 48GB Mac — and a memory map you can check against your own machine with our VRAM calculator.

What Kolibri-1 actually is

The architecture, from the model card and Aleph Alpha’s tech report:

  • 78.1B total parameters, 3.46B active per token. That’s a sparser ratio than almost anything else in the consumer-adjacent MoE class — 4.4% activation, versus ~8.6% for Qwen3.6-35B-A3B.
  • 50 layers, 384 routed experts per layer, top-6 sigmoid routing plus one shared expert.
  • Hybrid attention: sliding-window layers mixed with full-attention layers, which is how the long context stays affordable.
  • Context: trained to 262,144 tokens, validated by Aleph Alpha up to 1,048,576.
  • License: Apache 2.0 for the weights and configs. Training code and data recipes are not released.
  • Languages: German and English, by design. This is the first serious Apache 2.0 model where German is a first-class target rather than an afterthought.

Why the 3.46B active figure matters: decode speed on every machine is memory-bandwidth-bound, and a MoE only reads its active parameters per token. At Q4_K_M (~0.61 bytes per weight), Kolibri-1 reads roughly 2.1 GB per token generated — about the same traffic as Qwen3.6-35B-A3B. The physics says this model can be fast on modest hardware. The problem is fitting it there in the first place, and getting any software to load it.

The catch: nothing you already use can load it

Kolibri-1’s architecture is registered as kolibri1, and as of October 10, 2026, that architecture is not in mainline llama.cpp, not in any Ollama release, and not in LM Studio. Download one of the community GGUFs and point stock llama.cpp at it, and you get the classic failure:

$ ./llama-cli -m Kolibri-1-Q4_K_M.gguf -p "Hallo, wer bist du?"
llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'kolibri1'

That’s the same error class we covered in our unknown-architecture fix guide, and the usual fix — update llama.cpp — doesn’t work here, because there’s nothing to update to yet. Three real paths exist today:

1. The Hob-forge patch (the practical one). The Hob-forge/Kolibri-1-GGUF repo ships quantizations from Q2_K to Q8_0 plus a kolibri1-llama.cpp.patch that applies against upstream commit 836d571, and a prebuilt release if you don’t want to compile. The converter’s author has validated CPU inference; CUDA, Vulkan, and Metal are explicitly listed as untested by that project. A separate Prompt48 Q4_K_M build (47.45 GB) was likewise validated CPU-only.

2. The CWBudde port (the rigorous one). An independent effort, kolibri-llama-cpp, keeps its changes as ordered patches against a pinned upstream commit, with tokenizer golden tests and a goal of bit-exact parity with Aleph Alpha’s vLLM reference. It already runs the full 50-layer graph and has produced the fastest local number anyone has published so far (more below). Numerical validation against the vLLM reference is still open — treat outputs as provisional.

3. vLLM with Aleph Alpha’s plugin (the official one). Aleph Alpha’s supported serving path is vLLM via their aleph-alpha-inference package; a native vLLM PR has been open since October 5. The launch command per their docs:

vllm serve Aleph-Alpha/Kolibri-1 --tensor-parallel-size 2 \
  --kv-cache-dtype fp8 --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 --enable-auto-tool-choice

The official minimum for the FP8 checkpoint (~78 GB of weights) is 2× A100 80GB, 2× H100, 1× H200, or one B200/B300. That’s datacenter hardware — the vLLM path is not a home-lab path, which is why the patched llama.cpp builds matter. If you want the vLLM background anyway, aifoss.dev’s vLLM review covers the server side.

There is no hosted Aleph Alpha API for Kolibri-1 either. Open weights or nothing — which is honest, but it means the patch situation above is the entire local story right now.

Memory map: what fits where

File sizes are from the Hob-forge and Prompt48 model cards and the CWBudde repo; what-fits verdicts assume you leave room for context (the KV cache is small per token thanks to the hybrid attention, but 64K+ contexts still want several GB).

QuantizationFile sizeRuns onVerdict
BF16149 GiB (156 GB checkpoint)4× A100 80GB classDatacenter only
Q8_079 GiB96GB RTX PRO 6000, 128GB unified memoryFits, pointless for most
Q4_K_M47.45 GB64GB+ RAM (CPU), 64GB/128GB unified memoryThe sane default
CWBudde mix (IQ3_XXS/IQ4_XS experts, Q8_0 rest)33.1 GiB48GB Macs, 48GB dual-GPUTightest working fit
Q2_K~26 GB32GB unified / 2×16GBExists; quality unverified

Read that table against the cards people actually own and the shape of the problem is clear:

  • A 24GB card — used RTX 3090 or RTX 4090 — cannot hold any usable quant. Even Q2_K plus context overflows 24GB. You’d be in CPU-offload territory, reading experts from system RAM, at which point the GPU is barely participating.
  • 48GB is the entry ticket, and today that means a 48GB Mac (where the 33.1 GiB mix is proven to run) or dual 24GB cards (where nobody has published a verified CUDA run — remember, CUDA is listed untested). Our 48GB VRAM guide covers what else that tier buys you.
  • 64GB of system RAM runs it CPU-only — genuinely, not theoretically, with measured numbers below. The sting: a DDR5 64GB kit costs $680–$1,070 as of September 2026 in the middle of the DRAM supply crisis, so “just add RAM” is no longer the cheap advice it was in 2025.
  • 128GB unified-memory machines — GMKtec EVO-X2 ($3,499–$3,649), Framework Desktop ($3,449), Mac Studio M5 Max (from $2,499; $5,099 with 128GB) — hold Q4_K_M with room for the full trained context. That’s the comfortable tier, covered in our 128GB unified memory guide.

Measured speeds, and the ceiling check

Every number here is from a published run, with the bandwidth ceiling (bandwidth ÷ bytes-read-per-token) sanity-checked the way we always do:

$ python3 -c "
bw = {'DDR5 dual-ch desktop': 89.6, 'Strix Halo': 256, 'M5 Max': 614, 'RTX 3090': 936}
gb_tok = 2.1  # 3.46B active params at Q4_K_M, ~0.61 bytes/weight
for k, v in bw.items(): print(f'{k}: decode ceiling ~{v/gb_tok:.0f} tok/s')"
DDR5 dual-ch desktop: decode ceiling ~43 tok/s
Strix Halo: decode ceiling ~122 tok/s
M5 Max: decode ceiling ~292 tok/s
RTX 3090: decode ceiling ~446 tok/s

(The RTX 3090 row is the MoE punchline: the bandwidth is there for 400+ tok/s, but 24GB can’t hold the weights. Kolibri-1 is capacity-bound, not speed-bound, on every consumer GPU.)

HardwareSetupMeasured decodeSource
Ryzen 7 7800X3D, 64GB DDR5, no GPUQ4_K_M, patched llama.cpp, 8 threads11.9–14.9 tok/s (83 tok/s prompt, 46.6GB RAM used)Hob-forge card
48GB unified-memory Mac33.1 GiB quant mix, Metal, 64K context64 tok/sCWBudde port
Intel Arc Pro B70, VulkanQ4_K_M, partial offload22.4 tok/s (runs: 22.5 / 24.0 / 20.6)intelinside PR #98
2× H100 (vendor, serving)FP8, vLLM plugin~18 concurrent 256K-token requestsAleph Alpha tech report

All measured figures sit comfortably under their ceilings, so none of them trip our fake-benchmark alarm. The 64 tok/s Mac number is the one that should make 48GB-Mac owners sit up: that is a very usable speed for a 78B-class model, from a port that didn’t exist two weeks ago. It’s also provisional — that project itself says output parity with the vLLM reference hasn’t been confirmed yet.

One number circulating on social media fails the check: a claimed ~90 tok/s on a single RTX 3090 with a 4-bit quant. A 47.5 GB file does not fit in 24GB, so that run — if real — was mostly reading experts from system RAM, where the ceiling math above caps decode in the low 40s. We couldn’t trace the claim to a reproducible setup, so we’re not citing it as a data point. Treat it as noise.

The benchmark problem: Qwen already does this, smaller

Aleph Alpha’s own reported numbers, from their harness: SWE-bench Verified 66.4, MMLU-Pro 80. Respectable — and the model card’s own comparison shows Qwen3.6-35B-A3B at 73.8 on the same SWE-bench Verified run, with Qwen3.5-35B-A3B at 71.6. Kolibri-1 also trails Qwen3.6-35B-A3B on closed-book knowledge and multi-turn tool use in Aleph Alpha’s published tables. Credit where due: a vendor publishing a comparison its own model loses is rare, and it makes these numbers more trustworthy, not less.

But the implication for buyers is brutal. Qwen3.6-35B-A3B scores higher on coding, runs at 107 tok/s in stock Ollama on a used RTX 3090 (120+ on a 4090), fits in 24GB, and needs zero patches. Kolibri-1 needs twice the memory, a patched build, and loses the English benchmark race. If you’re assembling a local coding stack — say, Continue.dev against a local backend — Kolibri-1 is not the model that earns the hardware.

Where it genuinely has no local rival: German. Kolibri-1 was trained bilingual from scratch, and no Apache 2.0 model at any runnable size treats German as a primary language. If your prompts, documents, or users are German — a real consideration for the EU slice of the home-lab world — the Qwen comparison stops being the relevant one. The same goes if “trained outside the US and China, Apache 2.0, no API dependency” is itself the requirement, which for some European businesses it legally is.

What to actually buy

Prices as of October 2026, all verified in the sections above:

Your situationThe machinePriceWhere
You want this class of MoE running tonight, in stock Ollama, with better coding scoresUsed RTX 3090 24GB + Qwen3.6-35B-A3B$1,399–$1,450Check price
You need Kolibri-1 itself (German work, sovereignty requirement) in one quiet boxGMKtec EVO-X2 128GB$3,499–$3,649Check price
You already own a 48GB+ Mac or a 64GB-RAM desktopNothing — apply the patch and run$0Hob-forge GGUF
You want to test it before spending anythingRented GPUs, by the hour3090s from $0.07/hr (you’ll need two, or one 80GB card)Vast.ai

One honesty note on the EVO-X2 row: Kolibri-1 on Strix Halo means Vulkan or ROCm through the patched builds, and both are currently listed untested by the Hob-forge converter. The Arc B70 Vulkan run (22.4 tok/s, partial offload) suggests the Vulkan path works, but if you buy that machine today, you’re buying the proven 128GB-class workloads — gpt-oss-120b at ~31 tok/s, the big Qwen MoEs — with Kolibri-1 as an expected-soon bonus, not a guarantee. If none of these rows fit, start from the GPU buying guide instead of forcing this model into your budget.

FAQ

Will Ollama and LM Studio support Kolibri-1? Both need the kolibri1 architecture to land in their bundled llama.cpp first, and mainline llama.cpp support is still at the community-patch stage as of October 10, 2026. History says weeks-to-months: popular architectures with this much attention (two independent ports, multiple GGUF uploads in the first week) tend to get merged. Nothing is announced.

Can I run it on a single RTX 4090 or RTX 5090? No quant that preserves usable quality fits in 24GB or 32GB. With experts offloaded to system RAM you’re capped around the low-40s tok/s by DDR5 bandwidth before overheads — and the CUDA path in the community patches is untested anyway. This model wants unified memory or 48GB+.

Is the 1M-token context real on local hardware? The weights were trained to 262K and validated by Aleph Alpha to 1M — on datacenter serving. Locally, the proven configuration so far is 64K context on a 48GB Mac with ~2 GiB to spare. Long-context KV cache is the next thing to eat your memory headroom after the weights; size it with the VRAM calculator.

Why does a 78B model decode faster than a 22B dense model? It only reads its 3.46B active parameters per token (~2.1 GB at Q4_K_M), while a dense 22B at Q4 reads ~13 GB. Decode speed tracks bytes-read-per-token, not total parameters. Capacity (fitting all 47.5 GB of weights) is what you pay for — once they fit, the speed comes almost free.

Should German-speaking users buy hardware for this? If you were already shopping in the 128GB-unified-memory tier, Kolibri-1 strengthens that case meaningfully — it’s the first Apache 2.0 model where German isn’t a compromise. If you weren’t, rent two 3090s (or one 80GB card) for an evening on Vast.ai and see whether the German quality difference matters for your documents before committing $3,500.

  • Used RTX 3090 24GB — $1,399–$1,450 used; still the value play for the Qwen3.6-35B-A3B class that beats Kolibri-1 on coding
  • GMKtec EVO-X2 128GB — $3,499–$3,649; the capacity box where Q4_K_M fits with full-context headroom
  • Mac Studio M5 Max — from $2,499 ($5,099 at 128GB); 614 GB/s of unified bandwidth and the platform where the fastest Kolibri-1 run so far happened

Sources

Last updated October 10, 2026. Prices and specs change; verify current rates before purchasing. Kolibri-1 toolchain status (llama.cpp mainline support, Ollama/LM Studio availability) is moving fast — check the linked repos for the current state.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.