Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters

Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face. The target is sovereign deployment in regulated sectors such as public administration, industry and aerospace.

Is it deployable? Yes. The FP8 checkpoint is about 78GB and runs on a single B200, B300 or H200, or on 2 H100 SXM5 GPUs, served through vLLM with dedicated Kolibri reasoning and tool-call parsers.

What is Kolibri?

Kolibri (Kolibri-1) is a bilingual English-German MoE transformer developed end to end by teams in Germany. According to the technical research report, Aleph Alpha team controlled the full pipeline: data, architecture, training infrastructure, post-training and evaluation. Training ran on infrastructure in Germany and Finland. The design targets the EU General-Purpose AI Code of Practice, the EU AI Act and GDPR. Aleph Alpha is a signatory of that Code, and its data pipeline redacts personal data before training.

Architecture: Sparse Experts and Hybrid Attention

Kolibri stacks 50 transformer blocks with a model width of 2,560. Every MoE layer scores all 384 routed experts with a sigmoid router, sends each token to the top 6, and always runs 1 shared expert. Expert load is balanced with Exact Quantile Balancing and Load-Error Injection.

Attention uses grouped-query attention with 48 query heads and 4 KV heads. Every fifth block uses full attention without positional encoding. The other 40 blocks use sliding-window attention over the 512 preceding tokens, with RoPE. Sliding-window layers hold a fixed-size KV cache, so only 10 layers grow with context length. At matched compute, Aleph Alpha team reports the hybrid supports sequences 4 times longer than a full-attention model.

A Tokenizer Built for German

The 128,000-token vocabulary is trained with UniBPE, which builds merges like BPE but scores each merge by Unigram loss. On German text it reaches 4.90 bytes per token, versus 4.35 for the GPT-5 tokenizer. That means 11.2% fewer tokens on German web text. In English, Kolibri reaches 4.58 bytes per token against 4.67 for GPT-5.

Training: 24T Tokens, Then SFT and RL

Pre-training covered 20T tokens on 768 NVIDIA B200 GPUs, followed by 3.44T mid-training tokens at 65,536 sequence length. A 201B-token long-context stage then trained on 262,144-token sequences. Aleph Alpha added more than 2T German tokens it curated from the web or generated synthetically. Post-training combined supervised fine-tuning, mixed with MergeMix, with reinforcement learning on more than 1.2M internal tasks. The Merlin-Arthur protocol trains the model to abstain when retrieved context does not support an answer.

Interactive Explainer: How Kolibri Works