Aleph Alpha has released Kolibri, an open-weight large language model designed for German and English, under the Apache 2.0 license. The model became available on October 3, 2026, with weights hosted on Hugging Face.

Kolibri is a 78.1 billion parameter mixture-of-experts model that activates only 3.46 billion parameters per token—about 4.4% of total capacity. This design allows it to compute like a much smaller model while maintaining performance. The architecture includes 50 layers, each containing 384 expert sub-networks plus one shared expert. A router directs each token to 6 of the 384 specialists.
According to Aleph Alpha, the model was “built in Germany, trained on infrastructure in Germany and Finland, under European and German law, with no foreign control.” The company framed this “sovereign” approach as enabling “full freedom of deployment and intellectual-property safety” for organizations that deploy it on their own servers.
The model demonstrates several technical optimizations for German-language processing. Its custom tokenizer, trained with a new algorithm called UniBPE, requires 11.2% fewer tokens for German text than GPT-5’s tokenizer. The source text includes a test on the German constitution showing Kolibri needed 35,190 tokens compared to 41,482 for GPT-5’s tokenizer—a 17.9% reduction.
Kolibri supports a native context window of 262,144 tokens and has been tested up to 1,048,576 tokens. Forty of its 50 layers use sliding-window attention that looks back only 512 tokens, with full-attention layers interspersed every fifth layer. This design reduces computational requirements while enabling extended context. On the RULER long-context benchmark, Kolibri’s base model scored 63.2 compared to 57.5 for Qwen3.5 35B-A3B’s base model.
The model was trained on approximately 24 trillion tokens—more than a fifth in German—using 768 NVIDIA B200 GPUs. Its knowledge cutoff is June 18, 2026. Aleph Alpha signed the European Union’s General-Purpose AI Code of Practice. The company’s technical report, model card, and launch materials provide additional details about the model’s architecture and performance.
Key facts
- Kolibri is a 78.1 billion parameter mixture-of-experts model released October 3, 2026 under Apache 2.0 license
- The model activates only 3.46 billion parameters per token, reducing computational requirements
- It was trained on infrastructure in Germany and Finland with no foreign control
- Kolibri’s custom tokenizer requires 11.2% fewer tokens for German than GPT-5’s tokenizer
- The model supports native context of 262,144 tokens and has been tested up to 1,048,576 tokens
- It was trained on approximately 24 trillion tokens using 768 NVIDIA B200 GPUs
