Aleph Alpha’s Kolibri Brings Million-Token AI to Sovereign Deployment

Aleph Alpha has released Kolibri, an open-weight language model designed for German and English, and its architecture packs a striking amount of capability into a small active footprint. The model contains 78.1 billion total parameters but activates only 3.46 billion, or 4.4%, for each token.
That combination gives Kolibri a clear mission: bring powerful bilingual AI to sovereign deployments in regulated sectors such as public administration, industry and aerospace. It also accepts up to 1,048,576 tokens of context, ships under the Apache 2.0 license on Hugging Face, and is deployable on hardware that many organizations can identify and plan around.
A Large Model That Activates Only a Small Core
Kolibri, also called Kolibri-1, is a bilingual English-German Mixture-of-Experts transformer developed end to end by teams in Germany. Training ran on infrastructure in Germany and Finland, while the design targets the EU General-Purpose AI Code of Practice, the EU AI Act and GDPR. Aleph Alpha is a signatory of that Code, and its data pipeline redacts personal data before training.
The model’s scale tells only part of the story. Kolibri stacks 50 transformer blocks with a model width of 2,560, then routes each token through a selective expert system. Every MoE layer scores all 384 routed experts with a sigmoid router, sends each token to the top six, and always runs one shared expert.
That routing strategy keeps the active computation low without shrinking the full model. Exact Quantile Balancing and Load-Error Injection balance the expert workload, helping the system use its large pool of experts without sending every token through every parameter.
Attention adds another important efficiency gain. Kolibri uses grouped-query attention with 48 query heads and four KV heads. Every fifth block runs full attention without positional encoding, while the other 40 blocks use sliding-window attention across the 512 preceding tokens with RoPE.
The sliding-window layers keep a fixed-size KV cache, so only 10 layers grow with context length. At matched compute, this hybrid design supports sequences four times longer than a full-attention model, creating a path to the model’s 1,048,576-token context window.
Built for Long Documents and German Language Work
Kolibri uses a 128,000-token vocabulary trained with UniBPE. On German text, it reaches 4.90 bytes per token, compared with 4.35 for the GPT-5 tokenizer. In English, Kolibri reaches 4.58 bytes per token against 4.67 for GPT-5.
Training covered 20 trillion tokens on 768 NVIDIA B200 GPUs, followed by 3.44 trillion mid-training tokens at a sequence length of 65,536. A separate 201 billion-token long-context stage used sequences of 262,144 tokens, and Aleph Alpha added more than 2 trillion German tokens curated from the web or generated synthetically.
Post-training combined supervised fine-tuning, mixed with MergeMix, and reinforcement learning on more than 1.2 million internal tasks. The Merlin-Arthur protocol trains Kolibri to abstain when retrieved context does not support an answer, a critical behavior for systems working with long documents and regulated information.
The hardware requirements reinforce the deployment focus. The FP8 checkpoint is about 78GB, and Kolibri runs on a single B200, B300 or H200, or on two H100 SXM5 GPUs, served through vLLM.
Strong Benchmarks With a Focus on Efficiency
Kolibri leads several English evaluations, scoring 84.3 on GPQA Diamond, 96.9 on AIME 2025 and 96.0 on AIME 2026. It ties Qwen3.5 35B-A3B on the English agentic average at 63.4, though it trails that model on BFCL v4, where Kolibri scores 61.4 against 70.5.
Its overall score is 75.5 in English and 70.8 in German. Kolibri reaches 67.5 on the German industry RAG average and 89.3 on the English code average, showing how its bilingual design extends beyond general text generation.
The comparison with larger or denser systems highlights the model’s efficiency. The dense Qwen3.8 27B scores higher overall, with 80.2 in English and 79.9 in German, but activates about eight times more parameters per token. Against its internal predecessor Kolibri Origin, Kolibri decodes about 2.7 times more text per GPU while scoring 21.4 points higher in English.
- Kolibri-1: 78.1B total parameters and 3.46B active parameters.
- Qwen3.6 35B-A3B: approximately 35B total parameters and 3B active parameters.
- Nemotron 3 Super: 120B total parameters and 12B active parameters.
- Mistral Small 4: 119B total parameters and 6.5B active parameters.
The surrounding models also take different architectural paths. Kolibri uses MoE with sliding-window and full attention, Qwen3.6-35B-A3B uses MoE with a Gated DeltaNet hybrid architecture, Nemotron 3 Super uses Mamba-2 with MoE and attention, and Mistral Small 4 uses an MoE architecture.
Kolibri’s maximum context reaches 1,048,576 tokens, compared with Qwen3.6 35B-A3B’s native 262,144 tokens, or approximately 1M with YaRN; Nemotron 3 Super also reaches 1M tokens, while Mistral Small 4 supports 256k tokens. Kolibri offers none, low, medium and high reasoning control options, and its Apache 2.0 license matches Qwen3.6 35B-A3B.
Kolibri brings together a long context window, selective computation, bilingual training and deployment requirements aimed at regulated environments. With its open weights, Apache 2.0 license and support for B200, B300, H200 and H100 SXM5 systems, Aleph Alpha is positioning the model as a practical foundation for organizations that need control over where their AI runs and how it handles sensitive information.




