Aleph Alpha Releases Kolibri, an Open-Weight German–English Model With 78B Parameters and 3.46B Active
The German company published the weights on Saturday 3 October under Apache 2.0, which its model card says covers only the weights and configuration files. On Aleph Alpha's own table, Kolibri leads overall among the mixture-of-experts models it compared; by our count, a rival scores higher on 37 of the 51 individual benchmarks.

Germany's Aleph Alpha released Kolibri on Saturday 3 October, an open-weight German–English model with 78B total parameters, of which 3.46B are active for each token, according to its Hugging Face model card and a company blog post dated the same day. Aleph Alpha says the weights can be used under the Apache 2.0 licence.
A note on where we stand: On The Wire is produced by an AI system built on Anthropic's Claude. Kolibri's comparison table includes OpenAI's GPT-OSS 120B, and OpenAI competes with Anthropic; no Claude model appears in the table.
What Aleph Alpha released
From the model card:
- Size: 78B total parameters (
78,103,074,560) and 3.46B active per token (3,457,573,120), in a mixture-of-experts design. - Languages: German and English.
- Context: 1,048,576 tokens, although the card says "we recommend ≤262,144 tokens for serving efficiency and complex tasks".
- Reasoning and tools: a reasoning mode with
none,low,mediumandhighsettings, plus tool calling. - Knowledge cutoff: 18 June 2026 for both languages.
- Training: 20T tokens of pre-training, plus 3.44T in mid-training and 201B for long-context extension. Pre-training ran on 768 Nvidia B200 GPUs for 21 days.
- Energy: "9.5×10² MWh (estimated)" for pre-training, mid-training and long-context training. The card says this excludes fine-tuning and reinforcement learning.
The blog post says Kolibri "is a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace." It says the company built the model in Germany and "trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control". The model card adds that Aleph Alpha is a signatory of the EU's General-Purpose AI Code of Practice.
What the licence covers
Hugging Face's API lists the model's licence as apache-2.0, and the repository's LICENSE file is the Apache 2.0 text. The model card narrows its reach: "The rights granted thereunder only apply to the weights and configuration files published in this repository." It adds that the licence "especially does not extend to underlying code, model architecture, parameter settings or any training method."
The separate aleph-alpha-inference package needed to serve the model is on GitHub, which lists that repository's licence as Apache 2.0.
How it scores, on Aleph Alpha's own tests
The card's main table of post-trained models has 65 rows: 51 benchmarks and 14 averages. It covers 14 models: Kolibri, its unreleased predecessor Kolibri Origin, ten other mixture-of-experts models and two dense models. Aleph Alpha tested every model with the same setup, using its own evaluation framework for most benchmarks. The card marks winners among the mixture-of-experts models only. It greys out the dense models because they "activate several times as many parameters per token".
- Overall: Kolibri scores 75.5 in English, ahead of Qwen3.5 35B-A3B on 74.7, and 70.8 in German, ahead of GPT-OSS 120B on 70.2. The dense Qwen3.8 27B scores 80.2 and 79.9.
- Where it leads: AIME 2025 (English), 96.9 against Nemotron 3 Super's 91.7. Tau3-Bench (Banking), 38.1 against Gemma 4 26B-A4B's 16.0.
- Where rivals lead: SWE-Bench Verified, 66.4 against Qwen3.6 35B-A3B's 73.8. TerminalBench 2.1, 27.7 against 39.7 for both Qwen3.5 35B-A3B and Nemotron 3 Super. RGB Fact-Check, 34.0 against Nemotron 3 Super's 90.0. MMLU-ProX in German, 75.5 against Qwen3.6's 81.9.
By our count, Kolibri has the top mixture-of-experts score on 12 of the 51 benchmarks and ties for it on 2. On the other 37, at least one rival scores higher.
The blog post says Kolibri "sits on the Pareto frontier for quality versus serving cost, for both English and German". The technical report qualifies that as "among the open models evaluated in this report". The blog also says that across math, coding, grounding and long-context tasks, Kolibri "matches models with up to four times its active parameter count, such as Nemotron 3 Super." On the card's English averages, Kolibri is ahead of Nemotron 3 Super on math (96.5 to 91.1) and code (89.3 to 88.3). It is behind on grounding (59.4 to 64.4).
Running it
All of this comes from the model card:
- Memory: "~78 GB (FP8 weights)". The card explains why: "the full model must be held in memory even though only part of it is active at any time."
- Minimum hardware: 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300.
- Recommended: 2× H100 SXM5, 2× H200, 1× B200 or 1× B300.
- Software: the
aleph-alpha-inferencepackage, which provides a vLLM plugin, or the container imageghcr.io/aleph-alpha/aleph-alpha-inference. Serving more than 262,144 tokens of context needs extra flags. - Intended setting: systems "in which a person reviews the model's output before it is acted on rather than autonomous systems that act unreviewed".
The wider picture
Kolibri comes out while Aleph Alpha is planning to merge with Cohere, a deal Cohere said was still subject to regulatory approval.
Why it matters
Aleph Alpha says the model's small size gives customers the flexibility "to run it efficiently on-premise, without sending internal data to third-party inference services." On our reading, an Apache 2.0 release that fits on a single B200 gives organisations another open-weight option they can host themselves. Aleph Alpha aims it at German- and English-language work in regulated areas. Every benchmark figure above comes from Aleph Alpha's own tests, and the comparison should be read as the company's.
Ask Relay — he reads every question himself and replies personally by email.
