# Aleph Alpha Kolibri: License, Specs and Benchmarks Checked

> Aleph Alpha's Kolibri checked against its model card: what Apache 2.0 covers, 78.1B/3.46B size, real context limits, hardware, and its own benchmark table.

Aleph Alpha, the Heidelberg AI company, released Kolibri on October 3, 2026, the Day of German Unity, with full weights on Hugging Face under Apache 2.0; the launch post drew more than 500 points on Hacker News. This report checks the release against the model card, the license file and the 189-page technical report: what the license covers, the real size and context, what it takes to run, and where Aleph Alpha's own tables put it ahead of and behind other open-weight models.

## TL;DR

- **License:** the LICENSE file in the Hugging Face repository is the standard Apache 2.0 text (we compared it with the Apache Software Foundation's copy; they are identical). The model card adds that the grant covers only "the weights and configuration files published in this repository", not Aleph Alpha's training code, architecture, parameter settings or methods.
- **Size and context:** 78.1 billion total parameters, 3.46 billion active per token (4.4%, per the technical report). The headline "1M tokens" is reached by extrapolation; the card calls 262,144 tokens the native length and recommends staying at or below it for complex tasks.
- **Benchmarks are Aleph Alpha's own runs.** In its post-training table, Kolibri has the highest English overall score (75.5) among the mixture-of-experts models compared. A dense 27B model, Qwen3.8 27B, scores higher at 80.2 in the same table, and Kolibri trails several rivals on coding agents and multi-turn tool calling.

## What Aleph Alpha released

Aleph Alpha describes Kolibri as "an English-German Mixture-of-Experts Transformer" built for "sovereign mission-critical work in regulated areas including public administration, industrials and aerospace." Two repositories went up on Hugging Face: `Kolibri-1` with FP8 weights and `Kolibri-1-BF16`, both tagged Apache 2.0 and neither gated. Kolibri needs Aleph Alpha's `aleph-alpha-inference` package, a vLLM plugin, which GitHub lists under Apache 2.0 as well.

The release follows an internal predecessor, Kolibri Origin (30.6B total, 3.27B active), which the blog lists with "no public release". Aleph Alpha says work on the training pipeline began in January, Origin finished pre-training on June 11, and Kolibri finished on September 11.

| Item | Kolibri (as stated by Aleph Alpha) |
|---|---|
| Total / active parameters | 78.1B / 3.46B per token |
| Experts | 384 routed plus 1 shared per layer, 6 routed per token |
| Layers and attention | 50 layers; 40 use a 512-token sliding window, 10 use full attention |
| Context | 262,144 native; 1,048,576 "by extrapolation" (model card) |
| Languages | English and German |
| Reasoning effort | none, low, medium, high |
| Knowledge cutoff | June 18, 2026 (English and German) |
| Weight memory (FP8) | about 78 GB |
| Minimum hardware (card) | Two A100 80 GB or two H100 SXM5 GPUs; or a single H200, B200 or B300 |
| Training tokens | 20T pre-training, 3.44T mid-training, 201B long-context |
| Training hardware | 768 NVIDIA B200 GPUs; pre-training took 21 days and 392k GPU hours |

## What the license and model card say

**Apache 2.0 in plain words.** Section 2 of the license grants a "perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable" copyright license to copy, modify and distribute the work. Nothing in it limits commercial use, user counts or company size. The conditions sit in Section 4: pass on a copy of the license, mark files you changed, and keep existing notices. Section 3's patent license ends for anyone who sues claiming the work infringes a patent, and Section 6 grants no trademark rights.

**What the grant does not cover.** The model card's "License and terms" section limits the grant to the weights and configuration files in the repository and states that Aleph Alpha "retains all rights to its artifacts, code, model architecture, training methods, parameter settings and intellectual property rights." In other words, you can run, fine-tune and ship the weights, but the release does not open-source the training pipeline.

**Requests, not conditions.** A separate "Responsible Use" section lists uses Aleph Alpha asks people to avoid, including practices banned by Article 5 of the EU AI Act and unlawful activities "risking death or harm, including those related to military or nuclear applications." The card frames this as encouragement, not a license term, and notes the license is permissive. It also says Aleph Alpha has signed the EU General-Purpose AI Code of Practice.

**The 1M-token figure.** The card calls 262,144 tokens Kolibri's native length, says the context can be extended past it, and says Aleph Alpha has "validated" quality up to 1,048,576 tokens. Serving past 262,144 needs extra vLLM flags. On RULER, a long-context test, the card's base-model table shows Kolibri's score generally falling as prompts grow, with a small uptick from 128k (67.9) to 256k (69.8).

<figure class="inline-figure">
<a href="/blog-media/aleph-alpha-kolibri-open-weight-model-fig-1.svg" target="_blank" rel="noopener" title="Open full-size figure"><img src="/blog-media/aleph-alpha-kolibri-open-weight-model-fig-1.svg" alt="Bar chart of Kolibri Base RULER scores per Aleph Alpha: 86.9 at 4k, 80.9 at 16k, 76.3 at 32k, 72.2 at 64k, 67.9 at 128k, 69.8 at 256k, 65.5 at 512k, 63.2 at 1M." width="1200" height="675" loading="lazy" decoding="async" /></a>
<figcaption>In Aleph Alpha's tests, Kolibri Base scores 63.2 on RULER at 1M tokens, down from 86.9 at 4k. Source: Kolibri-1 model card (Hugging Face), as of Oct 4, 2026.</figcaption>
</figure>

At 1M tokens, the same table puts Kolibri Base (63.2) ahead of Nemotron 3 Nano 30B-A3B Base (58.5) and Qwen3.5 35B-A3B Base (57.5); at every length from 4k to 512k, at least one of those two scores higher.

## Benchmarks, as Aleph Alpha reports them

All scores below come from Aleph Alpha's post-training table on the model card. The card says every model ran on the same setup (Aleph Alpha's open-source eval-framework, plus Harbor for TerminalBench and SWE-Bench), each with its own documented context window and sampling settings, and Kolibri at reasoning effort high. We found no independent benchmark results as of October 4.

| Benchmark (Aleph Alpha's table) | Kolibri | Nemotron 3 Super (120B, 12B active) | Qwen3.6 (35B, 3B active) | Mistral Small 4 (119B, 6.5B active) | Qwen3.8 27B (dense) |
|---|---|---|---|---|---|
| Overall (English) | 75.5 | 73.0 | 71.4 | 63.1 | 80.2 |
| Overall (German) | 70.8 | 67.9 | 67.3 | 61.4 | 79.9 |
| AIME 2026 (English) | 96.0 | 90.4 | 91.0 | 83.1 | 97.7 |
| GPQA Diamond (English) | 84.3 | 78.0 | 83.4 | 74.7 | 89.2 |
| SWE-Bench Verified | 66.4 | 60.2 | 73.8 | 60.8 | 72.6 |
| TerminalBench 2.1 | 27.7 | 39.7 | – | 21.0 | 76.8 |
| BFCL v3 (multi-turn) | 39.8 | 44.6 | 53.5 | 36.2 | 42.5 |
| AA-Omniscience Accuracy | 14.8 | 26.7 | 21.0 | 25.0 | 19.0 |

<figure class="inline-figure">
<a href="/blog-media/aleph-alpha-kolibri-open-weight-model-fig-2.svg" target="_blank" rel="noopener" title="Open full-size figure"><img src="/blog-media/aleph-alpha-kolibri-open-weight-model-fig-2.svg" alt="English overall scores in Aleph Alpha's tests: Qwen3.8 27B (dense) 80.2, Kolibri 75.5, Qwen3.5 74.7, Nemotron 3 Super 73.0, GPT-OSS 72.3, Gemma 4 71.9, Qwen3.6 71.4, Mistral Small 4 63.1." width="1200" height="675" loading="lazy" decoding="async" /></a>
<figcaption>Kolibri tops the mixture-of-experts models in Aleph Alpha's English overall score; the dense Qwen3.8 27B is higher. Source: Kolibri-1 model card, post-training table (Aleph Alpha's runs), as of Oct 4, 2026.</figcaption>
</figure>

Three things the table shows that the launch headline does not:

- **Where Kolibri trails.** By Aleph Alpha's own numbers, Qwen3.6 35B-A3B leads it on SWE-Bench Verified and BFCL v3 multi-turn tool calling, and Kolibri has the lowest AA-Omniscience Accuracy (factual-knowledge accuracy) of the five models above.
- **Grounding is the stated focus.** Aleph Alpha says Kolibri abstains instead of answering wrong on 44% of AA-Omniscience items, against 15% for Kolibri Origin, which it attributes to abstention data and its "Merlin-Arthur" training procedure. In the same row, it lists Qwen3.6 35B-A3B at 56.7.
- **The blog and the card differ on a few rival numbers.** For example, Qwen3.6's AA-Omniscience Index is −15.3 in the launch post and −12.5 on the model card; Kolibri's own figures match in both.

Aleph Alpha also says Kolibri "sits on the Pareto frontier for quality versus serving cost" in both languages, measuring throughput with vLLM on eight B200 GPUs per model.

## How it compares on paper

Size, context and license for the models Aleph Alpha compares against, taken from each publisher's own model card:

| Model | Total / active parameters | Context (per own card) | License (per own card) |
|---|---|---|---|
| Aleph Alpha Kolibri | 78.1B / 3.46B | 262,144 native; 1,048,576 by extrapolation | Apache 2.0 |
| Qwen3.6 35B-A3B | 35B / 3B | 262,144 native; extensible to 1,010,000 | Apache 2.0 |
| Gemma 4 26B A4B | 25.2B / 3.8B | 256K | Apache 2.0 |
| gpt-oss-120b | 117B / 5.1B | Not stated on card | Apache 2.0 |
| Mistral Small 4 | 119B / 6.5B | 256k | Apache 2.0 |
| Nemotron 3 Super | 120B / 12B | Up to 1M | NVIDIA Nemotron Open Model License |

Aleph Alpha built Kolibri for two languages: its blog and technical report put German at 21.3% of pre-training tokens, while the model card lists about 23.9%. Kolibri holds more than twice as many total parameters as Qwen3.6 35B-A3B (78.1B against 35B) while activating a similar number per token, and the card notes that "the full model must be held in memory even though only part of it is active at any time." For more on how licenses differ across open models, see our [open-weight model glossary entry](/glossary/open-weight-model/), our [Qwen 3.6 27B license and memory check](/blog/qwen-3-6-27b-the-sweet-spot-for-powerful-local-ai-development/) and our [Gemma 4 Apache 2.0 report](/blog/google-unleashes-gemma-4-advanced-open-models-redefine-on-device-ai-with-apache-/).

## What "sovereign" means here

Aleph Alpha's post defines sovereignty by "how we built the model, and how it transfers to our customers." It says the model was built in Germany and trained in Germany and Finland "with no foreign control," and that customers can run it on-premise instead of sending data to outside inference services.

Several facts from Aleph Alpha's own documents sit alongside that claim:

- **Inputs from other developers' models.** The model card says English web text was rephrased with Google's Gemma-4-26B-A4B, German text with Mistral-NeMo-12B, and quality-filter labels were generated with Qwen3-32B. It also says the training data "contains material generated with Chinese language models," which Aleph Alpha says it filtered for political bias. Training ran on NVIDIA B200 GPUs.
- **Ownership is about to change.** After announcing a planned combination in April, Cohere and Aleph Alpha on September 16 signed a definitive business combination agreement under which the company will operate globally as Cohere, with plans for dual headquarters in Berlin and Toronto. The release says the deal "remains subject to final regulatory approvals" and is expected to close later this year, and that the combined company's structure "incorporates safeguards and oversight mechanisms" for both markets.
- **Leadership.** On September 28, Aleph Alpha said Co-CEO Reto Spörri had left and Ilhan Scheer continues as CEO and sole managing director. Scheer is set to become Cohere's COO when the deal closes.

<figure class="inline-figure">
<a href="/blog-media/aleph-alpha-kolibri-open-weight-model-fig-3.svg" target="_blank" rel="noopener" title="Open full-size figure"><img src="/blog-media/aleph-alpha-kolibri-open-weight-model-fig-3.svg" alt="Timeline: Jan pipeline work begins; Apr Cohere plan announced; Jun 11 Origin pre-training done; Sep 11 Kolibri pre-training done; Sep 16 Cohere deal signed; Sep 28 Co-CEO leaves; Oct 3 release." width="1200" height="675" loading="lazy" decoding="async" /></a>
<figcaption>Kolibri's training and release ran in parallel with Aleph Alpha's planned combination with Cohere. Source: Aleph Alpha Kolibri blog post (Oct 3, 2026); Aleph Alpha news releases of Sep 16 and Sep 28, 2026.</figcaption>
</figure>

## Early reaction

TestingCatalog noted that the results are vendor-run. On Hacker News, one commenter running the FP8 weights on an RTX Pro 6000 setup reported about 170 tokens per second but said the model "spends way too many tokens on overthinking." More model coverage sits in our [AI and semiconductors section](/category/ai-chips/).

## What could go right / What could go wrong

**What could go right**
- Organizations that need German-language models on their own servers get Apache 2.0 weights with no usage-based conditions, plus a detailed model card and technical report.
- If the grounding results Aleph Alpha reports hold up in outside testing, the abstention behavior targets a common problem in document question-answering systems.

**What could go wrong**
- Every benchmark so far is Aleph Alpha's own run, and a few rival figures differ between the blog and the model card.
- The model needs about 78 GB of memory for its FP8 weights and Aleph Alpha's vLLM plugin to serve, and quality past 262,144 tokens is lower in Aleph Alpha's own RULER results.
- The "no foreign control" description refers to how Kolibri was built; the company that built it has signed an agreement to combine with Cohere, pending regulatory approvals.

## FAQ

<p class="post-faq-q"><span class="post-faq-mark">Q</span> Can Aleph Alpha Kolibri be used commercially?</p>
<p class="post-faq-a"><span class="post-faq-mark">A</span> Yes. The weights use the standard Apache 2.0 license, which allows commercial use, modification and redistribution if you pass on the license, mark changed files and keep notices. The grant covers the weights and configuration files only, not Aleph Alpha's training code or methods.</p>

<p class="post-faq-q"><span class="post-faq-mark">Q</span> How big is Kolibri and what hardware does it need?</p>
<p class="post-faq-a"><span class="post-faq-mark">A</span> Kolibri has 78.1 billion total parameters with 3.46 billion active per token. The model card puts the FP8 weights at about 78 GB and lists a minimum of two A100 80 GB or two H100 SXM5 GPUs, or one H200, B200 or B300.</p>

<p class="post-faq-q"><span class="post-faq-mark">Q</span> Does Kolibri really have a 1 million token context window?</p>
<p class="post-faq-a"><span class="post-faq-mark">A</span> Aleph Alpha says it has validated up to 1,048,576 tokens by extrapolation, but the native trained length is 262,144 tokens, and the card recommends staying at or below that for complex tasks.</p>

<p class="post-faq-q"><span class="post-faq-mark">Q</span> How does Kolibri compare with Qwen3.6 and other open models?</p>
<p class="post-faq-a"><span class="post-faq-mark">A</span> In Aleph Alpha's own tests, Kolibri scores 75.5 English overall against 71.4 for Qwen3.6 35B-A3B and 73.0 for Nemotron 3 Super, but trails Qwen3.6 on SWE-Bench Verified and multi-turn tool calling. The dense Qwen3.8 27B scores higher overall at 80.2.</p>

<p class="post-faq-q"><span class="post-faq-mark">Q</span> Is Aleph Alpha part of Cohere?</p>
<p class="post-faq-a"><span class="post-faq-mark">A</span> Not yet. The two companies signed a definitive business combination agreement on September 16, 2026, under which the combined company will operate as Cohere. The deal remains subject to regulatory approvals and is expected to close later in 2026.</p>

---
Canonical: https://www.signalstack.news/blog/aleph-alpha-kolibri-open-weight-model/
Published: 2026-10-04
