Kolibri Activates Just 3.46B of Its 78.1B Parameters

October 4, 2026 • news
Mixture of ExpertsOpen Weights

Aleph Alpha has released Kolibri-1, an open-weight mixture-of-experts model for German and English. It contains 78.1B total parameters but activates only 3.46B, or 4.4%, per token. The Apache 2.0 model supports a maximum context of 1,048,576 tokens and lets users select no reasoning or low, medium and high reasoning effort for each request.

The FP8 checkpoint is approximately 78GB. MarkTechPost reports that it can run on one NVIDIA B200, B300 or H200, or on two H100 SXM5 GPUs. Training used infrastructure in Germany and Finland, while Aleph Alpha positions the model for sovereign deployments in regulated sectors including public administration, industry and aerospace.

Sparse experts and hybrid attention

Kolibri has 50 transformer blocks and a model width of 2,560. Each MoE layer scores 384 routed experts through a sigmoid router, sends every token to the top six and always activates one shared expert. Aleph Alpha uses Exact Quantile Balancing and Load-Error Injection to balance expert utilisation.

Its attention stack combines grouped-query attention with 48 query heads and four KV heads. Forty blocks use RoPE and sliding-window attention over the preceding 512 tokens; every fifth block instead applies full attention without positional encoding. Consequently, only 10 layers have KV caches that grow across the complete sequence, while the other 40 retain fixed-size caches.

Aleph Alpha reports that this hybrid supports sequences four times longer than a full-attention model at matched compute. The released model defaults to a 262,144-token context, with 1,048,576 tokens available through an explicit serving configuration.

Tokenizer and training pipeline

Kolibri uses a 128,000-token UniBPE vocabulary. UniBPE constructs merges like BPE but evaluates each merge using Unigram loss. On German web text, Aleph Alpha reports 4.90 bytes per token, compared with 4.35 for the GPT-5 tokenizer, corresponding to 11.2% fewer tokens. Its reported English result is 4.58 bytes per token versus 4.67 for the GPT-5 tokenizer.

Pre-training covered 20T tokens on 768 NVIDIA B200 GPUs. Aleph Alpha then used 3.44T mid-training tokens at a sequence length of 65,536 and a 201B-token long-context stage with 262,144-token sequences. More than 2T German tokens were curated from the web or generated synthetically by Aleph Alpha.

Post-training combined supervised fine-tuning, MergeMix and reinforcement learning across more than 1.2M internal tasks. The Merlin-Arthur protocol was used to train the model to abstain when retrieved context does not support an answer. Aleph Alpha says its data pipeline redacts personal data before training and targets the EU General-Purpose AI Code of Practice, EU AI Act and GDPR.

Deployment and vendor-reported results

Kolibri can be served through aleph-alpha-inference and vLLM. Its dedicated options include --reasoning-parser kolibri1, --tool-call-parser kolibri1 and --enable-auto-tool-choice. Reaching the full context requires --max-model-len 1048576 plus a max_position_embeddings override. Aleph Alpha recommends temperature 1.0, top-p 0.97 and top-k 128.

SpecificationKolibri-1
Total / active parameters78.1B / 3.46B
Routed experts384, with top six selected per token
Attention layers40 sliding-window, 10 full-attention
Maximum context1,048,576 tokens
Default context262,144 tokens
Reasoning controlNone, low, medium or high
LicenseApache 2.0

In its own evaluation harness, Aleph Alpha reports overall scores of 75.5 in English and 70.8 in German, plus an English code average of 89.3. It also reports GPQA Diamond at 84.3, AIME 2025 at 96.9 and AIME 2026 at 96.0. On BFCL v4, however, Kolibri scored 61.4 against 70.5 for Qwen3.5 35B-A3B.

Aleph Alpha reports that the dense Qwen3.8 27B reached higher overall scores of 80.2 in English and 79.9 in German while activating approximately eight times as many parameters per token. Against its internal Kolibri Origin predecessor, the company says Kolibri-1 decoded about 2.7 times more text per GPU and improved its English score by 21.4 points. These results are vendor-reported and have not been independently reproduced in the supplied source.

AI Mastery analysis

Kolibri’s key engineering advantage is not merely its 4.4% active-parameter ratio. Because only 10 of 50 attention layers grow their KV caches across the entire input, its million-token context avoids the memory profile of a model using full attention in every block. The German-focused tokenizer can further reduce the number of inference steps required for German-heavy workloads.

The trade-off appears in tool use: Aleph Alpha’s own BFCL v4 result trails Qwen3.5 35B-A3B by 9.1 points. Buyers should therefore test real retrieval and tool-calling workloads rather than infer agentic reliability from aggregate scores. Kolibri nonetheless combines sparse routing, request-level reasoning controls and single-GPU deployment in a package relevant to the broader shift toward constrained inference for predictable speed and cost.

Sources

Frequently asked questions

How many parameters does Kolibri-1 activate per token?

Kolibri-1 has 78.1B total parameters but activates 3.46B, or 4.4%, for each token. Every MoE layer routes the token to six of 384 routed experts and also runs one shared expert.

Does Kolibri-1 run on a single GPU?

MarkTechPost reports that the approximately 78GB FP8 checkpoint can run on one NVIDIA B200, B300 or H200. It can alternatively be served on two H100 SXM5 GPUs.

What is Kolibri-1’s maximum context length?

Kolibri-1 accepts up to 1,048,576 tokens, although its default serving context is 262,144 tokens. Enabling the full window requires `--max-model-len 1048576` and a `max_position_embeddings` override.

How was Kolibri-1 trained?

Pre-training covered 20T tokens on 768 NVIDIA B200 GPUs, followed by 3.44T mid-training tokens and a 201B-token long-context stage. Post-training included supervised fine-tuning and reinforcement learning across more than 1.2M internal tasks.

What license does Kolibri-1 use?

Kolibri-1 is available under the Apache 2.0 license on Hugging Face. Its open-weight release includes an FP8 checkpoint of approximately 78GB.

Free interactive tools for the decisions this piece raises.

Related Reading