Tooliax Logo
ExploreCompareCategoriesSubmit Tool
News
Tooliax Logo
ExploreCompareCategoriesSubmit Tool
News
Aleph Alpha Unveils Kolibri: A Breakthrough Bilingual MoE Model for Secure Enterprise AI
Back to News
Monday, October 5, 20266 min read

Aleph Alpha Unveils Kolibri: A Breakthrough Bilingual MoE Model for Secure Enterprise AI

Introducing Kolibri: A New Standard in Efficient Bilingual AI

Aleph Alpha, a prominent European AI company, has announced the release of Kolibri, an open-weight Mixture-of-Experts (MoE) language model designed for exceptional performance in both German and English. Kolibri distinguishes itself by possessing 78.1 billion total parameters, yet it activates only 3.46 billion, or approximately 4.4%, of these parameters per token during operation. The model supports an extensive context window of up to 1,048,576 tokens and allows users to dynamically adjust reasoning effort for each request. Distributed under the permissive Apache 2.0 license on Hugging Face, Kolibri is specifically aimed at sovereign deployments within highly regulated sectors, including public administration, industrial applications, and aerospace.

Deployment Capabilities and Infrastructure

Kolibri is engineered for practical deployment. Its FP8 checkpoint, weighing approximately 78GB, can operate on a single NVIDIA B200, B300, or H200 GPU. Alternatively, it can be served using two H100 SXM5 GPUs through vLLM, which includes specialized parsers for Kolibri's reasoning and tool-calling functions.

The Essence of Kolibri: German Engineering Meets Global Ambition

Developed entirely by teams within Germany, Kolibri-1 represents a complete, end-to-end bilingual English-German MoE transformer. According to its technical research report, Aleph Alpha maintained full control over the entire development pipeline, encompassing data curation, architectural design, training infrastructure, post-training optimization, and comprehensive evaluation. The training processes were executed on infrastructure situated in Germany and Finland. The model's design reflects a strong commitment to compliance with key European regulations, including the EU General-Purpose AI Code of Practice, the EU AI Act, and GDPR. Aleph Alpha, a signatory to the Code, ensures that personal data is meticulously redacted from its data pipeline prior to training.

Innovative Architecture: Sparse Experts and Hybrid Attention

Kolibri's architecture features 50 transformer blocks, each with a model width of 2,560. Each MoE layer employs a sigmoid router to score all 384 available experts, subsequently directing each token to the top six experts while always engaging one shared expert. Expert load balancing is managed through Exact Quantile Balancing and Load-Error Injection mechanisms. The attention mechanism utilizes grouped-query attention with 48 query heads and 4 KV heads. Every fifth block incorporates full attention without positional encoding, while the remaining 40 blocks employ sliding-window attention over the preceding 512 tokens, enhanced with RoPE. These sliding-window layers maintain a fixed-size KV cache, resulting in only 10 layers expanding with context length. At equivalent computational resources, the hybrid attention design reportedly supports sequences four times longer than a full-attention model.

A Tokenizer Tailored for German Fluency

The model's 128,000-token vocabulary is built using UniBPE, a method that generates merges similar to BPE but scores each merge based on Unigram loss. This specialized tokenizer achieves 4.90 bytes per token on German text, offering an 11.2% reduction in token count compared to the GPT-5 tokenizer's 4.35 bytes per token for German web content. For English text, Kolibri achieves 4.58 bytes per token, marginally outperforming GPT-5's 4.67 bytes per token.

Rigorous Training Regimen

Kolibri underwent an extensive training process, starting with pre-training on 20 trillion tokens using 768 NVIDIA B200 GPUs. This was followed by 3.44 trillion mid-training tokens with a sequence length of 65,536. A subsequent 201 billion-token long-context stage involved training on sequences of 262,144 tokens. Aleph Alpha incorporated over 2 trillion German tokens, both curated from the web and synthetically generated. Post-training involved a combination of supervised fine-tuning, integrated with MergeMix, and reinforcement learning across more than 1.2 million internal tasks. The Merlin-Arthur protocol further trained the model to abstain from providing answers when retrieved context does not adequately support a response.

Impressive Benchmark Performance

Evaluated using a consistent framework, Kolibri has demonstrated leading performance in English benchmarks, scoring 84.3 on GPQA Diamond, 96.9 on AIME 2025, and 96.0 on AIME 2026. It achieved a tie with Qwen3.5 35B-A3B on the English agentic average at 63.4, though it trailed on BFCL v4 with 61.4 against Qwen3.5's 70.5. While the dense Qwen3.8 27B showed higher overall scores (80.2 EN, 79.9 DE), it activates approximately eight times more parameters per token. Notably, Kolibri decodes about 2.7 times more text per GPU than its predecessor, Kolibri Origin, while simultaneously improving English scores by 21.4 points.

Comparative Edge Against Competitors

In a direct comparison with leading MoE models, Kolibri-1 exhibits compelling advantages:

  • Developer: Aleph Alpha (Germany)
  • Total / Active Parameters: 78.1B / 3.46B (highly efficient)
  • Architecture: MoE, combining sliding-window and full attention for versatility.
  • Max Context: 1,048,576 tokens (industry-leading).
  • Reasoning Control: Adjustable across none, low, medium, and high settings.
  • License: Apache 2.0.
  • Overall Score (EN / DE): 75.5 / 70.8 (leading among compared models).
  • Agentic Avg (EN): 63.4.
  • Industry RAG Avg (DE): 67.5.
  • Code Avg (EN): 89.3.

Kolibri's superior overall scores in both English and German, coupled with its significantly lower active parameter count compared to models like Nemotron 3 Super and Mistral Small 4, highlight its efficiency and performance advantages in key areas.

Simplified Kolibri Deployment

Deployment of Kolibri is streamlined through the aleph-alpha-inference package and vLLM. Users can install the necessary package and serve the model with fp8 KV cache, specifying kolibri1 for reasoning and tool-call parsers, and enabling auto-tool choice. The default context length is 262,144 tokens, with an option to extend to the full 1,048,576-token window via the --max-model-len parameter and a max_position_embeddings override. Reasoning effort can be adjusted using chat_template_kwargs. Recommended inference parameters include a temperature of 1.0, top-p of 0.97, and top-k of 128.

Key Highlights of Kolibri's Release

  • Kolibri features 78.1 billion total parameters, with only 3.46 billion active per token, leveraging 6 routed experts and 1 shared expert.
  • Its hybrid attention mechanism, comprising 40 sliding-window and 10 full-attention layers, facilitates an affordable 1 million-token context.
  • The model was trained on an extensive dataset of 24 trillion tokens, with German content representing over 20% of the mix.
  • Kolibri achieved the highest Overall score in both English (75.5) and German (70.8) among the 12 MoE models evaluated by Aleph Alpha.
  • The FP8 weights are released under Apache 2.0 license and can be served on a single B200 or H200 GPU using vLLM.

This article is a rewritten summary based on publicly available reporting. For the original story, visit the source.

Source: MarkTechPost
Share this article

Latest News

AI Elevates Group Chat: Instinct Launches Collaborative Assistant Feature

AI Elevates Group Chat: Instinct Launches Collaborative Assistant Feature

Oct 6

NYC Council Confronts AI Giants: Demands Accountability on Safety Practices

NYC Council Confronts AI Giants: Demands Accountability on Safety Practices

Oct 6

OpenAI Leads EU AI Act Compliance with Invisible Text Signatures

OpenAI Leads EU AI Act Compliance with Invisible Text Signatures

Oct 6

AI Titans Propel Market to New Heights, Shrugging Off Rising Yields

AI Titans Propel Market to New Heights, Shrugging Off Rising Yields

Oct 6

Over 300 Publishers Push Congress for Federal 'Stealth Bot' Transparency Law

Over 300 Publishers Push Congress for Federal 'Stealth Bot' Transparency Law

Oct 6

View All News

More News

AI Flood Forces Google to Suspend Open-Source Bug Bounty Program

October 5, 2026

AI Flood Forces Google to Suspend Open-Source Bug Bounty Program

AI Spam Overload: Google Halts Critical Bug Bounty Program, Raising Industry Alarms

October 5, 2026

AI Spam Overload: Google Halts Critical Bug Bounty Program, Raising Industry Alarms

Beyond Features: How Behavioral Flags Revolutionize Progressive Rollouts for Quality

October 5, 2026

Beyond Features: How Behavioral Flags Revolutionize Progressive Rollouts for Quality

Tooliax LogoTooliax

Your comprehensive directory for discovering, comparing, and exploring the best AI tools available.

Quick Links

  • Explore Tools
  • Compare
  • Submit Tool
  • About Us

Legal

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • Contact

© 2026 Tooliax. All rights reserved.