# Aleph Alpha Kolibri 1: Germany's Open-Weight MoE Reasoning Model

> Source: <https://www.mindstudio.ai/blog/aleph-alpha-kolibri-1-sovereign-moe/>
> Published: 2026-10-07 00:00:00+00:00

# Aleph Alpha Kolibri 1: Germany's Open-Weight MoE Reasoning Model

Kolibri 1 is Aleph Alpha's 78B mixture-of-experts model for German and English, with 1M-token context and Apache 2.0 licensing.

## What is Kolibri 1?

Kolibri 1 is a mixture-of-experts (MoE) reasoning model released by Aleph Alpha, a German AI company, on October 3, 2026. It has 78 billion total parameters but activates only 3.46 billion per token, supports explicit reasoning and tool calling, and handles context windows up to 1,048,576 tokens. It’s released under Apache 2.0 and focuses on just two languages, German and English, rather than chasing broad multilingual coverage.

## TL;DR

- **Kolibri 1** is a 78B-parameter MoE model from Aleph Alpha with only 3.46B active parameters per token, built for efficient inference rather than raw scale.
- The model was trained on **20 trillion tokens** of bilingual data (roughly 62.5% English, 23.9% German, 13.6% code), plus additional mid-training and long-context extension phases.
- It supports a native context length of 262,144 tokens, extendable to **1,048,576 tokens** , thanks to an architecture where positional encoding only applies in sliding-window attention layers.
- Kolibri ships with **explicit reasoning modes** (low, medium, high effort, or disabled) and Hermes-style tool calling, both controllable through the chat template.
- The model uses **FP8 weight precision** and runs on hardware as modest as two A100 80GB GPUs, or more comfortably on H100s, H200s, or B200/B300 chips.
- It’s released under **Apache 2.0** , positioning it as a fully open, sovereign alternative to closed frontier models for European enterprises and developers.
- Aleph Alpha frames the bilingual focus (German plus English) as a deliberate trade-off: depth in two languages over shallow coverage across dozens.

### Built like a system. Not vibe-coded.

Remy manages the project — every layer architected, not stitched together at the last second.

## What are the technical specs behind Kolibri 1?

Kolibri 1 is a 50-layer transformer MoE model using a 4:1 ratio of sliding window attention (SWA) to grouped-query attention (GQA). Each layer routes tokens across 384 experts, with 1 shared expert and 6 routed experts active per token. Training used Muon optimization and a technique called Exact Quantile Balancing to keep expert utilization even across the large expert pool.

The headline numbers: 78,103,074,560 total parameters, and 3,457,573,120 active parameters per token. That ratio, roughly 22.6 total parameters for every one activated, is the core efficiency play. A model this size, if dense, would require far more compute per token. By routing through a sparse expert pool, Kolibri delivers large-model capacity at a fraction of the inference cost, though it still needs the full 78GB (at FP8) resident in memory regardless of how many experts are active on a given forward pass.

Weights are stored in `float8_e4m3fn` format in 128x128 blocks with dynamically quantized activations, and the model has been evaluated with an FP8 KV cache. Embeddings, the LM head, normalization layers, and the MoE router stay in bfloat16 for stability. Minimum deployment hardware is two A100 80GB GPUs or two H100 SXM5 GPUs; a single H200, B200, or B300 can also run it, with two H100s or two H200s recommended for better throughput.

## How was Kolibri 1 trained?

Pre-training ran on 20 trillion tokens of filtered, bilingual data combining curated web text, synthetic rephrasings and translations, and high-quality curated sources. The language mix skews toward English (about 62.5%) but keeps a substantial German component (23.9%) and a meaningful code share (13.6%). On top of that, Aleph Alpha ran 3.44 trillion tokens of mid-training and 201 billion tokens of dedicated long-context extension training.

The sequence-length progression is notable: pre-training happened at 16,384 tokens, mid-training extended to 65,536, and the final long-context phase trained at 262,144 tokens, which is the model’s native context length. Because positional encoding is applied only in the sliding-window attention layers rather than globally, Aleph Alpha says the context can in principle be extended indefinitely without position scaling tricks. They validated quality and serving efficiency up to the full 1,048,576-token mark, though they recommend staying at or under 262,144 tokens for latency-sensitive or complex-task deployments.

Post-training combined supervised fine-tuning on a bilingual mix of open-source and synthetic data, followed by reinforcement learning across environments covering reasoning, agentic tasks, and instruction following.

Training ran on 768 NVIDIA B200 GPUs (96 nodes of 8 GPUs each) using expert-parallel, fully-sharded data-parallel configurations. Pre-training alone took 21 days (511 hours, 392,000 GPU-hours). Mid-training added 5 days (90,000 GPU-hours), and the long-context phase took 13 hours (10,000 GPU-hours) with a different parallelism setup. Total measured compute came to 6.4e23 FLOPs. Aleph Alpha also reports an estimated 950 megawatt-hours of energy consumption across pre-training, mid-training, and long-context training, including data center overhead, though this figure excludes SFT, RL, and idle/peak states.

## Why does Kolibri 1 only support two languages?

This is a deliberate design choice, not a limitation Aleph Alpha is apologizing for. Most open-weight models chase broad multilingual support, often at the cost of shallow performance in any single language outside English. Kolibri 1 goes the other direction: German and English only, with a tokenizer built specifically around German word structure (German’s long compound words and inflection patterns don’t tokenize efficiently under generic BPE schemes tuned for English).

The stated logic is “depth over breadth.” By narrowing scope, Aleph Alpha can dedicate more training signal, synthetic data, and evaluation effort to making both languages strong, rather than spreading capacity thin across dozens of languages most enterprise users won’t touch. For a European company targeting regulated industries and government use cases in German-speaking markets, this bet makes practical sense even if it limits the model’s appeal as a general-purpose global assistant.

## How do reasoning mode and tool calling work?

Kolibri 1 supports an explicit reasoning mode configured through the chat template rather than through prompt engineering. Developers can set `reasoning_effort` to `low`, `medium`, or `high` to control how much internal deliberation the model does before answering, or disable it entirely by setting it to `none` or passing `enable_thinking=false`, which produces an immediate response.

Tool calling uses a Hermes-style format, enabled at serving time with `--tool-call-parser kolibri1` and `--enable-auto-tool-choice` flags in vLLM. Function schemas are passed through the standard OpenAI-compatible `tools` field, and the model emits structured tool calls that the parser converts into executable function calls. Reasoning mode and tool calling can be combined, so the model can “think” through a multi-step problem and then issue a tool call based on that reasoning.

Aleph Alpha recommends sampling with `temperature=1.0`, `top_p=0.97`, and `top_k=128`. The model is served through an OpenAI-compatible API, meaning existing client code built against OpenAI’s chat completions format needs minimal changes to point at a self-hosted Kolibri endpoint.

## Is Kolibri 1 worth considering over other open models?

That depends on what you need. Kolibri 1 sits in the same active-parameter class (around 3B active) as other recent open MoE models like GLM-4.7 Flash (30B-A3B), Nemotron 3 Nano (30B-A3B), and Qwen3.5 (35B-A3B), based on the comparison groupings Aleph Alpha uses in its own evaluation tables. Within that class, the differentiators are context length (Kolibri’s 1M-token ceiling is unusually high), bilingual depth in German, and full Apache 2.0 licensing with no gated access or usage restrictions beyond the license terms.

For teams building German-language products, regulated-industry tools in the EU, or systems that need an auditable, self-hostable model stack rather than a closed API, Kolibri 1 is a legitimate option. Aleph Alpha is also a signatory of the EU’s GPAI Code of Practice, which may matter for organizations tracking EU AI Act compliance. For purely English-language, general-purpose use cases, it competes directly against other open MoE models, and the narrower language focus becomes less of a selling point.

The hardware floor (two A100 80GB GPUs minimum) keeps it out of reach for hobbyist single-GPU setups, but it’s realistic for teams with existing multi-GPU inference infrastructure.

## Frequently Asked Questions

### What does “MoE” mean for Kolibri 1’s performance?

## Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

Mixture-of-experts means the model has a large total parameter count (78B) but only activates a small subset (3.46B) for each token processed. This keeps per-token compute low while retaining more model capacity than a dense model of similar inference cost, at the expense of needing the full model in memory.

### What languages does Kolibri 1 support?

German and English only. Aleph Alpha built a custom tokenizer optimized for German word structure and trained on a corpus that’s roughly 62.5% English and 23.9% German, deliberately choosing depth in two languages over broad multilingual coverage.

### How long a context window can Kolibri 1 handle?

Its native trained context length is 262,144 tokens, but Aleph Alpha has validated quality and serving efficiency up to 1,048,576 tokens due to an architecture where positional encoding applies only in sliding-window attention layers.

### What hardware do you need to run Kolibri 1?

The FP8 weights take up about 78GB of memory. Minimum supported hardware is two A100 80GB GPUs or two H100 SXM5 GPUs; a single H200, B200, or B300 chip can also run it, with two H100s or H200s recommended for better performance.

### Is Kolibri 1 free to use commercially?

Yes. It’s released under the Apache 2.0 license, which permits commercial use, modification, and redistribution without royalty obligations, subject to the license’s standard terms.
