{"slug": "how-to-run-kolibri-1-locally-aleph-alpha-s-german-english-moe-model", "title": "How to Run Kolibri-1 Locally: Aleph Alpha's German-English MoE Model", "summary": "Aleph Alpha Research released Kolibri-1, a 78B-parameter German-English mixture-of-experts language model with 3.46B active parameters per token, on October 3, 2026 under an Apache 2.0 license. The FP8 weights occupy about 78 GB, so Aleph Alpha lists minimum serving hardware as two A100 80GB or two H100 SXM5 GPUs, or a single H200, B200, or B300, running through vLLM with a dedicated Kolibri plugin. The model supports context lengths above 1 million tokens, though Aleph Alpha recommends staying at or under 262,144 tokens for serving efficiency.", "body_md": "# How to Run Kolibri-1 Locally: Aleph Alpha's German-English MoE Model\n\nKolibri-1 is Aleph Alpha's 78B MoE model with 3.46B active params. Here's what hardware you need and how to serve it with vLLM.\n\n## What is Kolibri-1?\n\nKolibri-1 is an open-weight mixture-of-experts (MoE) language model built by Aleph Alpha Research, released on October 3, 2026 under an Apache 2.0 license. It has 78 billion total parameters but only activates about 3.46 billion per token, which keeps inference costs down while still giving the model the capacity of a much larger dense network. Unlike most frontier open models that chase broad multilingual coverage, Kolibri-1 focuses specifically on German and English, with a tokenizer built around German word structure and a training mix that’s roughly 62.5% English, 24% German, and 13.6% code.\n\n## TL;DR\n\n- **Kolibri-1 is a 78B-parameter MoE model** from Aleph Alpha with only 3.46B active parameters per token, making it far cheaper to run than its total size suggests.\n- **It needs serious GPU hardware regardless of the small active-parameter count** , because the entire model has to sit in memory even though just a fraction fires on any given token.\n- **The minimum setup is two A100 80GB or two H100 SXM5 GPUs** , with one H200, one B200, or one B300 also sufficient since those chips have enough memory per card.\n- **Context length maxes out at over 1 million tokens** , though Aleph Alpha recommends staying at or under 262,144 tokens for serving efficiency and for complex tasks.\n- **Serving runs through vLLM** using a dedicated Kolibri plugin, with built-in support for an explicit reasoning mode and Hermes-style tool calling.\n- **The model ships in FP8 precision by default** , with weights quantized in 128x128 blocks and an FP8 KV cache option to cut memory further.\n- **It’s positioned as a “sovereign” model** , meaning it’s meant to be self-hosted by organizations that want European-made AI infrastructure they fully control rather than a hosted API.\n\n## Other agents ship a demo. Remy ships an app.\n\nReal backend. Real database. Real auth. Real plumbing. Remy has it all.\n\n## What makes Kolibri-1’s architecture different?\n\nKolibri-1 is a 50-layer transformer MoE model with 384 experts per layer, of which 1 is shared and 6 are routed for any given token. It uses a 4:1 ratio of sliding-window attention (SWA) to grouped-query attention (GQA), which is the trick that makes its long-context support affordable. Most layers only look at nearby tokens, and only a smaller subset of attention layers process the full context window. Since positional encoding is applied only in those sliding-window layers, Aleph Alpha says the effective context can in principle extend indefinitely without position scaling tricks.\n\nThe model was trained with Muon (an alternative to Adam-style optimizers that’s gained traction for large-scale training) and something Aleph Alpha calls Exact Quantile Balancing, aimed at keeping expert utilization even across the MoE layers. Pre-training covered 20 trillion tokens, followed by 3.44 trillion tokens of mid-training and 201 billion tokens dedicated to extending context length. Total compute came to about 6.4e23 FLOPs, run on 768 Nvidia B200 GPUs over roughly three weeks for the main pre-training phase.\n\n## What hardware do you need to run Kolibri-1?\n\nThe FP8 weights take up about 78 GB of memory, which sets the floor for any deployment. Aleph Alpha lists the minimum viable setups as two A100 80GB GPUs, two H100 SXM5 GPUs, one H200, one B200, or one B300. The single-GPU options work because the H200, B200, and B300 each ship with more memory per card than the 80GB A100 and H100 SXM5. For anything beyond minimum viability, Aleph Alpha’s own recommendation is two H100 SXM5s, two H200s, one B200, or one B300, giving headroom for KV cache and longer contexts.\n\nThis is the central trade-off of MoE architecture: compute per token is cheap because only 3.46B parameters activate, but memory footprint isn’t reduced at all, since the full 78B-parameter model has to be resident on the GPUs to route tokens to whichever experts get selected. Anyone used to thinking “small active parameters means small GPU” needs to recalibrate for MoE models specifically.\n\n## How do you serve Kolibri-1 with vLLM?\n\nAleph Alpha distributes Kolibri-1 through a dedicated package called `aleph-alpha-inference`, which includes a vLLM plugin for the model, plus a prebuilt container image (`ghcr.io/aleph-alpha/aleph-alpha-inference`). Installing it also pulls in the specific vLLM version the plugin was built against:\n\n```\npip install 'aleph-alpha-inference>=1'\n```\n\nFrom there, a standard serving command looks like this:\n\n```\nvllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \\\n  --reasoning-parser kolibri1 \\\n  --tool-call-parser kolibri1 \\\n  --enable-auto-tool-choice\n```\n\nThat single command turns on FP8 KV caching, the custom reasoning-output parser, and automatic tool-call detection. For context windows beyond the default 262,144-token recommendation, you add `--max-model-len 1048576 --hf-overrides '{\"max_position_embeddings\": 1048576}'` to push all the way to the validated 1,048,576-token ceiling. Aleph Alpha recommends sampling with `temperature=1.0`, `top_p=0.97`, and `top_k=128`.\n\nOnce the server is running, it exposes an OpenAI-compatible API, so any existing OpenAI client library works by just pointing the `base_url` at your local endpoint.\n\n## How does reasoning mode work in Kolibri-1?\n\nKolibri-1 supports an explicit “thinking” mode, configured per request through chat template parameters rather than a separate model checkpoint. The `reasoning_effort` field accepts `low`, `medium`, or `high`, controlling how much internal deliberation the model does before producing a final answer. Setting `reasoning_effort` to `none`, or passing `enable_thinking=false`, skips the reasoning step entirely and returns an immediate response. This is passed via the `extra_body.chat_template_kwargs` field in an OpenAI-style request, for example:\n\n```\nextra_body={\n    \"chat_template_kwargs\": {\n        \"reasoning_effort\": \"high\",\n        \"enable_thinking\": True,\n    }\n}\n```\n\nThis gives developers a dial between latency and answer quality on a per-request basis, which matters a lot for agentic pipelines where some calls need deep reasoning and others just need a quick lookup or formatting pass.\n\n## How does tool calling work?\n\nTool calling uses a Hermes-style format, enabled by the `--tool-call-parser kolibri1` and `--enable-auto-tool-choice` flags at server startup. Function schemas get passed through the standard `tools` field in a chat completion request, the same way they would for OpenAI or Anthropic function calling. When the model decides to call a tool, the parser extracts structured arguments that the client code can execute and feed back into the conversation as a `tool` role message. Reasoning mode and tool calling can be combined, so the model can think through which tool to call and with what arguments before emitting the call.\n\nAleph Alpha explicitly frames Kolibri-1 for human-reviewed agentic workflows rather than fully autonomous action. Their model card describes it as fitting orchestration layers that call APIs or run searches, provided the calling system validates results, and as an advisory component in decision-support systems rather than the final decision-maker.\n\n## Is Kolibri-1 worth running locally?\n\nFor teams specifically needing strong German-language performance alongside English, Kolibri-1 fills a gap that most general-purpose open models leave unaddressed, since it was trained with a tokenizer and data mix weighted toward German rather than treating it as one of dozens of supported languages. The MoE design also means inference compute cost per token is low relative to the model’s total capacity, which matters for throughput-sensitive deployments.\n\nThe catch is the hardware floor. You need at minimum two 80GB-class GPUs or one newer-generation single card with enough onboard memory, which puts this well outside hobbyist territory and squarely in enterprise or research-lab budgets. It’s also not a general multilingual model, so if your use case spans many languages beyond German and English, a broader model may fit better. Aleph Alpha also positions it as a “sovereign” model, aimed at organizations, particularly in Europe, that want full control over where and how their AI infrastructure runs rather than depending on a hosted API from a non-domestic provider.\n\n## Frequently Asked Questions\n\n### How many GPUs does Kolibri-1 need?\n\nAt minimum two A100 80GB or two H100 SXM5 GPUs, or a single H200, B200, or B300 given their larger per-card memory. Aleph Alpha’s recommended (not just minimum) setup is two H100 SXM5s, two H200s, one B200, or one B300.\n\n### What’s the maximum context length for Kolibri-1?\n\nIt supports up to 1,048,576 tokens, validated by Aleph Alpha for both quality and serving efficiency, though they recommend staying at 262,144 tokens or below for latency-sensitive or complex tasks.\n\n### Does Kolibri-1 support languages other than German and English?\n\nNo, it’s deliberately scoped to just those two languages. Aleph Alpha describes this as a choice of depth over breadth, aiming for stronger performance in German and English rather than broad multilingual coverage.\n\n## \nPlans first.\n*Then code.*\n\nRemy writes the spec, manages the build, and ships the app.\n\n### Can Kolibri-1 call external tools and APIs?\n\nYes. It supports Hermes-style tool calling enabled through vLLM serving flags, with structured function schemas passed via the standard OpenAI-compatible `tools` field, and this can be combined with its reasoning mode.\n\n### What license is Kolibri-1 released under?\n\nApache 2.0, which permits commercial use, modification, and redistribution without the restrictions found in some other open-weight model licenses.", "url": "https://wpnews.pro/news/how-to-run-kolibri-1-locally-aleph-alpha-s-german-english-moe-model", "canonical_source": "https://www.mindstudio.ai/blog/aleph-alpha-kolibri-1-local/", "published_at": "2026-10-04 00:00:00+00:00", "updated_at": "2026-10-04 20:13:41.223572+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-products", "ai-tools"], "entities": ["Aleph Alpha Research", "Kolibri-1", "vLLM", "Nvidia", "A100", "H100", "H200", "B200"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-to-run-kolibri-1-locally-aleph-alpha-s-german-english-moe-model", "markdown": "https://wpnews.pro/news/how-to-run-kolibri-1-locally-aleph-alpha-s-german-english-moe-model.md", "text": "https://wpnews.pro/news/how-to-run-kolibri-1-locally-aleph-alpha-s-german-english-moe-model.txt", "jsonld": "https://wpnews.pro/news/how-to-run-kolibri-1-locally-aleph-alpha-s-german-english-moe-model.jsonld"}}