# How to Run Kolibri-1 Locally: VRAM and Hardware Requirements

> Source: <https://www.mindstudio.ai/blog/run-kolibri-1-locally/>
> Published: 2026-10-04 00:00:00+00:00

# How to Run Kolibri-1 Locally: VRAM and Hardware Requirements

Kolibri-1 needs roughly 78GB of FP8 weights and multi-GPU setups to serve. Here's the full vLLM, VRAM, and hardware guide.

## What do you need to run Kolibri-1 locally?

Kolibri-1 is a 78-billion-parameter mixture-of-experts model from Aleph Alpha, shipped in FP8 precision with a memory footprint of about 78GB for the weights alone. In practice, once you add KV cache, activation memory, and serving overhead, real-world deployments land closer to 140GB of VRAM. Aleph Alpha’s own model card lists the minimum hardware as two A100 80GB GPUs, two H100 SXM5 GPUs, one H200, one B200, or one B300. The model runs through vLLM using a dedicated inference plugin.

## TL;DR

- **Kolibri-1** is a 78B-parameter MoE model from Aleph Alpha with only about 3.46B active parameters per token, which keeps inference cheap despite the large total size.
- **FP8 quantization** brings the weight footprint down to roughly 78GB, but real-world serving (observed at around 140GB VRAM across two GPUs) needs more headroom for KV cache and activations.
- **Minimum supported hardware** per the official model card is 2x A100 80GB, 2x H100 SXM5, or a single H200, B200, or B300, with 2x H100 SXM5 or 2x H200 recommended for better throughput.
- **Serving requires the aleph-alpha-inference package** , which bundles a compatible vLLM build and the Kolibri-specific reasoning and tool-call parsers.
- **Context length defaults to 262,144 tokens** for serving efficiency, though the model has been validated up to just over 1 million tokens if you override the config.
- **Reasoning mode is adjustable on the fly** through four effort levels (none, low, medium, high) passed via chat template arguments, not through separate model variants.
- **Tool calling works out of the box** with Hermes-style function calling, and it can be combined with reasoning mode in the same request.

## Other agents start typing. Remy starts asking.

Scoping, trade-offs, edge cases — the real work. Before a line of code.

## How much VRAM does Kolibri-1 actually need?

The raw number to anchor on is 78GB, which is the size of the FP8-quantized weights. Aleph Alpha stores most of the model in `float8_e4m3fn` format using 128x128 blocks with dynamically quantized activations, while embeddings, the LM head, normalization layers, and the MoE router stay in `bfloat16` for stability. That weight footprint alone already rules out a single consumer GPU.

Once you serve the model for actual inference, VRAM usage climbs past the raw weight size. KV cache, especially with long-context requests, activation buffers, and vLLM’s own memory management add significant overhead. One practical test running the model across two GPUs reported total usage around 140GB of VRAM, which aligns with the gap between “weights on disk” and “weights plus everything needed to generate tokens at a usable context length.”

If you enable the FP8 KV cache option (recommended, and shown in Aleph Alpha’s serving command), memory pressure from the cache itself is reduced compared to running it in full precision, but it does not eliminate the need for multiple high-memory GPUs.

## What hardware configurations actually work?

Aleph Alpha’s model card specifies two tiers:

**Minimum:** 2x A100 80GB, 2x H100 SXM5, 1x H200, 1x B200, or 1x B300.

**Recommended:** 2x H100 SXM5, 2x H200, 1x B200, or 1x B300.

The pattern here is straightforward: you either need two 80GB-class data center GPUs working together, or a single newer-generation GPU with enough onboard memory (H200, B200, B300) to hold the model on its own. There is no officially supported single-GPU path on older or smaller cards, and nothing in the 24GB-48GB consumer or prosumer range is listed as viable. This is a model built for data center inference, not a local desktop rig, despite “only” 3.46B active parameters per token.

For anyone without owned hardware in this class, renting cloud GPU instances is the realistic path, since the minimum configuration already requires enterprise-grade accelerators.

## How do you install and serve Kolibri-1 with vLLM?

Kolibri-1 is not served with a plain, off-the-shelf vLLM install. It depends on the `aleph-alpha-inference` package, which provides a vLLM plugin specific to Kolibri’s architecture (its mixed sliding-window and full-attention layers, its MoE routing, and its custom reasoning and tool-call formats).

Install it with:

```
pip install 'aleph-alpha-inference>=1'
```

A pre-built container, `ghcr.io/aleph-alpha/aleph-alpha-inference`, is also available if you want to skip local dependency management.

To serve the model with both reasoning and tool calling active:

```
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice
```

By default this serves at the recommended context length. To push beyond 262,144 tokens up to the validated maximum of 1,048,576, add:

```
--max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'
```

Aleph Alpha recommends sampling with `temperature=1.0`, `top_p=0.97`, and `top_k=128`. Once the server is running, it exposes an OpenAI-compatible API, so existing clients built against the OpenAI SDK work with minimal changes, just pointing the base URL at your local server.

## Why does context length matter for hardware sizing?

## Other agents ship a demo. Remy ships an app.

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

Kolibri-1 was pre-trained at 16,384 tokens, mid-trained at 65,536, and went through a final long-context training phase at 262,144 tokens, which Aleph Alpha calls its native context length. Because positional encoding only applies to the sliding-window attention layers (a 4:1 ratio of sliding-window to full attention across the model’s 50 layers), context can theoretically extend further without position scaling tricks. Aleph Alpha has validated quality and serving efficiency up to 1,048,576 tokens.

This matters for VRAM planning because KV cache size scales with context length, and Kolibri’s architecture is specifically designed to make long context cheaper than it would be in a model with full attention on every layer. Still, Aleph Alpha explicitly recommends staying at or under 262,144 tokens for latency-sensitive or complex-reasoning workloads. Pushing to a million tokens is possible but adds memory and latency cost that most deployments do not need.

## Is Kolibri-1 worth setting up locally?

For teams that specifically need strong German and English language handling, long-context document processing, or agentic tool-calling in a model with an Apache 2.0 license, Kolibri-1 is a reasonable candidate. Its MoE design (384 experts per layer, with 1 shared and 6 routed) means only about 3.46B parameters activate per token, which keeps per-token compute low relative to a dense model of similar total size. The tradeoff, as Aleph Alpha itself states, is memory: the full 78B parameters have to sit in VRAM even though a small fraction is doing the work on any given token.

That tradeoff is the entire story for local deployment. If you already have access to two A100s, two H100s, or a single H200/B200/B300, running Kolibri-1 is a matter of installing the right package and running one vLLM command. If you don’t have that hardware, renting cloud GPU capacity is the practical alternative, since nothing about this model is sized for consumer cards.

## Frequently Asked Questions

### Can Kolibri-1 run on a single consumer GPU?

No. The minimum supported configurations listed by Aleph Alpha start at two 80GB data center GPUs (A100 or H100) or a single H200, B200, or B300. No consumer GPU has enough VRAM to hold the 78GB of FP8 weights plus KV cache and activation overhead.

### Why does a model with only 3.46B active parameters need so much VRAM?

Because it’s a mixture-of-experts model with 78B total parameters across 384 experts per layer. Even though only a small subset of experts activates per token, all the weights have to be resident in memory to route tokens correctly, so the memory requirement tracks the total parameter count, not the active count.

### What’s the difference between the minimum and recommended hardware setups?

The minimum setups (2x A100 80GB, 2x H100 SXM5, or single H200/B200/B300) will run the model. The recommended setups (2x H100 SXM5, 2x H200, or single B200/B300) offer more headroom for throughput and longer context without hitting memory ceilings as quickly.

### Do you need a special version of vLLM to serve Kolibri-1?

Yes. Kolibri-1 requires the `aleph-alpha-inference` package, which installs a compatible vLLM build along with Kolibri-specific reasoning and tool-call parsers (`kolibri1`). A standard vLLM install without this plugin will not correctly handle the model’s reasoning and tool-calling format.

### How do you control reasoning effort when serving Kolibri-1?

## Remy is new. The platform isn't.

Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.

Reasoning effort is set per request through chat template arguments, not through separate model files. You pass `reasoning_effort` as `none`, `low`, `medium`, or `high` in the `chat_template_kwargs` of your API call, letting you trade off latency against reasoning depth on a per-query basis.
