# Mistral Large 4 vs Beam vs Kolibri: which open weights you can download, and what your hardware can serve

> Source: <https://stackness.dev/blog/mistral-large-4-vs-beam-vs-kolibri-which-open-weights-you-can-download-and-what-your-hardware-can-se>
> Published: 2026-10-07 18:57:28+00:00

# Mistral Large 4 vs Beam vs Kolibri: which open weights you can download, and what your hardware can serve

Mistral Large 4 vs Beam vs Kolibri, as of 7 October 2026: only [Kolibri](https://stackness.dev/tools/kolibri) has weights you can download. Aleph Alpha put the 78.1B-parameter model, 3.46B of them active per token, on [Hugging Face](https://stackness.dev/tools/hugging-face-hub) under Apache 2.0 on 3 October, and community 4-bit builds run it on a 64GB Mac at about 50 tokens a second. Reflection AI's [Beam](https://stackness.dev/tools/reflection-beam) (501B total, 23B active) and [Mistral Large 4](https://stackness.dev/tools/mistral-large) (1.05T total, 52B active) are API previews whose weights are promised for later this month. On anything smaller than a data-center node, Kolibri is the only one of the three you can serve today.

## Three MoE launches in four days, and one download

The three were announced on 3, 5 and 6 October 2026, each as an open-weight mixture-of-experts model, and each drew a long Hacker News thread: more than 1,100 comments on [Mistral Large 4](https://news.ycombinator.com/item?id=49977979), 336 on [Kolibri](https://news.ycombinator.com/item?id=49942706) and 172 on [Beam](https://news.ycombinator.com/item?id=49969183), read 7 October. Only Kolibri shipped weights with its announcement.

I picked them because they share the launch week and the claim. I left out models that shipped weights earlier, such as GLM 5.3 Flash in August, dense models such as Qwen3.8 27B, and the small decision models covered in the [System One runtime post](https://stackness.dev/blog/how-do-you-run-system-one-decision-models-locally-ollaya-laya-mlx-and-the-runtime-slot). Every benchmark score below is vendor-reported unless a third party is named. Every memory figure marked as arithmetic is parameters times bytes per weight, which I calculated and nobody measured.

## Total parameters decide memory, active parameters decide speed

A mixture-of-experts model has to keep every expert in memory, because the router picks different experts for every token. Total parameters set how much memory the weights take. Active parameters set how much of that memory the runtime reads per token, which on a laptop or a single GPU is what limits speed. Kolibri's model card: "the full model must be held in memory even though only part of it is active at any time."

The arithmetic for the three, weights only, before KV cache and runtime overhead:

|  | Kolibri 1 | Beam | Mistral Large 4 | 
|---|---|---|---|
| Total parameters | 78.1B | 501B | 1.05T | 
| Active per token | 3.46B | 23B | 52B | 
| Weights at 16 bits | 156GB | 1,002GB | 2,100GB | 
| Weights at 8 bits | 78GB | 501GB | 1,050GB | 
| Weights at 4 bits | 39GB | 251GB | 525GB | 
| Weights at 2 bits | 20GB | 125GB | 263GB | 
| Read per token at 4 bits | 1.7GB | 11.5GB | 26GB | 

Real files run larger, because embeddings, norms and routers stay at higher precision: Kolibri's Q4_K_M GGUF is 47.5GB, not 39GB. The last row explains Kolibri's speed on modest hardware. Hob-forge measured 18.0 tokens a second with all its routed experts in system RAM and 6.4GB on a 12GB GPU, because each token touches so little of the model.

## Kolibri, Beam and Mistral Large 4, one fact line each

Kolibri 1 ships Apache 2.0 weights today, Beam is a waitlisted beta API with Apache 2.0 weights promised this month, and Mistral Large 4 is a public preview API with weights due by the end of October and no published licence. Each entry gives the same facts: size, context, licence, weights, price and what was measured.

### Kolibri 1

- **Developer and date:**[Aleph Alpha](https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/) , Heidelberg, 3 October 2026.
- **Size:** 78.1B total, 3.46B active; 384 routed experts per layer with 6 active per token, plus 1 shared expert, over 50 layers ([model card](https://huggingface.co/Aleph-Alpha/Kolibri-1) ).
- **Context:** 262,144 tokens native, validated to 1,048,576. English and German.
- **Licence:** Apache 2.0 on the weights and config files. The card says it "does not extend to underlying code, model architecture, parameter settings or any training method".
- **Weights:** on Hugging Face, not gated: FP8 at 78.86GB and BF16 at 156.23GB. No official GGUF, MLX or 4-bit build; community conversions fill the gap. The FP8 repo showed 5,775 downloads on 7 October.
- **API:** none found, and[OpenRouter](https://stackness.dev/tools/openrouter) does not list it.
- **Vendor hardware:** for FP8, at least 2 A100 80GB or 2 H100, or 1 H200, B200 or B300.
- **Vendor scores:** SWE-Bench Verified 66.4 and Terminal-Bench 2.1 27.7, run in Aleph Alpha's own harnesses. In the same table the dense Qwen3.8 27B beats it overall, 80.2 to 75.5 in English.
- **Pick it when:** you want weights on your own machine this week, need German, or need a licence your legal team already knows.

### Beam

- **Developer and date:**[Reflection AI](https://reflection.ai/blog/introducing-beam) , 5 October 2026.
- **Size:** 501B total, 23B active, 52 layers. Reflection does not give the expert count.
- **Context:** the launch post says 1M tokens. The[API docs](https://developers.reflection.ai/models) say 262,144 with 131,072 output, and that "the context window may change during the beta".
- **Licence:** Apache 2.0, announced: "This month, we will release the weights under an Apache 2.0 license".
- **Weights:** none, and no Hugging Face repo. "Early access" is a waitlist for the beta API, model`Beam-501B-A23B` behind an OpenAI-compatible endpoint, not a download.
- **Price:** none published.
- **Vendor scores:** SWE-bench Verified 80.9 and Terminal-Bench 2.1 80.1, with competitor scores taken from Artificial Analysis and DataCurve.
- **Pick it when:** the weights land and you have about 250GB for a 4-bit build. Until then the only option is the waitlist.

### Mistral Large 4

- **Developer and date:**[Mistral AI](https://mistral.ai/news/mistral-large-4/) , public preview on 6 October 2026.
- **Size:** 1.05T total and 52B active, plus a 1.6B vision encoder, per Mistral's[docs](https://docs.mistral.ai/models/mistral-large-4) . The launch post rounds to 1T and 49B active. The 675B and 41B active figures on some pages are Mistral Large 3's spec from December 2025.
- **Context:** 1M in the docs; OpenRouter and Artificial Analysis list 524,288. Text and image in, text out.
- **Licence:** not published; the docs say "Coming soon". VentureBeat[reports](https://venturebeat.com/technology/mistral-debuts-large-4-le-chonk-a-1-trillion-parameter-text-output-model-with-high-benchmarks-planned-for-open-weights-release) a custom Mistral licence. Large 3 shipped under Apache 2.0, and Medium 3.5 under a modified MIT licence with an exception for large companies.
- **Weights:** "Weights drop end of this month", says Mistral. VentureBeat gives 27 October; Mistral names no date.
- **API:**`mistral-large-4` at a list price of $1.36 per million input tokens and $4.18 per million output, currently on sale at $0.68 and $2.09 with no end date. The same sale price applies on OpenRouter and Ollama's cloud tier.
- **Measured:**[Artificial Analysis](https://artificialanalysis.ai/models/mistral-large-4) scores it 38 on its index, 64th of 225, at 116.1 output tokens a second and 1.46 seconds to first token on Mistral's API. It lists the model as proprietary until the weights ship.
- **Pick it when:** you want the model now and an API is fine. Self-hosting waits for the weights and about 525GB at 4-bit.

## Which local runtime serves each model?

Kolibri runs on patched or out-of-tree builds of vLLM, llama.cpp, MLX and ds4 as of 7 October 2026, and no stock release of any local runtime loads it yet. Beam and Mistral Large 4 run on no local runtime, because there are no weights. Ollama's `mistral-large-4:cloud` is a hosted model billed per token, not a download.

| Runtime | Kolibri 1 today | Evidence | 
|---|---|---|
| [vLLM](https://stackness.dev/tools/vllm) | Works with Aleph Alpha's plugin | [aleph-alpha-inference](https://stackness.dev/tools/aleph-alpha-inference) 1.0.0, pinned to vLLM 0.29.0; native support in[PR #60026](https://github.com/vllm-project/vllm/pull/60026) , open | 
| [llama.cpp](https://stackness.dev/tools/llama-cpp) | Community patch only | [Issue #29922](https://github.com/ggml-org/llama.cpp/issues/29922) open; patched builds from Hob-forge | 
| [MLX](https://stackness.dev/tools/mlx) | Out-of-tree launchers | mlx-lm [PR #1945](https://github.com/ml-explore/mlx-lm/pull/1945) open | 
| [Ollama](https://stackness.dev/tools/ollama) | No | [PR #18780](https://github.com/ollama/ollama/pull/18780) open; a community upload exists, unverified | 
| [LM Studio](https://stackness.dev/tools/lm-studio) | No | Not in its catalog | 
| [ds4](https://stackness.dev/tools/ds4) | Fork only, Metal and CPU | [PR #1183](https://github.com/antirez/ds4/pull/1183) open | 
| [SGLang](https://stackness.dev/tools/sglang) ,[Strata](https://stackness.dev/tools/strata) ,[Ollaya](https://stackness.dev/tools/ollaya) ,[Magnitude](https://stackness.dev/tools/magnitude) | No | No mention in code or docs | 

Strata serves one model and Ollaya serves decision models; the [llama.cpp alternatives post](https://stackness.dev/blog/llama-cpp-alternatives-for-running-local-agents-ds4-magnitude-and-strata-against-llama-cpp-vllm-and) covers what each engine is for.

What people measured on Kolibri in its first four days, each one person on one machine:

| Who | Hardware | Build | Speed | Peak memory | 
|---|---|---|---|---|
| [velaia](https://huggingface.co/velaia/Kolibri-1-MLX-4bit) | M1 Max, 64GB | MLX 4-bit | 52-56 tok/s | 44-48GB | 
| [eins78](https://huggingface.co/eins78/Kolibri-1-mlx-mixed-4-8-bit) | M4 Pro, 64GB | MLX mixed 4/8-bit | 50 tok/s | 45.4GB | 
| [audreyt](https://huggingface.co/audreyt/Kolibri-1-MLX-6bit) | M5 Max, 128GB | MLX 6-bit | 116.3 tok/s | 63.6GB | 
| [Hob-forge](https://huggingface.co/Hob-forge/Kolibri-1-GGUF) | RTX 5070 12GB, 128GB DDR5 | Q4_K_M, experts in RAM | 18.0 tok/s | 6.4GB VRAM | 
| Hob-forge | Same machine, CPU only | Q4_K_M | 12.3 tok/s | 46.6GB RAM | 
| [martianvoid](https://news.ycombinator.com/item?id=49943869) on HN | RTX Pro 6000 | FP8 | about 170 tok/s | Not given | 
| [ModelFit](https://modelfit.io/blog/kolibri-1-aleph-alpha-mac-memory-requirements/) | M4 MacBook Pro, 16GB | MLX 2-bit | Did not load | Needs 24.3GB | 

On Stackness, as of 7 October 2026, one real profile lists Ollama, one lists LM Studio, and none lists llama.cpp, vLLM, MLX, ds4 or any of the three models ([data sources](https://stackness.dev/about/data-sources)). The numbers are small; the [language models developers list on Stackness](https://stackness.dev/categories/llms) will show when that changes.

## Kolibri, Beam and Mistral Large 4 side by side

Kolibri 1 is the smallest of the three and the only one with weights out. Beam and Mistral Large 4 are six to thirteen times larger in total parameters and are still API previews, with weights due by 31 October 2026. Self-hosting either will take hundreds of gigabytes.

|  | Kolibri 1 | Beam | Mistral Large 4 | 
|---|---|---|---|
| Announced | 3 October 2026 | 5 October 2026 | 6 October 2026 | 
| Total / active | 78.1B / 3.46B | 501B / 23B | 1.05T / 52B | 
| Licence | Apache 2.0, shipped | Apache 2.0, announced | Not published | 
| Weights today | Yes, FP8 and BF16 | No, "later this month" | No, "end of this month" | 
| Use it today via | Your own hardware | Waitlisted beta API | Public API | 
| Context | 262K native, 1M validated | 1M claimed, 256K in the API | 1M in docs, 524K on OpenRouter | 
| Local runtimes | Patched vLLM, llama.cpp, MLX, ds4 | None | None | 
| Weights at 4-bit | about 39GB (47.5GB Q4_K_M file) | about 251GB | about 525GB | 

The DeepSWE comparison that has circulated since 6 October, 61.7 for Mistral Large 4 against 44.4 for Beam, joins two vendors' tables rather than one test run, as [CellCog](https://cellcog.ai/blog/mistral-large-4/) notes. No third party has scored all three on one benchmark yet.

## Which one for which machine

Kolibri is the only one of the three that runs on a machine below a data-center node as of 7 October 2026. Beam and Mistral Large 4 will need hundreds of gigabytes once their weights ship. The figures below are peak memory measured by the people in the table above, or arithmetic where marked on paper.

- **16GB to 32GB Mac or laptop:** none of the three. Kolibri's smallest build is 25.5GB, and it did not load on a 16GB Mac.
- **36GB to 48GB Mac:** Kolibri in MLX 2-bit or 3-bit, 26GB to 39GB peak.
- **64GB Mac:** Kolibri in MLX 4-bit, about 50 tokens a second and 44GB to 48GB peak.
- **128GB Mac:** Kolibri at 6-bit or 8-bit; 116 tokens a second at 6-bit on an M5 Max.
- **12GB to 24GB GPU with 64GB or more of RAM:** Kolibri Q4_K_M on patched llama.cpp with the experts in RAM, about 18 tokens a second.
- **96GB GPU, or one H200 or B200:** Kolibri FP8 as Aleph Alpha ships it, on vLLM with its plugin.
- **512GB Mac Studio:** wait for Beam, about 251GB at 4-bit on paper. Mistral Large 4 at 4-bit, about 525GB, does not fit, and a 3-bit build, about 394GB, would.
- **One 8xH100 node, 640GB:** Beam at FP8, about 501GB, on paper; Mistral Large 4 needs 4-bit or a bigger node.
- **No hardware, want the biggest model now:** the Mistral Large 4 API at $0.68 and $2.09 per million tokens while the sale lasts.

## Tools in this post

## Use any of these tools?

Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.

[Show my stack](https://stackness.dev/register)
