cd /news/large-language-models/mistral-large-4-vs-beam-vs-kolibri-w… · home › topics › large-language-models › article
[ARTICLE · art-147124] src=stackness.dev ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Mistral Large 4 vs Beam vs Kolibri: which open weights you can download, and what your hardware can serve

Aleph Alpha released Kolibri 1, a 78.1B-parameter mixture-of-experts model with 3.46B active parameters, under Apache 2.0 on Hugging Face on 3 October 2026, making it the only one of three open-weight models announced that week whose weights are downloadable. Reflection AI's Beam (501B total, 23B active) and Mistral Large 4 (1.05T total, 52B active) remain API previews with weights promised later in October, and community 4-bit builds of Kolibri run on a 64GB Mac at about 50 tokens a second. Kolibri's Q4_K_M GGUF file is 47.5GB, and Hob-forge measured 18.0 tokens a second with routed experts in system RAM and 6.4GB on a 12GB GPU.

by read10 min views2 publishedOct 7, 2026
Mistral Large 4 vs Beam vs Kolibri: which open weights you can download, and what your hardware can serve
Image: Stackness (auto-discovered)

Mistral Large 4 vs Beam vs Kolibri, as of 7 October 2026: only Kolibri has weights you can download. Aleph Alpha put the 78.1B-parameter model, 3.46B of them active per token, on Hugging Face under Apache 2.0 on 3 October, and community 4-bit builds run it on a 64GB Mac at about 50 tokens a second. Reflection AI's Beam (501B total, 23B active) and Mistral Large 4 (1.05T total, 52B active) are API previews whose weights are promised for later this month. On anything smaller than a data-center node, Kolibri is the only one of the three you can serve today.

Three MoE launches in four days, and one download #

The three were announced on 3, 5 and 6 October 2026, each as an open-weight mixture-of-experts model, and each drew a long Hacker News thread: more than 1,100 comments on Mistral Large 4, 336 on Kolibri and 172 on Beam, read 7 October. Only Kolibri shipped weights with its announcement.

I picked them because they share the launch week and the claim. I left out models that shipped weights earlier, such as GLM 5.3 Flash in August, dense models such as Qwen3.8 27B, and the small decision models covered in the System One runtime post. Every benchmark score below is vendor-reported unless a third party is named. Every memory figure marked as arithmetic is parameters times bytes per weight, which I calculated and nobody measured.

Total parameters decide memory, active parameters decide speed #

A mixture-of-experts model has to keep every expert in memory, because the router picks different experts for every token. Total parameters set how much memory the weights take. Active parameters set how much of that memory the runtime reads per token, which on a laptop or a single GPU is what limits speed. Kolibri's model card: "the full model must be held in memory even though only part of it is active at any time."

The arithmetic for the three, weights only, before KV cache and runtime overhead:

Kolibri 1 Beam Mistral Large 4
Total parameters 78.1B 501B 1.05T
Active per token 3.46B 23B 52B
Weights at 16 bits 156GB 1,002GB 2,100GB
Weights at 8 bits 78GB 501GB 1,050GB
Weights at 4 bits 39GB 251GB 525GB
Weights at 2 bits 20GB 125GB 263GB
Read per token at 4 bits 1.7GB 11.5GB 26GB

Real files run larger, because embeddings, norms and routers stay at higher precision: Kolibri's Q4_K_M GGUF is 47.5GB, not 39GB. The last row explains Kolibri's speed on modest hardware. Hob-forge measured 18.0 tokens a second with all its routed experts in system RAM and 6.4GB on a 12GB GPU, because each token touches so little of the model.

Kolibri, Beam and Mistral Large 4, one fact line each #

Kolibri 1 ships Apache 2.0 weights today, Beam is a waitlisted beta API with Apache 2.0 weights promised this month, and Mistral Large 4 is a public preview API with weights due by the end of October and no published licence. Each entry gives the same facts: size, context, licence, weights, price and what was measured.

Kolibri 1

- **Developer and date:**[Aleph Alpha](https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/) , Heidelberg, 3 October 2026.
- **Size:** 78.1B total, 3.46B active; 384 routed experts per layer with 6 active per token, plus 1 shared expert, over 50 layers ([model card](https://huggingface.co/Aleph-Alpha/Kolibri-1) ).
  • Context: 262,144 tokens native, validated to 1,048,576. English and German.
  • Licence: Apache 2.0 on the weights and config files. The card says it "does not extend to underlying code, model architecture, parameter settings or any training method".
  • Weights: on Hugging Face, not gated: FP8 at 78.86GB and BF16 at 156.23GB. No official GGUF, MLX or 4-bit build; community conversions fill the gap. The FP8 repo showed 5,775 downloads on 7 October.
  • API: none found, andOpenRouter does not list it.
  • Vendor hardware: for FP8, at least 2 A100 80GB or 2 H100, or 1 H200, B200 or B300.
  • Vendor scores: SWE-Bench Verified 66.4 and Terminal-Bench 2.1 27.7, run in Aleph Alpha's own harnesses. In the same table the dense Qwen3.8 27B beats it overall, 80.2 to 75.5 in English.
  • Pick it when: you want weights on your own machine this week, need German, or need a licence your legal team already knows.

Beam

  • Developer and date:Reflection AI , 5 October 2026.
  • Size: 501B total, 23B active, 52 layers. Reflection does not give the expert count.
  • Context: the launch post says 1M tokens. TheAPI docs say 262,144 with 131,072 output, and that "the context window may change during the beta".
  • Licence: Apache 2.0, announced: "This month, we will release the weights under an Apache 2.0 license".
  • Weights: none, and no Hugging Face repo. "Early access" is a waitlist for the beta API, modelBeam-501B-A23B behind an OpenAI-compatible endpoint, not a download.
  • Price: none published.
  • Vendor scores: SWE-bench Verified 80.9 and Terminal-Bench 2.1 80.1, with competitor scores taken from Artificial Analysis and DataCurve.
  • Pick it when: the weights land and you have about 250GB for a 4-bit build. Until then the only option is the waitlist.

Mistral Large 4

  • Developer and date:Mistral AI , public preview on 6 October 2026.
  • Size: 1.05T total and 52B active, plus a 1.6B vision encoder, per Mistral'sdocs . The launch post rounds to 1T and 49B active. The 675B and 41B active figures on some pages are Mistral Large 3's spec from December 2025.
  • Context: 1M in the docs; OpenRouter and Artificial Analysis list 524,288. Text and image in, text out.
  • Licence: not published; the docs say "Coming soon". VentureBeatreports a custom Mistral licence. Large 3 shipped under Apache 2.0, and Medium 3.5 under a modified MIT licence with an exception for large companies.
  • Weights: "Weights drop end of this month", says Mistral. VentureBeat gives 27 October; Mistral names no date.
  • API:mistral-large-4 at a list price of $1.36 per million input tokens and $4.18 per million output, currently on sale at $0.68 and $2.09 with no end date. The same sale price applies on OpenRouter and Ollama's cloud tier.
  • Measured:Artificial Analysis scores it 38 on its index, 64th of 225, at 116.1 output tokens a second and 1.46 seconds to first token on Mistral's API. It lists the model as proprietary until the weights ship.
  • Pick it when: you want the model now and an API is fine. Self-hosting waits for the weights and about 525GB at 4-bit.

Which local runtime serves each model? #

Kolibri runs on patched or out-of-tree builds of vLLM, llama.cpp, MLX and ds4 as of 7 October 2026, and no stock release of any local runtime loads it yet. Beam and Mistral Large 4 run on no local runtime, because there are no weights. Ollama's mistral-large-4:cloud is a hosted model billed per token, not a download.

Runtime Kolibri 1 today Evidence
vLLM Works with Aleph Alpha's plugin aleph-alpha-inference 1.0.0, pinned to vLLM 0.29.0; native support inPR #60026 , open
| [llama.cpp](https://stackness.dev/tools/llama-cpp) | Community patch only | [Issue #29922](https://github.com/ggml-org/llama.cpp/issues/29922) open; patched builds from Hob-forge | 
| [MLX](https://stackness.dev/tools/mlx) | Out-of-tree launchers | mlx-lm [PR #1945](https://github.com/ml-explore/mlx-lm/pull/1945) open | 
| [Ollama](https://stackness.dev/tools/ollama) | No | [PR #18780](https://github.com/ollama/ollama/pull/18780) open; a community upload exists, unverified | 
| [LM Studio](https://stackness.dev/tools/lm-studio) | No | Not in its catalog | 
| [ds4](https://stackness.dev/tools/ds4) | Fork only, Metal and CPU | [PR #1183](https://github.com/antirez/ds4/pull/1183) open | 

| SGLang ,Strata ,Ollaya ,Magnitude | No | No mention in code or docs |

Strata serves one model and Ollaya serves decision models; the llama.cpp alternatives post covers what each engine is for.

What people measured on Kolibri in its first four days, each one person on one machine:

| Who | Hardware | Build | Speed | Peak memory |

|---|---|---|---|---|
| [velaia](https://huggingface.co/velaia/Kolibri-1-MLX-4bit) | M1 Max, 64GB | MLX 4-bit | 52-56 tok/s | 44-48GB | 
| [eins78](https://huggingface.co/eins78/Kolibri-1-mlx-mixed-4-8-bit) | M4 Pro, 64GB | MLX mixed 4/8-bit | 50 tok/s | 45.4GB | 
| [audreyt](https://huggingface.co/audreyt/Kolibri-1-MLX-6bit) | M5 Max, 128GB | MLX 6-bit | 116.3 tok/s | 63.6GB | 
| [Hob-forge](https://huggingface.co/Hob-forge/Kolibri-1-GGUF) | RTX 5070 12GB, 128GB DDR5 | Q4_K_M, experts in RAM | 18.0 tok/s | 6.4GB VRAM | 

| Hob-forge | Same machine, CPU only | Q4_K_M | 12.3 tok/s | 46.6GB RAM | | martianvoid on HN | RTX Pro 6000 | FP8 | about 170 tok/s | Not given |

| ModelFit | M4 MacBook Pro, 16GB | MLX 2-bit | Did not load | Needs 24.3GB | On Stackness, as of 7 October 2026, one real profile lists Ollama, one lists LM Studio, and none lists llama.cpp, vLLM, MLX, ds4 or any of the three models (data sources). The numbers are small; the language models developers list on Stackness will show when that changes.

Kolibri, Beam and Mistral Large 4 side by side #

Kolibri 1 is the smallest of the three and the only one with weights out. Beam and Mistral Large 4 are six to thirteen times larger in total parameters and are still API previews, with weights due by 31 October 2026. Self-hosting either will take hundreds of gigabytes.

Kolibri 1 Beam Mistral Large 4
Announced 3 October 2026 5 October 2026 6 October 2026
Total / active 78.1B / 3.46B 501B / 23B 1.05T / 52B
Licence Apache 2.0, shipped Apache 2.0, announced Not published
Weights today Yes, FP8 and BF16 No, "later this month" No, "end of this month"
Use it today via Your own hardware Waitlisted beta API Public API
Context 262K native, 1M validated 1M claimed, 256K in the API 1M in docs, 524K on OpenRouter
Local runtimes Patched vLLM, llama.cpp, MLX, ds4 None None
Weights at 4-bit about 39GB (47.5GB Q4_K_M file) about 251GB about 525GB

The DeepSWE comparison that has circulated since 6 October, 61.7 for Mistral Large 4 against 44.4 for Beam, joins two vendors' tables rather than one test run, as CellCog notes. No third party has scored all three on one benchmark yet.

Which one for which machine #

Kolibri is the only one of the three that runs on a machine below a data-center node as of 7 October 2026. Beam and Mistral Large 4 will need hundreds of gigabytes once their weights ship. The figures below are peak memory measured by the people in the table above, or arithmetic where marked on paper.

  • 16GB to 32GB Mac or laptop: none of the three. Kolibri's smallest build is 25.5GB, and it did not load on a 16GB Mac.
  • 36GB to 48GB Mac: Kolibri in MLX 2-bit or 3-bit, 26GB to 39GB peak.
  • 64GB Mac: Kolibri in MLX 4-bit, about 50 tokens a second and 44GB to 48GB peak.
  • 128GB Mac: Kolibri at 6-bit or 8-bit; 116 tokens a second at 6-bit on an M5 Max.
  • 12GB to 24GB GPU with 64GB or more of RAM: Kolibri Q4_K_M on patched llama.cpp with the experts in RAM, about 18 tokens a second.
  • 96GB GPU, or one H200 or B200: Kolibri FP8 as Aleph Alpha ships it, on vLLM with its plugin.
  • 512GB Mac Studio: wait for Beam, about 251GB at 4-bit on paper. Mistral Large 4 at 4-bit, about 525GB, does not fit, and a 3-bit build, about 394GB, would.
  • One 8xH100 node, 640GB: Beam at FP8, about 501GB, on paper; Mistral Large 4 needs 4-bit or a bigger node.
  • No hardware, want the biggest model now: the Mistral Large 4 API at $0.68 and $2.09 per million tokens while the sale lasts.

Tools in this post #

Use any of these tools? #

Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.

Show my stack

── more in #large-language-models 4 stories · sorted by recency
── more on @aleph alpha 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mistral-large-4-vs-b…] indexed:0 read:10min 2026-10-07 · —