# Mistral Large 4 explained: the real price, the hardware to run 1 trillion parameters, and the 93% Cybench claim

> Source: <https://theainewsreport.com/2026-10-06-mistral-large-4-price-hardware-cybench-explained.html>
> Published: 2026-10-06 14:07:49+00:00

# Mistral Large 4 explained: the real price, the hardware to run 1 trillion parameters, and the 93% Cybench claim

Mistral put Mistral Large 4 into public preview on October 6 and says the open weights follow by the end of the month. This page reconciles the two prices on Mistral's own pages, works the cost against Claude, GPT-6.1 Sol and an open rival, does the memory math for running it yourself, and explains how a model can top a hacking benchmark while also refusing more cyber requests than any other open model.

**This explains reporting by**

[Mistral AI, "Introducing Mistral Large 4" (October 6, 2026)](https://mistral.ai/news/mistral-large-4/).
Read the original first:

[https://mistral.ai/news/mistral-large-4/](https://mistral.ai/news/mistral-large-4/)

## In one minute

- Mistral Large 4 is a mixture-of-experts model with 1.05 trillion total parameters, 49 billion active per token, a 1.6 billion parameter vision encoder and a 1 million token context window. It is in public preview on Mistral's API now.
- Mistral says it will release the weights by the end of October. Outlets including The Next Web report October 27. The license has not been published.
- Mistral's post lists $1.36 per million input tokens and $4.18 output. Its docs page shows those prices struck through and $0.68 and $2.09 charged instead. Mistral does not say there how long the lower price lasts.
- Mistral says it solves 93% of Cybench, 40 professional capture-the-flag tasks, while Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse. It also says its refusal rate on malicious cyber prompts is the highest of any open model it tested.
- Those two claims fit together: Cybench tasks look like sanctioned practice, while the refusal sets are plain malicious requests. Once the weights are public, anyone can retrain the refusals away.
- For a security team or an MSP, this is a strong defensive tool that you can run privately, and a reason to assume attackers will have the same thing in weeks.

## What Mistral shipped on October 6

Mistral Large 4, nicknamed "le Chonk" in the post, is in public preview through Mistral's API and Studio. The docs page lists the model id as mistral-large-4 and gives the size as 49 billion active parameters, 1.05 trillion total, plus a 1.6 billion parameter vision encoder. The context window is 1 million tokens.

Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, and serves the preview from the same machines. During reinforcement learning, it says one training run produced about 33 billion tokens a day.

The weights are promised "by the end of the month." The Next Web and other outlets report October 27, and Alpha Signal reports FP8 and FP4 checkpoints. Mistral has not published the license.

Until then, Mistral says it is red-teaming the model with cybersecurity leaders, vetted partners and state authorities, who get "the same model with reduced moderation and expanded cyber capabilities."

## Two prices on two Mistral pages

The announcement lists $1.36 per million input tokens and $4.18 per million output tokens. The docs page shows the same two numbers struck through, next to $0.68 and $2.09. Cached input shows $0.14 struck through and $0.07 charged. In other words, the docs page charges exactly half of list.

Neither page, as of October 6, says why or for how long. Treat the half price as a preview offer that can end, and budget on the list price.

Here is one sample job to make the numbers concrete: 10 million input tokens and 2 million output tokens, roughly a busy week of an agent reading code and writing patches. A token is about three quarters of a word.

- Mistral Large 4 at the docs price: $6.80 in plus $4.18 out, about $11.
- Mistral Large 4 at list: $13.60 in plus $8.36 out, about $22.
- GPT-6.1 Sol at $2 and $10: $20 plus $20, $40.
- Claude Sonnet 5.5 at $2 and $10: also $40.
- Claude Opus 5.5 at $4 and $20: $40 plus $40, $80.
- Xiaomi's open MiMo-V2.6-Pro at $0.435 and $0.87: $4.35 plus $1.74, about $6.

So at list price Mistral Large 4 costs a bit over half of Sol or Sonnet for this mix, and at the docs price about a quarter. It is still well above the cheapest Chinese open model. These are list prices only. They ignore caching, batch discounts and how many tokens each model spends thinking, which can swing the real bill a lot.

## What 1 trillion parameters with 49 billion active means

A mixture-of-experts model splits most of its layers into many smaller expert blocks. For each token, a router picks a few experts and skips the rest. Mistral Large 4 holds 1.05 trillion parameters but uses about 49 billion of them per token.

That split has two effects. The compute per token is close to a 49 billion parameter dense model, which is why it can be priced well below its size. But the memory bill is the full 1.05 trillion, because any expert might be needed for the next token.

The memory math for self-hosting, counting weights only:

- FP8 stores one byte per parameter, so about 1.05 TB of weights.
- FP4 stores half a byte per parameter, so about 525 GB.
- An NVIDIA B200 has roughly 180 to 192 GB of memory and a B300 about 288 GB.

At FP4, the weights alone fill about three B200s, and the cache that holds a long conversation needs more on top. That is why outlets report 4 to 8 B200 or B300 GPUs. At FP8, plan on a full 8-GPU server. For comparison, Aleph Alpha's Kolibri-1, released October 3, needs about 78 GB at FP8 and fits on one H200 or B200.

If you are an MSP thinking about private hosting for clients, the honest reading is that this is a datacenter model. The interesting part for most shops is the API price today and the option to move to a private cloud host later.

## The cyber claim, and what Cybench measures

Cybench is a public benchmark from Stanford researchers: 40 professional-level capture-the-flag tasks taken from four competitions. The model works in a real environment through shell commands and network calls and must recover a hidden answer, the flag.

Mistral says Large 4 solves 93% of those tasks and calls it one of the highest scores reported for an open-weight model. It adds that "several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task." On a separate Artificial Analysis test that asks a model to reproduce a real vulnerability in open-source software and then patch it, Mistral says it scores 82%, the highest of any model.

The same post also says Large 4's refusal rate on cyber prompts from JailbreakBench, StrongREJECT and AgentHarm is higher than every open model Mistral compared. Those sets are written as clearly harmful requests.

Both can be true. A capture-the-flag task reads like an authorized exercise: here is a target, find the flag. A jailbreak set reads like someone asking for harm. Mistral has tuned the model to say yes to the first kind and no to the second. Closed labs have drawn the line further back, so their models decline both.

The catch is that refusals are a trained behavior. Once the weights are downloadable, anyone with enough GPUs can fine-tune that behavior away. Mistral's argument is that defenders need this capability because attackers already jailbreak closed models. Whether you agree or not, the practical result is the same: by November, a strong open model with 93% on Cybench will be in many hands.

## Where it lands against other open models

Mistral's coding numbers, some of them run privately by Artificial Analysis before a public harness launch: 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4, for a combined Coding Agent Index of 49.8%, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.

The Next Web's chart puts its DeepSWE score at 62% next to GLM-5.3 at 61%, DeepSeek-V4-Pro at 57% and Qwen 3.8 Max at 51%. That is a narrow lead, not a gap.

In a blind Surge AI review of coding output, it ranked second of five at 3.74 out of 5, behind Claude Opus 5 at 4.22 and ahead of Kimi K3, GLM-5.3 and GLM-5.2. On AutomationBench, 657 business workflows across apps like Gmail, Sheets, Slack and Salesforce, Mistral reports 59.9%.

Vision is its strongest area by Mistral's account. It reports 42% on the Dense 200 visual grounding test against 41% for GPT-6 Astra.

Every one of these numbers comes from Mistral or from tests Mistral arranged. Independent leaderboards will fill in over the next few weeks.

## What this means for a security team or an MSP

If your current model refuses legitimate defensive work, such as reproducing a known flaw in your own test system, this is the first open model that claims to do that work well. You can try it in the preview today, and later run it inside your own network so client code never leaves.

The flip side matters more for most readers. A model that reproduces real vulnerabilities at this level will be downloadable this month. Patch windows that were tight are about to get tighter, and the side systems that attackers like, such as old portals and forgotten file servers, are where an automated tool earns its keep.

- Treat known-exploited flaws as same-week patches, not monthly ones.
- Keep a list of every internet-facing system you or your clients run, including the small ones.
- If you use the API, check Mistral's data and retention terms before you send client code.

## Who is affected

| Case | Status | 
|---|---|
| Security teams and MSPs | A strong defensive model you can run privately after the weights ship. Test it now in preview. | 
| Anyone with internet-facing systems | Expect attackers to have the same capability within weeks. Shorten patch windows. | 
| Teams paying for closed models on coding work | List price is about half of GPT-6.1 Sol or Claude Sonnet 5.5 for a typical mix. Benchmark it on your own tasks. | 
| Teams that want to self-host | Plan for a multi-GPU server: about 525 GB of weights at FP4 and about 1.05 TB at FP8. | 

## What to do

- Run one defensive task your current model refuses through the Mistral Large 4 preview, and record the result and the cost.
- Budget on the $1.36 and $4.18 list price, not the half price on the docs page.
- Read Mistral's data terms before sending client code to the API.
- Wait for the license before you plan a self-hosted deployment.
- Move known-exploited vulnerabilities to same-week patching on every client system you manage.

## What is still unknown

- The license. Mistral has not published the terms for the weights.
- The exact weight release date. Mistral says the end of October; outlets report October 27.
- How long the half price on the docs page lasts, and why it differs from the announcement.
- Independent cyber results. The 93% Cybench and 82% vulnerability reproduction numbers come from Mistral's post, and some coding scores were run privately by Artificial Analysis before a public harness launch.
- How easily the refusal behavior can be removed from the open weights. That is a general property of open models, not something Mistral has tested in public.
- The GPU count for self-hosting. The 4 to 8 GPU figure is from press coverage; Mistral's post does not give hardware requirements.

## Sources

[AI News Report](https://theainewsreport.com/)· every headline, every morning.
