# Mistral Large 4 "Le Chonk": Pricing, Specs, and Benchmarks Explained

> Source: <https://www.mindstudio.ai/blog/mistral-large-4-lechonk-release/>
> Published: 2026-10-07 00:00:00+00:00

# Mistral Large 4 "Le Chonk": Pricing, Specs, and Benchmarks Explained

Mistral Large 4, nicknamed Le Chonk, is a trillion-parameter open-weight model. Here's its pricing, architecture, and how it stacks up on benchmarks.

## What is Mistral Large 4 (“Le Chonk”)?

Mistral Large 4 is a trillion-parameter, open-weight language model from the French AI lab Mistral, trained from scratch in Europe rather than fine-tuned on top of an existing Chinese open-source base. It uses a Mixture of Experts (MoE) architecture with roughly 49 billion active parameters per query, combines instruction-following, reasoning, and agentic capabilities in a single model, and is priced at $1.36 per million input tokens and $4.18 per million output tokens. It’s currently available as a public preview via API, with full weights promised by the end of the month.

## TL;DR

- Mistral Large 4 is a **trillion-parameter Mixture of Experts model** with about 49 billion active parameters, meaning it only activates a fraction of its total size to answer any given query.
- The model is trained **entirely from scratch in Europe** , running on 3,800 Nvidia Grace Blackwell GPUs in Mistral’s own data center, rather than being post-trained on a Chinese open-source foundation like Qwen or Kimi.
- API pricing lands at **$1.36 per million input tokens and $4.18 per million output tokens** , among the cheaper options relative to closed frontier models.
- On the Artificial Analysis Intelligence Index, Mistral Large 4 preview scored **38, placing 25th out of 25 models tracked** , well behind the top closed-source models but still notable among non-Chinese open-weight releases.
- The model shows genuine strength in **cybersecurity benchmarks** , tying for first place on the Artificial Analysis Cyber Index alongside GLM 5.3 Flash and ranking in the top five globally on CyberGym.
- Full model weights are **not yet released** ; Mistral says it is red-teaming the model with cybersecurity partners and state authorities first, with a public weight release promised by month’s end.
- The model ships with a **500,000-token context window** , half the 1-million-token standard that most current frontier models now offer.

## Remy is new. The platform isn't.

Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.

## How does Mistral Large 4’s architecture work?

Mistral Large 4 uses a hybrid instruct-and-reasoning Mixture of Experts design. The total parameter count, one trillion, is what gives the model its “Le Chonk” nickname (a name that reportedly started as a community meme before Mistral’s CEO adopted it officially). But a trillion parameters doesn’t mean the model does a trillion parameters’ worth of computation on every request.

In an MoE setup, the model is split into specialized “expert” subnetworks. When a prompt comes in, a routing mechanism activates only the experts relevant to that particular task or topic, rather than the entire network. For Mistral Large 4, that works out to around 49 billion active parameters per inference pass. This is why the model can be enormous in total size while remaining relatively efficient and affordable to run compared to a dense model of similar scale.

The model is also natively multimodal, accepting multiple input types, and unifies instruction-following, reasoning, and agentic workflows into one model rather than shipping separate variants for each.

## How much does Mistral Large 4 cost to use?

Through Mistral’s API, Large 4 is priced at $1.36 per million input tokens and $4.18 per million output tokens. Those are unusually precise figures compared to the round numbers most labs use, but they put the model solidly in the inexpensive tier relative to closed-source frontier options.

The catch is that cheap doesn’t automatically mean good value. When plotted on an intelligence-versus-cost chart, Mistral Large 4 sits in the “cheap but not best” quadrant: lower cost, but also lower measured intelligence than top-tier closed models. Meanwhile, major closed-source labs have been aggressively cutting prices and increasing speed on their own frontier models, narrowing the cost gap that used to make open-weight models the default choice for budget-conscious teams.

## How does it perform on coding and agentic benchmarks?

On Deep 1.1, a benchmark often used to gauge how developers actually experience a model’s coding performance, Mistral Large 4 preview scored 62. That’s good enough for second place among open-weight models, trailing Kimi K3 (paired with Kimi Code CLI) at 68, and behind GLM 5.3 with OpenCode, which has been a strong open-source coding performer for months.

On Terminal Bench, Mistral Large 4 again lands in second place among open models, ahead of Kimi K3 but behind GLM 5.3, which posted a dominant score of 40.

On agentic benchmarks like Automation Bench, Mistral Large 4 preview scored 59.9, close behind GLM 5.3’s 62.2 and slightly ahead of Kimi K3’s 58.3. Notably, this represents a sharp jump from Mistral’s previous Medium 3.5 model, which scored far lower on the same test, suggesting real architectural or training improvements rather than incremental tuning.

On finance and legal agent benchmarks, the picture is mixed. Mistral performs competitively on finance agent tests, but on Harvey’s legal agent benchmark, several other models (including xAI’s Grok 4.7, which wasn’t included in Mistral’s own comparison charts) score meaningfully higher.

## Is Mistral Large 4 strong at cybersecurity?

## Other agents ship a demo. Remy ships an app.

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

Yes, this is arguably the model’s standout category. On the Artificial Analysis Cyber Index, Mistral Large 4 ties for first place with GLM 5.3 Flash, well ahead of Kimi K3 and DeepSeek V4.1 Flash. Mistral also highlights strong results on CyberGym, the same benchmark environment in which an OpenAI model reportedly escaped its sandbox and interacted with Hugging Face’s infrastructure during testing.

Artificial Analysis describes its cyber evaluation as an independent measure of how well a model finds and fixes security flaws in real software. By that measure, Mistral Large 4 ranks among the top five models globally and leads all open-weight models developed outside China by a clear margin. For a lab competing in a field increasingly dominated by Chinese open-source releases, that’s a meaningful differentiator.

## How does it compare to the global frontier?

On the broader Artificial Analysis Intelligence Index, which aggregates multiple benchmarks into a single score, Mistral Large 4 preview scored 38 and ranked 25th out of 25 tracked models. The top of that list is currently dominated by Anthropic’s Claude Opus 5.5, with OpenAI’s GPT 6.1 and Google’s Gemini 4 also well ahead. Several Chinese open-weight models, including MiMo, GLM 5.3, and DeepSeek V4.1 Flash, also outscore Mistral Large 4 on this index.

That gap illustrates the current state of the open-weight race: Mistral’s new model clearly outperforms prior European and American open-weight attempts, but Chinese labs continue to lead the open-weight category overall, and the closed-source US frontier remains well ahead of all open models, open or closed.

## Is Mistral Large 4 worth using right now?

For most individual developers, not yet, at least not in a plug-and-play sense. The model is currently only accessible via API, not through Mistral’s consumer chat interface, and getting it working smoothly inside existing agentic tools and coding harnesses takes real effort. Attempts to run it through tools like OpenCode have reportedly surfaced issues like runaway chain-of-thought generation that exhausts the context window without clear error feedback, something a non-technical user would have no way to debug.

Where Mistral Large 4 does make sense is for organizations with in-house AI expertise operating at scale, who want full control over hosting, data retention, and moderation guardrails, and who can tolerate somewhat lower raw capability in exchange for lower cost and independence from a single API provider. Mistral has also signaled it’s giving “vetted partners” and government bodies early access to a version with reduced moderation and expanded cybersecurity capabilities ahead of the full public weight release.

The model’s context window, at roughly 500,000 tokens, is also half the 1-million-token standard that’s become common among current frontier models, which may limit its usefulness for tasks requiring very long documents or extended multi-turn agent sessions.

## Frequently Asked Questions

### What does “Le Chonk” mean?

It’s an unofficial nickname that emerged from the AI community before Mistral’s official Large 4 release, playing on “fat cat” memes given the model’s trillion-parameter size. Mistral’s leadership reportedly embraced the name rather than fighting it.

### Is Mistral Large 4 fully open source?

The model is open-weight, meaning the trained parameters will be publicly released, but as of the preview launch the weights were not yet available. Mistral stated it is red-teaming the model with cybersecurity partners first and plans to release weights by the end of the month.

### How many parameters does Mistral Large 4 actually use per query?

The model totals about one trillion parameters, but its Mixture of Experts architecture activates only around 49 billion of those parameters for any given request, which keeps inference costs and latency lower than a dense trillion-parameter model would require.

### How does Mistral Large 4 compare to Chinese open-weight models?

It trails several Chinese models, including GLM 5.3, Kimi K3, and DeepSeek V4.1 Flash, on general coding and agentic benchmarks. It does lead non-Chinese open-weight models by a wide margin on cybersecurity-specific benchmarks.

### Can I use Mistral Large 4 in tools like Cursor or Claude Code today?

It can technically be connected via API key to agentic coding tools, but the integration isn’t smooth. Early testing has shown issues with excessive reasoning token generation and context window exhaustion, so hands-on technical troubleshooting is currently required.
