# Mistral ships a trillion-parameter model, Large 4, in API preview, with weights due at month's end

> Source: <https://www.thenewway.ai/mistral-ships-a-trillion-parameter-model/>
> Published: 2026-10-07 16:31:49+00:00

# Mistral ships a trillion-parameter model, Large 4, in API preview, with weights due at month's end

A trillion-parameter model in preview, scoring 61.7% on the DeepSWE coding benchmark, with weights due this month. Plus Codex users vote for a usage reset.

Mistral put a public preview of Large 4 on Mistral Studio Tuesday: a trillion-parameter model whose weights it says drop at the end of the month, scoring 61.7% on the DeepSWE coding benchmark. Separately, OpenAI's Codex team reset usage limits after 76% of 74,565 poll voters asked for it, on day two of a 28-day pledge to ship an improvement or a reset every day. GitHub shipped two changes to how teams build: a fix for a metrics bug that was hiding Copilot's agent activity, and general availability for stacked pull requests. And Claude Code's newest patch closes a real permission-prompt gap, buried in 92 bullets of plumbing.

## What changed this week

### Mistral ships a trillion-parameter model, Large 4, in API preview, with weights due at month's end

Mistral put a public preview of Large 4, "le Chonk," on Mistral Studio Tuesday: a 1-trillion-parameter, natively multimodal model with 49 billion active parameters. On the coding benchmarks Mistral reports, it scores 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4. In a blind human evaluation it ranked second of five, behind only Claude Opus 5. The louder claim is security. On one test in the independent Artificial Analysis Cyber Index, reproducing and patching a real vulnerability, it scores 82%, and Mistral says Opus 5.5 and GPT-6 Astra "score near zero" because they refuse. The same afternoon, [Anthropic widened access to Opus 5.5](https://x.com/AnthropicAI/status/2107546569654636883?ref=thenewway.ai) for verified security professionals doing defensive work. Weights drop at month's end.

### Codex users vote 76% for a usage reset, and OpenAI's Codex team grants it

Tibo Sottiaux of OpenAI's Codex team pledged Sunday that for 28 days, each day brings an improvement or "a full reset," meaning a reset of usage limits. Day two brought both. He listed four updates, among them the [Decisions API](https://developers.openai.com/api/docs/guides/decisions?ref=thenewway.ai), which returns a probability, a category or a score "about 10x faster than the Responses API" at $0.10 per million input tokens. He polled users, and [76% of 74,565 voted](https://x.com/thsottiaux/status/2107576143285219799?ref=thenewway.ai) "needs a reset." His reply: "the reset has been processed."

### GitHub says a Copilot SDK migration undercounted agent activity in usage metrics

GitHub explained why some teams watched Copilot's agent activity fall in their usage dashboards while overall usage kept growing. "Several IDEs recently moved Copilot agent sessions to the Copilot SDK," and those sessions "didn't identify which IDE they came from," so the metrics dropped most of that activity or miscounted it as CLI use. VS Code 1.139 has the fix now; Visual Studio, JetBrains, Eclipse and Xcode follow through November. The fix doesn't back-fill history. GitHub's own warning is blunt: "their activity can't be recovered later."

### GitHub's stacked pull requests reach general availability — approvals now survive a rebase

GitHub's stacked pull requests, the feature for splitting a large change into smaller reviewable ones, are generally available after a preview in which "repositories using stacks have seen a 9% increase in merged code compared to peers" and the top 1% of repos saw "a 5% improvement in time-to-merge." The GA release fixes a real workflow gap: rebasing a stack used to risk dismissing reviewers' approvals, and now keeps them in place even in repos that dismiss stale approvals on every other push. That's the real fix.

### Claude Code 2.1.292 fixes a permission-prompt bypass for network file reads

Claude Code 2.1.292 closes a real gap: a security fix stops "PreToolUse hook approvals and auto mode" from "bypassing the permission prompt for file reads from network (UNC) paths." The same release fixes a tampered on-disk settings cache that could have disabled Claude Code's built-in policy plugin, and a sandboxed command that could read staged /ultrareview upload files it shouldn't see. The other 89 bullets are mod and plugin-hook plumbing. Not security fixes.

Know someone who'd want this in their inbox? Forward it — that's how this grows. And if we got something wrong, or you think we buried the real story today, hit reply. A person reads every one.

*The New Way is written with AI. It gathers the day's stories, checks them against their sources and drafts every summary. A person decides what runs and reviews every issue before we hit send.*
