Mistral Large 4 (Le Chat): Hands-On Testing of the New Flagship Mistral AI released Mistral Large 4, its new flagship mixture-of-experts model with roughly a trillion total parameters and about 50 billion active at inference, in public preview via its API and Le Chat, with open weights promised by the end of the month. Hands-on testing found the model tied GLM 5.3 Flash for the top spot on a cybersecurity index, trailed only Kimi K3 on a SWE-bench-style coding benchmark, and beat most open models, but fell behind GLM 5.3 on terminal-based tasks and cross-app business workflows. Multilingual translation held up for Mandarin, Arabic, Japanese and Thai but broke down on low-resource languages such as Gujarati, and image-grounded HTML/CSS generation from a photo was convincing overall though fire intensity did not fully match the reference image. Mistral Large 4 Le Chat : Hands-On Testing of the New Flagship Mistral Large 4, aka Le Chat, tested on reasoning, vision, code security, and multilingual tasks against rivals like GLM 5.3. What is Mistral Large 4? Mistral Large 4 is the latest flagship model from Mistral AI, released as a public preview under the internal nickname “Le Chonk.” It is a large mixture-of-experts model, roughly a trillion parameters total with about 50 billion active at inference time, trained in Mistral’s own European data centers. It is available now through Mistral’s API and Le Chat, with open weights promised for release by the end of the month. TL;DR - Le Chonk is Mistral’s new flagship, a roughly trillion-parameter mixture-of-experts model with about 50 billion active parameters, currently in public preview via API with open weights coming later. - On vulnerability detection , the model found a planted IDOR insecure direct object reference bug in a Flask app, correctly diagnosing that a delete endpoint checked ownership while a get endpoint didn’t, then flagged plaintext passwords and a hardcoded session secret on top of it. - Independent benchmark charts shown in testing put the model at or near the top on SWE-bench style coding tasks and a cybersecurity index, while GLM 5.3 led on terminal-based tasks and cross-app business workflows. - Visual and scientific reasoning held up well: the model correctly solved a multi-person electrical shock puzzle from an image, including a trick case most people get wrong. - Social reasoning from a screenshot was a standout result, correctly untangling a WhatsApp conversation where timestamps and slang created a misunderstanding between a boss and an employee. - Multilingual translation was strong for high-resource languages Mandarin, Arabic, Japanese, Thai but broke down on low-resource languages like Gujarati, suggesting multilinguality outside major languages is not yet a strength. - Image-grounded generation recreating a rotisserie scene as HTML/CSS from a photo produced a convincing result overall, though fine details like fire intensity didn’t fully match the reference image. Everyone else built a construction worker. We built the contractor. One file at a time. UI, API, database, deploy. How big is Mistral Large 4 and how is it deployed? Mistral Large 4 follows a mixture-of-experts design: around a trillion parameters in total, with only about 50 billion active for any given token, which keeps inference costs down relative to a dense model of similar total size. It was trained in Mistral’s own European infrastructure, a detail the company has leaned on before as a differentiator from US and Chinese labs. The release is staged. The model is live now in public preview through Mistral’s API and in Le Chat, Mistral’s consumer chat product. Open weights, which would let developers self-host and fine-tune the model, are expected by the end of the month. That lag between API availability and weight release is standard practice for labs that want early feedback before committing to a permanent open download. How does it perform on coding and security benchmarks? Testing referenced several benchmark categories that matter for developers: On a software engineering benchmark resembling SWE-bench, which checks whether a model can fix real-world coding issues, Mistral Large 4 beat most open models in the comparison, trailing only Kimi K3. On a terminal-based benchmark that measures how well a model handles command-line tasks, it performed well but was clearly behind GLM 5.3, and Chinese open models generally led this category. On a cybersecurity index measuring a model’s ability to find and fix security flaws in real software, Mistral Large 4 tied for the top spot with GLM 5.3 Flash. On a business workflow benchmark covering tasks across email, spreadsheets, and Slack-style apps, it landed just behind GLM 5.3 among the strongest open models. On a finance-agent benchmark involving research over company filings and financial data, it sat near the top, slightly behind GLM but slightly ahead of GPT-6 Astra in the same chart. The pattern across these charts: Mistral Large 4 is competitive at the top tier of open models, strong specifically in coding and security work, but not uniformly ahead of rivals like GLM 5.3, which led in terminal tasks and workflow automation. Can it actually find security vulnerabilities in code? This was tested directly rather than just taken from a leaderboard. A Flask-based client portal app was set up locally with a SQLite database, multiple users, and invoices tied to each account. The app had a deliberately planted bug: logging in as one user and requesting another user’s invoice by ID returned that invoice’s full data, a classic insecure direct object reference IDOR and broken access control flaw. Remy doesn't write the code. It manages the agents who do. Remy runs the project. The specialists do the work. You work with the PM, not the implementers. Pointed at the running app through an agent setup, Mistral Large 4 identified the IDOR quickly and correctly explained the root cause: the delete endpoint checked object ownership before acting, but the get endpoint didn’t perform the same check, which is exactly the inconsistency the bug was built around. It didn’t stop there. The model also surfaced plaintext password storage, a hardcoded session secret key, and a committed cookie file containing a live session token, all issues that weren’t the primary target of the test but were real problems in the sample app. It then proposed and could apply a fix. This kind of result lines up with the cybersecurity benchmark score and suggests the model is genuinely useful for automated code review and security triage, not just pattern-matching on benchmark-style prompts. How well does it reason over images? Two image-based tests stand out. The first was a classic physics-style puzzle: an image showing four people touching hot and neutral wires under different conditions standing on insulating stools versus bare ground , asking the model to determine who gets shocked. Mistral Large 4 correctly identified both people at risk, including a case where one person touched both hot and neutral wires simultaneously, a detail designed to trip up the surface-level reasoning: that person isn’t saved by the usual insulation logic. It correctly reasoned that the person on an insulating stool was safe and that another person was at the same electrical potential as the ground. The second test was a social reasoning task: a WhatsApp screenshot where message timestamps didn’t line up cleanly with the conversation content, creating a misunderstanding between a boss and an employee. The model correctly interpreted workplace slang, identified the triggering message, and reconstructed how a wife reading over someone’s shoulder misread a work message as romantic, building a clear chain of inference from ambiguous text to the boss’s reaction. This is a harder kind of test than typical visual question answering because it requires inferring intent and context, not just transcribing text. A third test involved recreating a photographed rotisserie scene meat, drip tray, motor, fire as a single self-contained HTML file based only on the image and a text prompt. The output captured the layered meat, drip tray, and mechanical details convincingly, though the fire effect was noticeably weaker and less dynamic than in the source photo. Is Mistral Large 4 good at multiple languages? Here results were mixed. Given a Tamil newspaper front page and asked to translate the headline into roughly 75 languages and scripts, Mistral Large 4 handled high-resource languages well, producing clean and correctly scripted translations for Mandarin, Arabic, Japanese, and Thai. Performance dropped noticeably for low-resource languages such as Gujarati, where translation quality degraded. This suggests Mistral Large 4’s multilingual strength tracks closely with training data availability: it’s reliable for the world’s most widely used languages but not yet a strong choice for low-resource or less commonly digitized languages. Teams building for global, long-tail language coverage should verify specific language pairs before relying on it rather than assuming uniform quality across the roughly 75 languages tested. Is Mistral Large 4 worth using right now? Plans first. Then code. Remy writes the spec, manages the build, and ships the app. For coding and security-focused workloads, it looks like one of the stronger open-weight options available, competitive with or ahead of rivals on SWE-bench-style and cybersecurity benchmarks, and it backed that up in a live vulnerability-detection test rather than just a static score. For visual and even social reasoning from images, it performed well above a cursory glance would suggest. The weak points are specific: terminal-heavy agentic tasks and cross-app business workflows where GLM 5.3 currently leads, and translation into low-resource languages. Since it’s still a public preview with open weights not yet released, teams planning to self-host should wait for the weight drop before committing infrastructure decisions, while those using the API or Le Chat can evaluate it today. Frequently Asked Questions What does “Le Chonk” mean in the Mistral Large 4 release? It’s Mistral’s informal nickname for the model, playing on internet slang for something large and chunky, often used to describe big cats. Mistral leaned into the joke with a pixel-art cat on the model’s page, while the official name remains Mistral Large 4. How many parameters does Mistral Large 4 have? It is a mixture-of-experts model with roughly a trillion parameters in total, but only about 50 billion are active for any given inference pass, which keeps compute costs down relative to a dense model of the same total size. Is Mistral Large 4 open source? Open weights were promised for release by the end of the month following the preview announcement. At launch, the model was accessible via Mistral’s API and through Le Chat, but the downloadable weights for self-hosting weren’t yet public. How does Mistral Large 4 compare to GLM 5.3? It’s close but not uniformly ahead. Mistral Large 4 led or tied on coding and cybersecurity benchmarks, while GLM 5.3 was ahead on terminal-based tasks and cross-app business workflow automation. The two sit near the top of the open-model field on different strengths. Can Mistral Large 4 find security bugs in real code? Yes, in testing it correctly identified a planted insecure direct object reference IDOR vulnerability in a Flask application, explained the exact logic flaw causing it, and additionally found unrelated issues like plaintext password storage and a hardcoded session secret.