{"slug": "ai-news-report-october-6-mistral-large-4-solves-93-of-hacking-tests-that-claude", "title": "AI News Report, October 6: MISTRAL LARGE 4 SOLVES 93% OF HACKING TESTS THAT CLAUDE AND GPT-6 REFUSE", "summary": "Mistral released a public preview of Mistral Large 4, an open-weight model with about 1 trillion parameters that Mistral says solves 93% of Cybench, a set of 40 hacking-competition challenges, while Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse. The weights are due by the end of October, letting security teams and attackers run the model on their own servers. The report also covers Polars 2.0 shipping October 6, Kolibri-1's 78 billion parameters with 3.46 billion active, and CVE-2026-90970, which requires GitLab self-hosted upgrades to 19.2.4, 19.3.2 or 19.4.1.", "body_md": "The free, MIT-licensed engine runs Qwen3.8-Flash-Next and wrote 94 tokens a second on an RTX 5070 in its own tests. You need 32 GB of RAM and about 80 GB of disk, and nothing leaves your PC.\n\nPolars 2.0 shipped October 6 with spill-to-disk near 80% of RAM and a native Map type. Before you upgrade, joins and group-bys no longer keep row order unless you set maintain_order=True.\n\nKolibri-1 has 78 billion parameters with 3.46 billion active, an Apache 2.0 license and up to a 1 million token context. Its FP8 weights take about 78 GB, so one H200 or B200 can serve it.\n\nSkills reached Rapid Release Workspace domains on October 5, and you can stack several in one prompt. Business accounts lose Gems on March 1, 2027, so list the Gems your team relies on now.\n\nChakra ran more than 500,000 interviews in a six-month beta and folds a recruiter screen, a take-home and an engineer interview into one session. HackerRank says humans still make the hire.\n\nAhmad Al-Dahle says Airbnb ships nearly 80% more features a year and AI resolves about half of support tickets. With an internal tool, airport pickups took about six weeks to build, against eight or nine months for groceries.\n\nThe Techdirt writer lists Muse's problems, from a Mac zero-day to reading private messages and profiling your contacts. Read it before you let any agent pay bills for you.\n\nHe quotes Anthropic's Felix Rieseberg: the new Cowork runs the model and the VM in the cloud, with a sandbox per session. Work keeps going after you close the laptop.\n\nConor McCarthy traces the page every developer has pasted into a test, which now cycles through six languages every 5 seconds. A light read on the web's most boring page.\n\nOn 24 September 2026, a malicious cyber actor (MCA) used 149.104.78.141 to attempt zero-day exploitation against a Citrix NetScaler Gateway. At the… · GreyNoise\n\nCVE-2026-90970 lets a user with Duo Agent Platform access escape the prompt-template sandbox on a self-hosted AI Gateway. GitLab.com is already fixed; self-hosted shops need 19.2.4, 19.3.2 or 19.4.1.\n\nHorizon3 used Mythos to find CVE-2026-61500, which lets anyone forge an admin login on Rejetto HFS 3.0.0 to 3.2.0. VulnCheck honeypots now see probing, so upgrade to 3.2.1 or later.\n\nGoogle says most of the automated submissions were not valid, so the program is paused from October 1, with an update due in early 2027. Its other bug bounty programs stay open.\n\nThe Information reports Meta's Claude Code users fell from about 60,000 to 30,000, and Microsoft cut a $1 billion internal Claude estimate by more than a third. Both push staff to in-house tools.\n\nOpenAI will test labeled visual ads during image generation in the US later this month, on a service it says reaches 1.2 billion people a week. OpenAI says ads do not change ChatGPT's answers.\n\nNew Microsoft 365 Copilot Business licenses sold through CSP get usage-based billing on by default, capped at 4,000 Copilot Credits per user per month. MSPs can lower that cap before clients see a bill.\n\nReuters says Tencent and CATL lead commitments above 80 billion yuan, well past a 50 billion yuan target. DeepSeek is preparing for a possible STAR Market listing.\n\nSB 1246 fines operators up to $10,000 per car that blocks an emergency for more than 30 minutes after a technician is called. It also limits remote driving to licensed drivers in the US, from July 2028.\n\nThe BriefMistral released a public preview of Mistral Large 4, an open-weight model with about 1 trillion parameters, and says it solves 93% of Cybench, a set of 40 hacking-competition challenges, while Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse. The weights are due by the end of October, so security teams will be able to run it on their own servers, and so will attackers. If you defend networks, try it on defensive work now and plan as if every attacker will soon have the same tool.\n\nLevel UpPick one defensive chore your current AI model refuses or hedges on, such as confirming a known flaw in a test copy of your own software or writing a detection rule for it. Run the same task through the Mistral Large 4 preview, which Mistral's docs list at 68 cents in and $2.09 out per million tokens, marked down from $1.36 and $4.18. Write down the result and the cost before the weights arrive. →Mistral docs: Mistral Large 4 model page and pricing", "url": "https://wpnews.pro/news/ai-news-report-october-6-mistral-large-4-solves-93-of-hacking-tests-that-claude", "canonical_source": "https://theainewsreport.com/2026-10-06.html", "published_at": "2026-10-06 14:07:49+00:00", "updated_at": "2026-10-06 14:20:20.925617+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-products", "ai-policy"], "entities": ["Mistral", "Mistral Large 4", "Cybench", "Claude Opus 5.5", "GPT-6 Astra", "Polars 2.0", "Kolibri-1", "GitLab"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-news-report-october-6-mistral-large-4-solves-93-of-hacking-tests-that-claude", "markdown": "https://wpnews.pro/news/ai-news-report-october-6-mistral-large-4-solves-93-of-hacking-tests-that-claude.md", "text": "https://wpnews.pro/news/ai-news-report-october-6-mistral-large-4-solves-93-of-hacking-tests-that-claude.txt", "jsonld": "https://wpnews.pro/news/ai-news-report-october-6-mistral-large-4-solves-93-of-hacking-tests-that-claude.jsonld"}}