# AI News — September 11, 2026: Alibaba's 151M-Exchange Distillation Named, Second Mathematician Accuses OpenAI

> Source: <https://ai0.news/posts/2026-09-11-daily-digest/>
> Published: 2026-09-11 06:00:08+00:00

Good morning. Anthropic is naming names — Alibaba, Moonshot, and DeepSeek — in a report alleging massive distillation attacks against Claude, while the OpenAI-and-mathematicians saga picks up a third act. Cognition tried to launch a coding model into that news cycle and got a skeptical reception, and DeepSeek quietly shipped a V4.1 Flash that the HN crowd is already calling the best release of the week.

**Anthropic accuses three Chinese labs of large-scale Claude distillation.** In a new report covered by [TechCrunch](https://techcrunch.com/2026/09/10/anthropic-details-distillation-campaigns-from-alibaba-moonshot-ai-and-deepseek/), Anthropic says roughly 200 million exchanges were used to extract Claude’s reasoning traces for training competing models. The largest campaign, attributed to Alibaba, ran 151 million exchanges across 3,500 accounts between May and July, apparently feeding Qwen training. A separate Moonshot campaign allegedly routed some queries through Chinese military infrastructure, including surveillance-related prompts — a claim that, if it holds up, will complicate any argument that distillation is just a commercial nuisance.

**The OpenAI math controversy widens to a second mathematician.** Andreas Thom has now [publicly accused OpenAI](https://www.theverge.com/ai-artificial-intelligence/993263/where-does-openai-get-mathematics-training-data) of using his unpublished work on non-sofic groups, becoming the second researcher in a week to raise the alarm after Tristan Buckmaster. Thom notes OpenAI’s original announcement failed to credit him and Gábor Kun before quietly amending it under pressure, and his own [Mastodon thread](https://mathstodon.xyz/@andreasthom/117240535270608201) lays out why the resemblance to his approach looks suspicious. On HN, one commenter flagged the pattern that’s hardest to explain away: OpenAI generating 300 billion output tokens from a still-training model right after learning a major proof may have been in its data. Others pushed back that companies genuinely can’t audit every user’s data-sharing settings on demand.

**Cognition’s SWE-2 lands to a skeptical crowd.** Cognition [released SWE-2](https://cognition.com/blog/swe-2), post-trained from Kimi K3, claiming 50.0 on FrontierCode 1.1 and 92.8 on Terminal-Bench 2.1 at 64% lower cost than Fable 5.1. The problem, as [one HN thread](https://news.ycombinator.com/item?id=49645443) quickly noted, is that Terminal-Bench 2.1 is essentially saturated, and SWE-2’s score on the newer Terminal-Bench 4.0 is 27.3 — well behind competitors near 56-58. Combine that with Cognition’s history of misleading Devin demos and a lot of “just tried it, running in circles” comments, and the launch is landing as benchmaxing rather than a genuine leap. A [separate writeup on Tokenstead](https://tokenstead.ai/models/swe-2) is getting the same treatment.

**DeepSeek V4.1 Flash arrives at 552B parameters.** DeepSeek [announced V4.1 Flash](https://twitter.com/deepseek_ai/status/2097930608790167907), the smallest model in a new architecture family with native vision, and it’s already on HuggingFace. It’s roughly double the parameter count of the previous Flash, which explains the benchmark jumps but rules out most local deployments. The community reaction is unusually warm — Simon Willison ran it through his pelican SVG test across all reasoning levels, and multiple commenters singled out the $0.003-per-million cache hit price as potentially cheaper than the network cost of sending the tokens in the first place. The tech report’s willingness to actually explain the model, rather than pad with safety boilerplate, keeps getting called out favorably.

**OpenAI ships an Agents API and a Lean 4 proof.** OpenAI’s new [Agents API](https://developers.openai.com/api/docs/guides/agents-api/overview) offers managed sessions, orchestration, context compaction, and sandboxed code execution — essentially hosting the Codex harness so you don’t have to. Reaction is split between “finally, the harness problem is real” and “this is vendor lock-in wearing a bowtie,” with the self-hosted sandbox option softening the second complaint. Separately, John D. Cook [flagged](https://www.johndcook.com/blog/2026/09/09/formal-method-revolution/) that OpenAI’s Navier-Stokes release included a Lean 4 formal verification completed in 17 hours versus an estimated 132,800 person-hours by hand. HN commenters pointed out the 40-hours-per-page baseline is from 2005 and Lean’s automation has improved enormously since, so the “four orders of magnitude” framing is generous — but the underlying point about formal verification becoming practical is real.

**OpenAI also announced ChatGPT for Financial Services** ([link](https://openai.com/index/introducing-chatgpt-financial-services)) and a piece on [using Codex to hunt antimicrobial molecules](https://openai.com/index/using-codex-chatgpt-to-search-for-new-antimicrobials), though neither post had much substance beyond the headline when we checked.

That’s a lot of frontier lab drama for a Thursday. If Anthropic’s distillation report is accurate at the scale described, expect the “who trained on whom” conversation to eat the next news cycle whole.
