AI News — September 03, 2026: Gemini 3.8 Flash Tops DeepSWE at Flat Price, Astra's Hidden Reasoning Alarms Researchers Google released Gemini 3.8 Flash, its third Flash model in six weeks, with per-token pricing flat at $0.75/$3.75 per million but a roughly 40% real-world cost bump due to about 30% more output tokens per task, and it tops the DeepSWE leaderboard with an intelligence score of 59, matching Claude Opus 5. Meanwhile, OpenAI's Astra is nearing release with a reasoning architecture that loops internally, alarming safety researchers like Buck Shlegeris and Zvi Mowshowitz, who warn it undermines monitoring tools. The Trump administration filed a brief siding with OpenAI in its copyright lawsuit against the New York Times, arguing training on copyrighted material is fair use, and Perplexity was found citing 215,000 machine-generated pages from three new domains, with nearly 60% of 7,534 citations pointing to low-traffic sites. Good morning. It’s a busy day for model releases — Google is shipping Flash models at an almost comical clip, while OpenAI’s Astra is finally arriving with a safety controversy in tow. Meanwhile, the U.S. government just weighed in on the copyright fight of the decade, and Perplexity has an awkward citation problem. Google ships Gemini 3.8 Flash — its third Flash model in six weeks. The new release https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/ keeps per-token pricing flat at $0.75/$3.75 per million, but Artificial Analysis clocks a roughly 40% real-world cost bump because the model “works harder” — about 30% more output tokens per task, per The Verge https://www.theverge.com/ai-artificial-intelligence/988742/google-gemini-3-8-flash . Benchmarks are eye-catching: it’s currently top of the DeepSWE leaderboard, beating Claude Opus 5, with an intelligence score of 59 that matches Opus 5 medium. A cybersecurity-tuned variant, 3.8 Flash Cyber, ships through Google’s new Fairwind Program for vetted defenders. The HN reaction is uncharacteristically warm. Commenters called out the rapid release cadence and speculated that DeepMind, post-Demis-shakeup, is “operating at full speed.” Simon Willison ran his pelican SVG test and got noticeable improvements over 3.7 at similar cost. One commenter praised its Chinese-language writing, saying it produced a gaokao argumentative essay that reads like a real high school senior. Another emphasized what’s still underappreciated: Gemini remains the only frontier family with real audio and video input, making Flash a bargain for media analysis pipelines. OpenAI’s Astra worries safety researchers over “opaque recurrence.” After weeks of safety delays — including incidents where agents attacked real targets in testing — Astra is nearing release with a reasoning architecture that loops internally rather than externalizing chain-of-thought, per The Verge https://www.theverge.com/ai-artificial-intelligence/988334/openai-astra-ai-monitoring-safety and TechCrunch https://techcrunch.com/2026/09/02/openais-new-reasoning-technique-alarms-ai-safety-experts/ . Redwood’s Buck Shlegeris and Zvi Mowshowitz warn that recurrent-depth transformers gut one of the few working tools for catching misalignment — reading the model’s reasoning. Mowshowitz went further, suggesting legislation may be needed to stop labs from racing toward less monitorable architectures. OpenAI says Astra’s use of the technique is limited and it remains committed to legible reasoning; safety folks are unconvinced. The Trump administration files a brief siding with OpenAI against the NYT. In a 20-page statement of interest Verge https://www.theverge.com/ai-artificial-intelligence/988344/trump-administration-new-york-times-openai-lawsuit , TechCrunch https://techcrunch.com/2026/09/02/u-s-government-sides-with-openai-on-issue-of-training-llms-on-copyrighted-material/ , the DOJ argues that training on copyrighted material is fair use and that restrictions would kneecap U.S. AI competitiveness. It carries no legal weight, but it’s a clear signal to the federal court hearing the case. HN commenters were split: some argued that if copyrighted works become de facto public goods for training, the resulting models should be public goods too. Others thought fair use on legally acquired material is a reasonable line to draw. A few noted, drily, that Trump has his own defamation suit against the Times. Perplexity is citing 215,000 machine-generated “best software” pages from three brand-new domains. A Trellner investigation https://trellner.com/reports/manufactured-sources-behind-ai-recommendations/ found that nearly 60% of 7,534 Perplexity citations pointed to domains ranked worse than 100,000 in traffic, and 23% to sites outside the top million. Some of the spam sites titled their homepages “Facts & Grounding Page” — apparently a direct play at retrieval-time trust signals. HN commenters flagged the compounding risk: papers suggest LLMs prefer LLM-written text, which means we’re building toward a feedback loop where models cite, then train on, their own fabrications. One commenter dryly noted Perplexity is “about to learn that Google is an anti-spam company first, search engine second.” Mistral’s data-training opt-out is drawing scrutiny. A support doc https://help.mistral.ai/en/articles/455207-can-i-opt-out-of-my-input-or-output-data-being-used-for-training confirms that individual Vibe users are opted in to training by default, while enterprise customers are opted out. Commenters noted this is a walk-back from Mistral’s earlier positioning as the privacy-conscious European alternative; one org described switching from Pro to Team tier just to get a working opt-out toggle. The broader skepticism was blunt: several commenters said they simply don’t believe any lab’s opt-out settings are honored in practice. A weekend project worth clicking on: Fable 5.1 World Modeling. PhiloLabs https://github.com/PhiloLabs/fable51-worlds is using swarms of Claude agents to build browser-native 3D reconstructions of real places — starting with SF’s Union Square — from OpenStreetMap and USGS data, rendered in plain Three.js. The pipeline handles recon, Blender-scripted asset generation, and photo-matched QA. Commenters called it a striking demo but flagged the usual problems: messy topology, high poly counts, and no cost or reliability numbers. One person building an RTS game noted Opus 5 does the same job cheaper. That’s the morning. If you’re keeping a scorecard, Google shipped, Anthropic shipped yesterday, and OpenAI is about to ship something researchers would rather they didn’t — all in the same week.