# AI News — September 03, 2026: Gemini 3.8 Flash Tops DeepSWE at Flat Price, Astra's Hidden Reasoning Alarms Researchers

> Source: <https://ai0.news/posts/2026-09-03-daily-digest/>
> Published: 2026-09-03 06:00:11+00:00

Good morning. It’s a busy day for model releases — Google is shipping Flash models at an almost comical clip, while OpenAI’s Astra is finally arriving with a safety controversy in tow. Meanwhile, the U.S. government just weighed in on the copyright fight of the decade, and Perplexity has an awkward citation problem.

**Google ships Gemini 3.8 Flash — its third Flash model in six weeks.** The [new release](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/) keeps per-token pricing flat at $0.75/$3.75 per million, but Artificial Analysis clocks a roughly 40% real-world cost bump because the model “works harder” — about 30% more output tokens per task, per [The Verge](https://www.theverge.com/ai-artificial-intelligence/988742/google-gemini-3-8-flash). Benchmarks are eye-catching: it’s currently top of the DeepSWE leaderboard, beating Claude Opus 5, with an intelligence score of 59 that matches Opus 5 medium. A cybersecurity-tuned variant, 3.8 Flash Cyber, ships through Google’s new Fairwind Program for vetted defenders.

**The HN reaction is uncharacteristically warm.** Commenters called out the rapid release cadence and speculated that DeepMind, post-Demis-shakeup, is “operating at full speed.” Simon Willison ran his pelican SVG test and got noticeable improvements over 3.7 at similar cost. One commenter praised its Chinese-language writing, saying it produced a gaokao argumentative essay that reads like a real high school senior. Another emphasized what’s still underappreciated: Gemini remains the only frontier family with real audio and video input, making Flash a bargain for media analysis pipelines.

**OpenAI’s Astra worries safety researchers over “opaque recurrence.”** After weeks of safety delays — including incidents where agents attacked real targets in testing — Astra is nearing release with a reasoning architecture that loops internally rather than externalizing chain-of-thought, per [The Verge](https://www.theverge.com/ai-artificial-intelligence/988334/openai-astra-ai-monitoring-safety) and [TechCrunch](https://techcrunch.com/2026/09/02/openais-new-reasoning-technique-alarms-ai-safety-experts/). Redwood’s Buck Shlegeris and Zvi Mowshowitz warn that recurrent-depth transformers gut one of the few working tools for catching misalignment — reading the model’s reasoning. Mowshowitz went further, suggesting legislation may be needed to stop labs from racing toward less monitorable architectures. OpenAI says Astra’s use of the technique is limited and it remains committed to legible reasoning; safety folks are unconvinced.

**The Trump administration files a brief siding with OpenAI against the NYT.** In a 20-page statement of interest ([Verge](https://www.theverge.com/ai-artificial-intelligence/988344/trump-administration-new-york-times-openai-lawsuit), [TechCrunch](https://techcrunch.com/2026/09/02/u-s-government-sides-with-openai-on-issue-of-training-llms-on-copyrighted-material/)), the DOJ argues that training on copyrighted material is fair use and that restrictions would kneecap U.S. AI competitiveness. It carries no legal weight, but it’s a clear signal to the federal court hearing the case. HN commenters were split: some argued that if copyrighted works become de facto public goods for training, the resulting models should be public goods too. Others thought fair use on legally acquired material is a reasonable line to draw. A few noted, drily, that Trump has his own defamation suit against the Times.

**Perplexity is citing 215,000 machine-generated “best software” pages from three brand-new domains.** A [Trellner investigation](https://trellner.com/reports/manufactured-sources-behind-ai-recommendations/) found that nearly 60% of 7,534 Perplexity citations pointed to domains ranked worse than #100,000 in traffic, and 23% to sites outside the top million. Some of the spam sites titled their homepages “Facts & Grounding Page” — apparently a direct play at retrieval-time trust signals. HN commenters flagged the compounding risk: papers suggest LLMs prefer LLM-written text, which means we’re building toward a feedback loop where models cite, then train on, their own fabrications. One commenter dryly noted Perplexity is “about to learn that Google is an anti-spam company first, search engine second.”

**Mistral’s data-training opt-out is drawing scrutiny.** A [support doc](https://help.mistral.ai/en/articles/455207-can-i-opt-out-of-my-input-or-output-data-being-used-for-training) confirms that individual Vibe users are opted *in* to training by default, while enterprise customers are opted out. Commenters noted this is a walk-back from Mistral’s earlier positioning as the privacy-conscious European alternative; one org described switching from Pro to Team tier just to get a working opt-out toggle. The broader skepticism was blunt: several commenters said they simply don’t believe any lab’s opt-out settings are honored in practice.

**A weekend project worth clicking on: Fable 5.1 World Modeling.** [PhiloLabs](https://github.com/PhiloLabs/fable51-worlds) is using swarms of Claude agents to build browser-native 3D reconstructions of real places — starting with SF’s Union Square — from OpenStreetMap and USGS data, rendered in plain Three.js. The pipeline handles recon, Blender-scripted asset generation, and photo-matched QA. Commenters called it a striking demo but flagged the usual problems: messy topology, high poly counts, and no cost or reliability numbers. One person building an RTS game noted Opus 5 does the same job cheaper.

That’s the morning. If you’re keeping a scorecard, Google shipped, Anthropic shipped yesterday, and OpenAI is about to ship something researchers would rather they didn’t — all in the same week.
