AI Agents Weekly: GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, NVIDIA Buys Hugging Face, Grok Bot Design, FrontierHarness Eval, and More OpenAI released GPT-6 Astra, a model designed for computer operation, topping all seven rows of OpenAI's launch comparison with 41.4% on AutomationBench versus 18.1% for GPT-5.6 Sol and 31.4% for Claude Fable 5.1, and 57.9% on Terminal-Bench 4.0 versus 37.3% and 55.8%. The rollout begins with select organizations and expands to ChatGPT Plus, Pro, Business, Enterprise, the OpenAI API, and AWS. In today’s issue: - OpenAI ships GPT-6 Astra - Anthropic releases Claude Fable 5.1 and Mythos 5.1 - Google ships Gemini 3.8 Flash and a cyber variant - NVIDIA agrees to acquire Hugging Face - xAI publishes the design thinking behind Grok Bot - Meta releases Muse Spark 1.3 - FrontierHarness Eval compares nine harnesses on one model - Claude uses your computer in the background - Claude Code previews Function Hooks - Anthropic open-sources Claude Commerce Agents - Cursor runs cloud agents on your own machines - Cline migrates its extension onto a new SDK harness - FrontierSWE v2 extends coding agent runs to 20 hours - Cheating and whistleblowing emerge in a 100-agent research swarm And all the top AI dev news, papers, and tools. Top Stories OpenAI Ships GPT-6 Astra OpenAI released GPT-6 Astra, a model built for operating a computer rather than answering in a chat window. It tops all seven rows of OpenAI’s launch comparison, with the widest margins on the agentic and computer-workflow benchmarks. - Agentic workloads: In OpenAI’s launch comparison, Astra takes 41.4% on AutomationBench against 18.1% for GPT-5.6 Sol and 31.4% for Claude Fable 5.1, and 57.9% on Terminal-Bench 4.0 against 37.3% and 55.8%. - Science and math: 64.6% on Terminal-Bench Science 0.1 against 22.4% and 52.6%, and 97.6% on FrontierMath Tier 4 v2 against 83.0% and 87.8%. - ARC-AGI-3: 99.9% against 7.8% for GPT-5.6 Sol, with no reported Fable 5.1 score. - Computer workflows: OpenAI also claims state-of-the-art on Agents’ Last Exam and ScreenSpot Pro, its other computer workflow benchmarks. - Rollout: Limited to a set of organizations on day one, then rolling out to ChatGPT Plus, Pro, Business, and Enterprise, plus the OpenAI API and AWS. The Hacker News thread reached 2,082 points and 1,897 comments in a day.