# 🤖 AI Agents Weekly: Claude Opus 5, OpenAI x Hugging Face Security Incident, Gemini 3.6 Flash, Sakana Fugu-Ultra, Progressive Disclosure, Cursor Router, and More

> Source: <https://nlp.elvissaravia.com/p/ai-agents-weekly-claude-opus-5-openai>
> Published: 2026-07-25 17:20:58+00:00

# 🤖 AI Agents Weekly: Claude Opus 5, OpenAI x Hugging Face Security Incident, Gemini 3.6 Flash, Sakana Fugu-Ultra, Progressive Disclosure, Cursor Router, and More

### Claude Opus 5, OpenAI x Hugging Face Security Incident, Gemini 3.6 Flash, Sakana Fugu-Ultra, Progressive Disclosure, Cursor Router, and More

In today’s issue:

Anthropic ships Claude Opus 5

OpenAI models breach Hugging Face

Google launches Gemini 3.6 Flash

Sakana drops Fugu-Ultra v1.1

Study tests progressive disclosure

Cursor Router cuts costs 60%

Anthropic thins Claude Code prompts

Notion ships workspaces as code

Ant releases Ling-3.0-flash

Jack Dorsey launches Buzz

OpenAI unveils Presence for enterprises

METR proposes expenditure horizon

Papers probe agent memory and safety

And all the top AI dev news, papers, and tools.

## Top Stories

### Anthropic Ships Claude Opus 5

Anthropic released Claude Opus 5, a proactive frontier model it positions near Fable 5 intelligence at roughly half the price.

**State of the art:** New SOTA on coding and knowledge-work evals like Frontier-Bench and GDPval-AA, while still trailing on some cybersecurity tasks.**Effort control:** A new low, medium, and high effort toggle lets users trade cost against capability on a per-task basis.**Pricing:** Holds at 5 dollars per million input and 25 dollars per million output tokens, unchanged from Opus 4.8.**Availability:** Becomes the new default on Claude Max and the strongest model on Claude Pro, live in the API today.

### OpenAI Models Breach Hugging Face

OpenAI and Hugging Face disclosed that cyber-capable OpenAI models compromised Hugging Face production infrastructure during a benchmark evaluation.

**What happened:** The models breached production systems while being run through a capability evaluation rather than an isolated sandbox.**Joint response:** The two companies are sharing preliminary findings to help defenders understand emerging risks from autonomous cyber-capable models.**Why it matters:** Evaluation harnesses that grant models real tool access can themselves become an attack surface.**Builder takeaway:** A concrete reason to isolate eval environments and treat capable agents as untrusted during testing.
