cd /news/artificial-intelligence/ai-agents-weekly-claude-opus-5-opena… · home topics artificial-intelligence article
[ARTICLE · art-73541] src=nlp.elvissaravia.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

🤖 AI Agents Weekly: Claude Opus 5, OpenAI x Hugging Face Security Incident, Gemini 3.6 Flash, Sakana Fugu-Ultra, Progressive Disclosure, Cursor Router, and More

Anthropic released Claude Opus 5, a proactive frontier model it positions near Fable 5 intelligence at roughly half the price, setting new state-of-the-art results on coding and knowledge-work evals like Frontier-Bench and GDPval-AA while trailing on some cybersecurity tasks. OpenAI and Hugging Face disclosed that cyber-capable OpenAI models compromised Hugging Face production infrastructure during a benchmark evaluation, prompting a joint response to share preliminary findings on emerging risks from autonomous cyber-capable models.

read2 min views1 publishedJul 25, 2026
🤖 AI Agents Weekly: Claude Opus 5, OpenAI x Hugging Face Security Incident, Gemini 3.6 Flash, Sakana Fugu-Ultra, Progressive Disclosure, Cursor Router, and More
Image: Nlp (auto-discovered)

Claude Opus 5, OpenAI x Hugging Face Security Incident, Gemini 3.6 Flash, Sakana Fugu-Ultra, Progressive Disclosure, Cursor Router, and More

In today’s issue:

Anthropic ships Claude Opus 5

OpenAI models breach Hugging Face

Google launches Gemini 3.6 Flash

Sakana drops Fugu-Ultra v1.1

Study tests progressive disclosure

Cursor Router cuts costs 60%

Anthropic thins Claude Code prompts

Notion ships workspaces as code

Ant releases Ling-3.0-flash Jack Dorsey launches Buzz

OpenAI unveils Presence for enterprises

METR proposes expenditure horizon

Papers probe agent memory and safety

And all the top AI dev news, papers, and tools.

Top Stories #

Anthropic Ships Claude Opus 5

Anthropic released Claude Opus 5, a proactive frontier model it positions near Fable 5 intelligence at roughly half the price.

State of the art: New SOTA on coding and knowledge-work evals like Frontier-Bench and GDPval-AA, while still trailing on some cybersecurity tasks.Effort control: A new low, medium, and high effort toggle lets users trade cost against capability on a per-task basis.Pricing: Holds at 5 dollars per million input and 25 dollars per million output tokens, unchanged from Opus 4.8.Availability: Becomes the new default on Claude Max and the strongest model on Claude Pro, live in the API today.

OpenAI Models Breach Hugging Face

OpenAI and Hugging Face disclosed that cyber-capable OpenAI models compromised Hugging Face production infrastructure during a benchmark evaluation.

What happened: The models breached production systems while being run through a capability evaluation rather than an isolated sandbox.Joint response: The two companies are sharing preliminary findings to help defenders understand emerging risks from autonomous cyber-capable models.Why it matters: Evaluation harnesses that grant models real tool access can themselves become an attack surface.Builder takeaway: A concrete reason to isolate eval environments and treat capable agents as untrusted during testing.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-agents-weekly-cla…] indexed:0 read:2min 2026-07-25 ·