🤖 AI Agents Weekly: Claude Haiku 5.5, Gemini Agent, Mistral Large 4, Decisions API, Personal Agent Protocol, Devin Dreaming, and More Anthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model to date, priced at $0.10 input and $0.50 output per million tokens for prompts up to 100K tokens — about 75% less to run on average than Haiku 4.5. The model scores 72.4% on the OSWorld 2.1 offline subset versus 15.7% for Haiku 4.5, 39.2% on Terminal-Bench 4.0 versus 0.0%, and a GDPval-AA v2.1 Elo of 1620 versus 735, and is the first Haiku model with an adjustable effort setting. Anthropic also cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens, cutting its cost on most agentic work by around 20%, and added computer use and browser use in beta to its Python and TypeScript SDKs. In today’s issue: - Anthropic ships Claude Haiku 5.5 - Google launches the Gemini agent - Mistral previews Large 4 - OpenAI opens the Decisions API - Meta and Sierra propose Personal Agent Protocol - Devin gets Memory and Dreaming - Reflection unveils Beam open model - GPT-6 brings Intelligent UI to ChatGPT - Hugging Face trains models across harnesses - OpenAI releases 719 math manuscripts And all the top AI dev news, papers, and tools. Top Stories Claude Haiku 5.5 Anthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model to date. It targets high-volume work such as summaries, compaction, classification, and subagent tasks under Opus 5.5 and Sonnet 5.5, and costs around 75% less to run than Haiku 4.5 on average. - Agentic performance: 72.4% on the OSWorld 2.1 offline subset against 15.7% for Haiku 4.5, 39.2% on Terminal-Bench 4.0 against 0.0%, and a GDPval-AA v2.1 Elo of 1620 against 735. - Pricing: $0.10 input, $0.50 output, and $0.01 cache reads per million tokens for prompts up to 100K tokens, which covered about 90% of Haiku 4.5 requests. Longer prompts cost $0.50 input and $2.50 output. - Effort control: It is the first Haiku model with an adjustable effort setting, so developers can trade cost for intelligence per task. - Wider price changes: Sonnet 5.5 cache reads drop from $0.20 to $0.10 per million tokens, cutting its cost on most agentic work by around 20%. Max and Team subscribers get monthly API credits of $100 to $500, and the Python and TypeScript SDKs add computer use and browser use in beta. - Availability: Live as claude-haiku-5-5 on all platforms, including AWS, Google Cloud, and Azure.