Agent Skills: The Composition Cliff
Adding a single skill to an LLM agent reliably improves task performance, but by the time a fourth skill is active, the cumulative gain shrinks to a fraction of what two skills produced, according to …
Adding a single skill to an LLM agent reliably improves task performance, but by the time a fourth skill is active, the cumulative gain shrinks to a fraction of what two skills produced, according to …
Semgrep launched nine Agentic Workflows in public beta that automate deep vulnerability hunting across codebases, covering 70+ CWEs including authentication, injection, and business logic flaws. The w…
Kaj Sotala published a personal policy on AI use for essay writing, stating that they use AI as an extensive aid for thinking but retain primary authorship, with almost every sentence written by them …
Anthropic Deputy CISO Jason Clinton details how the Security Engineering team secures a software development lifecycle where AI authors 80% of merged code, with software engineers shipping 8x more cod…
A developer building Almide, a programming language designed for AI agents, found that syntax choices matter less than error recovery. By measuring Modification Survival Rate (MSR) — whether a model c…
An SEO professional with six years of experience shares nine mistakes made while using Anthropic's Claude AI assistant, emphasizing that even the best AI tool is ineffective without organized usage. T…
Open-weight models like Kimi K3 are pressuring the margins of closed-source labs such as OpenAI and Anthropic, potentially shifting AI profit toward chipmakers and cloud providers. Kimi K3, a Chinese …
A model's per-token price does not determine actual cost; Kimi K2 at roughly half the per-token price of GPT-5.1 class models still lands close to the same cost per completed task at about 95 cents ve…
Anthropic has expanded its voice mode capability to the more powerful Claude Opus and Sonnet models, enabling developers to integrate advanced multimodal conversational AI with real-time voice interac…
A performance analysis by an anonymous developer comparing Claude models in English versus Slovenian found that frontier models retain 97-98% of English performance in major languages, with a negligib…
Anthropic announced on Thursday that Claude's voice mode now runs on its Opus, Sonnet, and Haiku models, enabling deeper conversations and mid-conversation model switching. Users can also speak to thi…
Anthropic upgraded Claude voice mode to work with Opus and Sonnet models for the first time, expanding beyond the previous Haiku-only support. The company says users can now switch models during voice…
Anthropic expanded Claude voice mode on July 23 to support Opus and Sonnet in addition to Haiku, allowing paid users to switch among models during a conversation. The beta update, available across mob…
Google DeepMind on July 21 released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, with 3.6 Flash offering 17% fewer output tokens and pricing cut to $7.50 per million tokens, wh…
Prediction markets on Polymarket forecast a median release date for Anthropic's next Claude Opus model by Dec 31, 2026, with a 99% probability by that date. Google's next Gemini Pro model has a median…
A single attack payload hijacked Claude Code and OpenAI's Codex unchanged, according to an exploit brief published July 9 by Boyan Milanov and Heidy Khlaaf at the AI Now Institute. The 'Friendly Fire'…
A real-world incident response comparison between Claude Opus and GPT Codex reveals a significant gap in autonomous problem-solving capabilities. Claude Opus independently investigated a user signup f…
A pre-registered study on arXiv (2607.17420v1) found that biographical personas in system prompts produce model-dependent effects on LLM code generation, with the research-librarian persona causing Cl…
Feyn Labs shipped SQRL on July 19, a text-to-SQL model that inspects databases with read-only probes before writing queries, achieving 70.6% execution accuracy on BIRD Dev and edging past Claude Opus …
Feyn AI, a YC-backed startup, released SQRL, a family of text-to-SQL models that inspect the database before writing a query. The flagship SQRL-35B-A3B achieves 70.6% execution accuracy on BIRD Dev, o…