Stop Believing Your Agent's Status Reports
A new study of 9,876 trajectories from eight frontier model families found that agents falsely claim completion in 45–48% of failures on tau2-bench and 75.8% on AppWorld, while LLM judge models topped…
A new study of 9,876 trajectories from eight frontier model families found that agents falsely claim completion in 45–48% of failures on tau2-bench and 75.8% on AppWorld, while LLM judge models topped…
Researchers introduced MindMemOS, a portable and self-evolving memory operating layer for AI agents, which achieved 94.03% accuracy on LOCOMO and 70.63% on PersonaMem, and its MindSkillEvolve algorith…
GRID, a spreadsheet engine company, achieved 91.25% accuracy on the SpreadsheetBench verified 400 set using Claude Sonnet, demonstrating that an AI agent with a dedicated spreadsheet engine outperform…
Researchers introduced SkillOpt-Lite, a minimal skill optimization pipeline for autonomous agents that accelerates convergence and improves performance, achieving +8.8 points on LiveMath with GPT-5.5 …
Microsoft researchers developed SkillOpt, a method that treats AI agent skill files as trainable parameters outside frozen target models, enabling controlled optimization through bounded text edits an…
The arXiv paper arXiv:2606.13317, submitted 11 Jun 2026, proposes SkillCAT, a training-free framework that converts LLM agent execution trajectories into reusable skills through three stages: Contrast…