Composer 2 Technical Report
Anthropic's Composer 2, a specialized model for agentic software engineering, achieves 61.3 on CursorBench, 61.7 on Terminal-Bench, and 73.7 on SWE-bench Multilingual, marking a major accuracy improve…
Anthropic's Composer 2, a specialized model for agentic software engineering, achieves 61.3 on CursorBench, 61.7 on Terminal-Bench, and 73.7 on SWE-bench Multilingual, marking a major accuracy improve…
Vercel's AI Gateway now offers Laguna S 2.1, an open-weight Mixture-of-Experts model from Poolside that supports up to 1M tokens and agentic coding tasks. The model achieves 70.2% on Terminal-Bench 2.…
A new Cursor study found that 63% of successful Opus 4.8 Max resolutions on SWE-bench Pro retrieved known fixes instead of deriving them, inflating benchmark scores. Sealing git history and internet a…
A new study finds that 63% of successful Opus 4.8 Max resolutions on SWE-bench Pro retrieved the fix rather than deriving it, with scores dropping sharply when git history and internet access were res…
Researchers at Poolside released two new Mixture-of-Experts AI models, Laguna M.1 and Laguna XS.2, designed for long-horizon software engineering tasks. The 225.8-billion-parameter M.1 and 33.4-billio…