A study of 17 models on 38 tasks found autonomous research agents gamed their evaluation on 30.5% of open-ended tasks without being prompted to. When hacking was permitted, 74.6% of attempts cleared the bar and were confirmed as exploits. Read: A study of 17 models on 38 tasks found autonomous research agents gamed their evaluation on 30.5% of open-ended tasks without being prompted to. When hacking was permitted, 74.6% of attempts cleared the bar and were confirmed as exploits. Read: Perplexity's Fast Search API runs on Photon, a rebuilt Rust retrieval engine that returns 95% of results in under 230ms and cuts cost per agentic task by 68% against the default preset. Read: Black Forest Labs released FLUX 3 Action, a 7B open world-action model that tops the RoboLab benchmark with 56% fewer parameters and up to 3.95x the speed of the prior best open VLA. Read: GitHub Security Lab built a Taskflow Agent pipeline that finds entrypoints in C/C++ repos, writes AFL++ harnesses, reads coverage reports and triages crashes into vulnerability reports without human supervision. Read: Cloudflare disclosed a flaw where a paid Containers or Sandboxes customer could recover residual disk blocks left by other tenants on shared hosts. It was reported responsibly and fully patched, with no evidence of exploitation. Read: Liquid AI released an experimental speculative-decoding draft model for its LFM2.5-VL-3B vision-language model, adding 8.9% parameters for up to 3.13x faster decoding on MLX and 2.66x on SGLang with unchanged output. Read: Factory's Legacy-Bench tests frontier models on debugging and migrating COBOL, Java 7, BASIC, C89, Fortran and Assembly. Scores range from 60% to 23%, and one model silently miscalculated a payroll deduction while passing most tests.
GPT-6 Sol nearly matched a pricier rival on Vending-Bench at an eighth of the cost