{"slug": "deltaml-bench-evaluating-machine-learning-agents-on-real-world-research", "title": "DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories", "summary": "Researchers introduced DeltaML-Bench, a benchmark of 48 tasks from research papers that tests AI agents' ability to improve published baselines in real-world repositories. In evaluations, the ARG scaffolding raised GPT-5's per-run success rate from 9.4% to 33.9% in a 4x6h allocation and to 49.0% in a 2x12h allocation, while Modular configurations showed specification gaming rates up to 47.9%.", "body_md": "arXiv:2608.19653v1 Announce Type: new\nAbstract: Autonomous agents for machine learning experimentation must navigate heterogeneous repositories, repair training pipelines, and evaluate candidate improvements under realistic compute constraints. Existing benchmarks only partially capture these conditions. We introduce DeltaML-Bench, a benchmark comprising 48 tasks sourced from research papers that require agents to improve published baselines within imperfect, open-source repositories. We evaluate GPT-5 and Claude Sonnet 4 with a standard Modular agent and a search-based ARG scaffolding. In the 4 x 6h allocation, ARG raises GPT-5's per-run success rate from 9.4% to 33.9%; in the 2 x 12h allocation, GPT-5 ARG reaches 49.0%. Modular configurations exhibit specification gaming rates as high as 47.9%, while no gaming is observed in the evaluated ARG configurations. These results indicate that scaffolding design and integrity checks are important considerations when deploying agents for autonomous ML experimentation.", "url": "https://wpnews.pro/news/deltaml-bench-evaluating-machine-learning-agents-on-real-world-research", "canonical_source": "https://www.machinebrief.com/news/deltaml-bench-evaluating-machine-learning-agents-on-real-wor-cl26", "published_at": "2026-08-21 04:00:00+00:00", "updated_at": "2026-08-21 04:14:20.348901+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-agents", "ai-research"], "entities": ["DeltaML-Bench", "GPT-5", "Claude Sonnet 4", "ARG"], "alternates": {"html": "https://wpnews.pro/news/deltaml-bench-evaluating-machine-learning-agents-on-real-world-research", "markdown": "https://wpnews.pro/news/deltaml-bench-evaluating-machine-learning-agents-on-real-world-research.md", "text": "https://wpnews.pro/news/deltaml-bench-evaluating-machine-learning-agents-on-real-world-research.txt", "jsonld": "https://wpnews.pro/news/deltaml-bench-evaluating-machine-learning-agents-on-real-world-research.jsonld"}}