DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories
Researchers introduced DeltaML-Bench, a benchmark of 48 tasks from research papers that tests AI agents' ability to improve published baselines in real-world repositories. In evaluations, the ARG scaf…