04:00
2026-07-24
arxiv.org
artificial-intelligence
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
Researchers introduced DataPrep-Bench, the first unified benchmark for evaluating how well large language models (LLMs) and agents prepare training data, covering both data construction and data qualiβ¦