ProgramDistill: AI Coding Agents Still Fail Half the Time ProgramDistill, a benchmark for real-world app reconstruction tasks, found that AI coding agents fail 51% of the time, including GPT-6 Astra, according to byteiota. The result highlights a persistent gap between AI coding agent claims and performance on practical software reconstruction work. ProgramDistill exposes a critical gap: even GPT-6 Astra fails 51% of real-world app reconstruction tasks. Here is what the benchmark means for developers using AI coding agents. The post ProgramDistill: AI Coding Agents Still Fail Half the Time appeared first on byteiota .