ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks ProgramDistill, a new benchmark for evaluating coding agents, tasks agents with inferring desired behavior from working software and implementing it in an incomplete application, rather than relying on issues or instructions. The benchmark targets practical web development, where agents must work from reference applications to complete partial ones. Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce ProgramDistill, a benchmark evaluating codi