04:01
2026-08-04
trae1oung.github.io
artificial-intelligence
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
Researchers introduced SWE-Touch, a benchmark evaluating coding agents' ability to repair software after users directly edit code, finding that most models' performance drops significantly when user eโฆ