Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? Repository-level coding benchmarks such as SWE-bench suffer from data leakage because they are built on popular open-source repositories repeatedly used for LLM training, meaning strong performance may reflect memorization of canonical solutions rather than learned software-engineering skill, according to an analysis titled "Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?" The finding matters because SWE-bench has become the standard for evaluating coding agents, so leaked training data can inflate reported capability. Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built upon popular open-source repositories repeatedly used for training. Consequently, strong performance may reflect memorization of canonical rep