AI agents write Ruby but can't navigate it: a 5-model, 13-codebase benchmark
A benchmark testing five AI models across 13 real Ruby codebases found that agents equipped with a structural code map achieved a mean cited recall lift of +0.26 for Claude Opus 4.8, with gains concenβ¦