cd /news/artificial-intelligence/repo2skill-evo-repository-skills-go-… · home topics artificial-intelligence article
[ARTICLE · art-109698] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Repo2Skill-Evo: Repository Skills Go Stale in Silence

A new arXiv preprint, Repo2Skill-Evo, finds that repository-specific skills used by large language model (LLM) agents become stale silently after software releases, and even frontier agents struggle to maintain them. Across 57 real-world repositories and 105 release transitions, every transition invalidated part of the V1 skill set, yet six frontier agents achieved only 29.9%-69.7% avg@3 macro F1 under a patch-grounded removal metric. The study highlights that incomplete coverage and overbroad editing are the primary errors, indicating that repository skills go stale without explicit signals.

read1 min views1 publishedAug 25, 2026

arXiv:2608.21964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over evolving software repositories, where success depends on repository-specific procedural knowledge: which APIs to call, which scripts to run, and which conventions the current release expects. Agent skills externalize this knowledge into reusable units, and prior work shows that they can improve agent performance. What remains unclear is whether that improvement is durable. The same version specificity that makes a skill useful also makes it fragile: after a release, it may become stale without raising any explicit signal, while continuing to provide obsolete guidance. Externalizing knowledge into a skill can therefore make its decay invisible. We study whether agents can keep this externalized knowledge current. Repo2Skill-Evo casts each release transition as a skill-maintenance task: given a V1 skill set and the official V1-to-V2 patch, an agent must update obsolete skill content while preserving guidance that remains valid. Across 57 real-world repositories and 105 selected release transitions, every evaluated transition invalidates part of the V1 skill set. Yet six frontier agents reach only 29.9%-69.7% avg@3 macro F1 under a patch-grounded removal metric that balances stale-content recall against over-editing precision. Across runs, two opposing errors dominate: incomplete coverage of affected files in the skill set leaves stale content untouched, while overbroad editing is associated with higher recall but lower precision. Repository skills go stale in silence, and even frontier agents cannot reliably maintain them.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/repo2skill-evo-repos…] indexed:0 read:1min 2026-08-25 ·