cd /news/large-language-models/when-synthetic-data-hurts-on-catastr… · home topics large-language-models article
[ARTICLE · art-126616] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents

A production skill router covering 34,396 skills shows that fine-tuning on synthetic data improves in-distribution skill retrieval for LLM agents but causes catastrophic forgetting on real and out-of-distribution data, according to an arXiv paper (2609.10750v1). Forgetting-mitigation approaches borrowed from continual learning — embedding-anchor regularization, Learning without Forgetting (LwF), Elastic Weight Consolidation (EWC), and L2-initialization — retained OOD retrieval performance and improved synthetic in-distribution retrieval by 13.98% for a 0.6B Qwen retriever and reranker. The authors present the results as a practical benchmark and fine-tuning recipe for scarce, multi-positive supervision.

by read1 min views2 publishedSep 11, 2026

arXiv:2609.10750v1 Announce Type: cross Abstract: LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large repositories a critical challenge. We present a production skill router over 34,396 skills and a large-scale study of skill retrieval using limited real supervision and synthetic data. We found that the synthetic-data fine-tuning improves in-distribution retrieval but it causes catastrophic forgetting on real and out-of-distribution (OOD) data. We evaluate several forgetting mitigation fine-tuning approaches inspired by continual learning, including embedding-anchor regularization, Learning without Forgetting (LwF), Elastic Weight Consolidation (EWC), and L2-initialization. The results show that these approaches not only retain the performance on OOD skills retrieval but also improve the retrieval on synthetic in-distribution skills by 13.98% for 0.6B Qwen retriever and reranker. Our results provide a practical benchmark and a robust fine-tuning recipe for scarce, multi-positive supervision.

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/when-synthetic-data-…] indexed:0 read:1min 2026-09-11 ·