{"slug": "gitskills-a-dataset-of-agent-skills-on-github", "title": "GitSkills: A Dataset of Agent Skills on GitHub", "summary": "Researchers released GitSkills, a dataset of 3,797,117 SKILL.md files collected from 282,200 public GitHub repositories in July 2026, capturing the population of agent skills introduced by Anthropic in October 2025. The dataset, stored in a single SQLite file, groups identical files into 1,877,981 distinct contents and enriches representatives with metadata and commit history to support research on adoption, reuse, structure, authorship, maintenance, and security of agent skills.", "body_md": "# Computer Science > Software Engineering\n\n[Submitted on 11 Aug 2026]\n\n# Title:GitSkills: A Dataset of Agent Skills on GitHub\n\n[View PDF](/pdf/2608.10906)\n\n[HTML (experimental)](https://arxiv.org/html/2608.10906v1)\n\nAbstract:An agent skill is a folder containing a[this http URL]file with instructions for a language-model agent, optionally accompanied by scripts and reference files. The agent loads the skill when it judges that a task matches the skill description. Anthropic introduced the format in October 2025 as an open specification. Nine months later, we find that skill files in the millions sit in public GitHub repositories. Skills are unlike the artifacts the SE research community usually mines: they are written mainly in natural language, a model selects them probabilistically at run time, and no compiler or type checker verifies the selection. They also have no central registry or package manager, so they spread by copying folders between repositories. How developers write, reuse, and maintain skills is therefore an empirical question, and no existing dataset records this population. We present GitSkills, a dataset of 3,797,117[this http URL]files collected from 282,200 public repositories in July 2026. The dataset retains every file occurrence with its repository, path, and content hash. It groups identical files into 1,877,981 distinct contents and enriches one representative per group with the full text, parsed front matter, folder contents, repository metadata, and, for a subset, the commit history of the file. A single self- contained SQLite file supports research on the adoption, reuse, structure, authorship, maintenance, and security of agent skills.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/gitskills-a-dataset-of-agent-skills-on-github", "canonical_source": "https://arxiv.org/abs/2608.10906", "published_at": "2026-08-13 09:19:47+00:00", "updated_at": "2026-08-13 09:40:59.147930+00:00", "lang": "en", "topics": ["ai-agents", "ai-research", "developer-tools"], "entities": ["GitSkills", "Anthropic", "GitHub", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/gitskills-a-dataset-of-agent-skills-on-github", "markdown": "https://wpnews.pro/news/gitskills-a-dataset-of-agent-skills-on-github.md", "text": "https://wpnews.pro/news/gitskills-a-dataset-of-agent-skills-on-github.txt", "jsonld": "https://wpnews.pro/news/gitskills-a-dataset-of-agent-skills-on-github.jsonld"}}