{"slug": "evaluating-skills-not-just-agents-agentic-continuous-evaluation-of-skills", "title": "Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills", "summary": "NVIDIA researchers introduced ACES (Agentic Continuous Evaluation of Skills), a repository-native framework that evaluates enterprise agent skills by running paired live trials with and without a target skill, reporting a mean composite Skill Lift of 0.2134 across 947 scored paired cases from 58 of 64 production skills. The framework, available as open-source in NVIDIA SkillEvaluator, found that scan-only gates measure complementary facets (structural versus LLM-judge Spearman ρ = 0.14) and that composite lift was positive in 72.8% of paired cases, with the largest gains in skill execution, behavior check, and skill efficiency.", "body_md": "# Computer Science > Artificial Intelligence\n\n[Submitted on 20 Aug 2026]\n\n# Title:Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills\n\n[View PDF](/pdf/2608.20614)\n\n[HTML (experimental)](https://arxiv.org/html/2608.20614v1)\n\nAbstract:Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, style, and security, but they do not answer the deployment question: does the capability package help a live agent complete enterprise tasks under the same model, sandbox, and grading policy?\n\nWe present ACES (Agentic Continuous Evaluation of Skills), a repository-native framework for evaluating skills and product capability packages as executable agent artifacts. ACES runs paired live trials with and without a target skill, normalizes trajectories into the Agent Trajectory Interchange Format (ATIF), grades six default runtime metrics, and reports Skill Lift: the target skill's added value for a fixed task, harness, workspace, and scorer. The same protocol supports product-owned task suites that compare baseline, skill, bundle, team-skill, and plugin targets.\n\nOn 145 real skills from internal enterprise repositories and public catalogs, scan-only gates surface useful authoring issues but measure complementary facets (structural versus LLM-judge Spearman $\\rho = 0.14$). Across 947 scored paired cases from 58 of 64 production skills and four primary harnesses, mean composite Skill Lift is 0.2134 (95\\% paired-case CI [0.1967, 0.2301]); mean outcome-only lift, the average of accuracy and goal accuracy, is 0.1799. Composite lift is positive in 72.8\\% of paired cases. The largest process-metric gains appear in skill execution, behavior check, and skill efficiency---signals about discovery, routing, workflow following, and tool use that document scans cannot observe. An open-source implementation of the methodology is available in NVIDIA SkillEvaluator.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/evaluating-skills-not-just-agents-agentic-continuous-evaluation-of-skills", "canonical_source": "https://arxiv.org/abs/2608.20614", "published_at": "2026-08-30 16:09:31+00:00", "updated_at": "2026-08-30 16:21:56.754779+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-tools", "ai-agents"], "entities": ["NVIDIA", "ACES", "NVIDIA SkillEvaluator", "Agent Trajectory Interchange Format (ATIF)"], "alternates": {"html": "https://wpnews.pro/news/evaluating-skills-not-just-agents-agentic-continuous-evaluation-of-skills", "markdown": "https://wpnews.pro/news/evaluating-skills-not-just-agents-agentic-continuous-evaluation-of-skills.md", "text": "https://wpnews.pro/news/evaluating-skills-not-just-agents-agentic-continuous-evaluation-of-skills.txt", "jsonld": "https://wpnews.pro/news/evaluating-skills-not-just-agents-agentic-continuous-evaluation-of-skills.jsonld"}}