cd /news/artificial-intelligence/self-supervised-skill-optimization · home topics artificial-intelligence article
[ARTICLE · art-84191] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Self-Supervised Skill Optimization

Researchers introduced Self-Supervised Skill Optimization (SSO), a framework that optimizes agent skills for frozen large language models using only unlabeled task instances, eliminating the need for ground-truth labels, rewards, or task-specific evaluators. SSO outperforms existing ground-truth-free prompt optimizers on closed-ended and open-ended tasks, and on closed-ended benchmarks it approaches or exceeds the strongest ground-truth-based skill optimizer without using any ground-truth feedback.

read1 min views1 publishedAug 3, 2026

arXiv:2607.28777v1 Announce Type: new Abstract: Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized with ground-truth (GT) feedback. Many applications, however, lack GT labels, task scores, rewards, or reliable task-specific evaluators. We therefore introduce Self-Supervised Skill Optimization (SSO), a comparative framework that learns a reusable skill from unlabeled task instances alone. At each step, SSO runs the current skill on an unlabeled batch, uses a subset of the resulting executions to generate complete skill probes, and runs the probes on the same batch. An LLM judge compares the resulting answers, trajectories, artifacts, or terminal states. A separate behavior extractor identifies behavioral differences without seeing the judge's decisions. SSO uses these decisions to aggregate evidence for and against the observed behaviors across instances. It then ranks the behaviors by the resulting evidence and renders a new complete skill from the highest-ranked behaviors. The update is accepted only if the new skill outperforms the current one on an unlabeled validation set. SSO outperforms existing GT-free prompt optimizers on both closed-ended and open-ended tasks. On closed-ended benchmarks, it approaches and sometimes exceeds the strongest GT-based skill optimizer without using any GT feedback.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @self-supervised skill optimization (sso) 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/self-supervised-skil…] indexed:0 read:1min 2026-08-03 ·