cd /news/artificial-intelligence/sbco-self-supervised-verifier-ground… · home topics artificial-intelligence article
[ARTICLE · art-93056] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents

Researchers introduced SBCO (Self-supervised Block Coordinate Optimizer), a verifier-grounded harness optimizer for planning agents that improves agent outputs from its own graded feedback without human labels or self-modification. In tests across two domains, SBCO matched or exceeded a customized self-modifying baseline while using 4-5.5 times less compute budget.

read1 min views1 publishedAug 12, 2026

arXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over time. Recently, methods like the Darwin G"odel Machine and the Huxley G"odel Machine have been proposed which enable open-ended, recursive self-improvement through self-reference where a coding agent edits its own code. Such self-referential self-improvement methods require that the competence required to perform the task coincides or aligns well with the competence required for self-modification which is the case for coding tasks. For domains or tasks, which do not satisfy the alignment needed, self-referential self-improvement is not available. In such cases, it is possible to adapt the above algorithms to other tasks by removing the self-referential aspect or introducing explicit self-modification of a meta-agent -- both computationally expensive, relying on population or self-modification search over many candidate agents. For planning tasks with explicit constraints, we propose a far cheaper alternative. We introduce SBCO (Self-supervised Block Coordinate Optimizer), a verifier-grounded harness optimizer in the same closed-loop, improve-from-experience family as the G"odel-machine methods, but self-supervised rather than self-referential. Given an agentic harness, SBCO learns a decomposed bank of verifiers and a harness policy via approximate block coordinate ascent, improving the agent's outputs from its own graded feedback---with a fixed meta-agent and no human labels. Across two domains SBCO matches or exceeds a customized self-modifying baseline while using 4-5.5 times less compute budget.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @sbco 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sbco-self-supervised…] indexed:0 read:1min 2026-08-12 ·