{"slug": "sbco-self-supervised-verifier-grounded-harness-optimization-for-planning-agents", "title": "SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents", "summary": "Researchers introduced SBCO (Self-supervised Block Coordinate Optimizer), a verifier-grounded harness optimizer for planning agents that improves agent outputs from its own graded feedback without human labels or self-modification. In tests across two domains, SBCO matched or exceeded a customized self-modifying baseline while using 4-5.5 times less compute budget.", "body_md": "arXiv:2608.10157v1 Announce Type: new\nAbstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over time. Recently, methods like the Darwin G\\\"odel Machine and the Huxley G\\\"odel Machine have been proposed which enable open-ended, recursive self-improvement through self-reference where a coding agent edits its own code. Such self-referential self-improvement methods require that the competence required to perform the task coincides or aligns well with the competence required for self-modification which is the case for coding tasks. For domains or tasks, which do not satisfy the alignment needed, self-referential self-improvement is not available. In such cases, it is possible to adapt the above algorithms to other tasks by removing the self-referential aspect or introducing explicit self-modification of a meta-agent -- both computationally expensive, relying on population or self-modification search over many candidate agents. For planning tasks with explicit constraints, we propose a far cheaper alternative. We introduce SBCO (Self-supervised Block Coordinate Optimizer), a verifier-grounded harness optimizer in the same closed-loop, improve-from-experience family as the G\\\"odel-machine methods, but self-supervised rather than self-referential. Given an agentic harness, SBCO learns a decomposed bank of verifiers and a harness policy via approximate block coordinate ascent, improving the agent's outputs from its own graded feedback---with a fixed meta-agent and no human labels. Across two domains SBCO matches or exceeds a customized self-modifying baseline while using 4-5.5 times less compute budget.", "url": "https://wpnews.pro/news/sbco-self-supervised-verifier-grounded-harness-optimization-for-planning-agents", "canonical_source": "https://arxiv.org/abs/2608.10157", "published_at": "2026-08-12 04:00:00+00:00", "updated_at": "2026-08-12 04:18:12.706995+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-agents"], "entities": ["SBCO", "Darwin Gödel Machine", "Huxley Gödel Machine"], "alternates": {"html": "https://wpnews.pro/news/sbco-self-supervised-verifier-grounded-harness-optimization-for-planning-agents", "markdown": "https://wpnews.pro/news/sbco-self-supervised-verifier-grounded-harness-optimization-for-planning-agents.md", "text": "https://wpnews.pro/news/sbco-self-supervised-verifier-grounded-harness-optimization-for-planning-agents.txt", "jsonld": "https://wpnews.pro/news/sbco-self-supervised-verifier-grounded-harness-optimization-for-planning-agents.jsonld"}}