cd /news/ai-agents/sage-a-statistical-acceptance-gate-f… · home › topics › ai-agents › article
[ARTICLE · art-142216] src=arxiv.org ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

SAGE: A Statistical Acceptance Gate for Self-Evolving Agents

A new arXiv paper (2609.36043v1) proposes SAGE, a statistical acceptance gate for self-evolving LLM agents that replaces the standard rule of keeping any edit which improves an aggregate validation score. SAGE uses a per-item paired comparison and a one-sided paired test to commit an edit only when its wins are statistically reliable, and across five benchmarks and four backbone LLMs under an equal-budget protocol it lowered the regression rate in 19 of 20 settings, including from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA with DeepSeek-V4. SAGE also attained the highest final score in all 20 settings, raising LiveMath from 34.15 to 48.78.

by read2 min views2 publishedSep 30, 2026

arXiv:2609.36043v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Second, it is vulnerable to the Optimizer's Curse, since the best observed score on a finite and noisy validation set is upward biased. To solve the above two limitations, we propose a statistical acceptance gate for self-evolving agents (SAGE). Compared with previous work, SAGE has two contributions. First, SAGE proposes a per-item paired comparison that evaluates the current skill and the edited skill on identical validation items, which exposes regressions that an aggregate score hides and penalizes them asymmetrically. Second, SAGE also employs a one-sided paired test that commits an edit only when its wins are statistically reliable against its losses, and it abstains otherwise. SAGE is a conservative refinement of the standard gate that recovers the baseline exactly at a boundary setting. It commits only a subset of the baseline's edits, filtering out those whose gains are unreliable or purchased by breaking already-solved items. Across five benchmarks and four backbone LLMs under an equal-budget protocol, SAGE lowers the regression rate in 19 of 20 settings and matches the baseline in the remaining one, for example from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA with DeepSeek-V4. SAGE also attains the highest final score in all 20 settings, raising LiveMath from 34.15 to 48.78.

── more in #ai-agents 4 stories · sorted by recency
── more on @sage 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sage-a-statistical-a…] indexed:0 read:2min 2026-09-30 · —