cd /news/ai-agents/traverse-learning-when-to-remember-r… · home › topics › ai-agents › article
[ARTICLE · art-142272] src=machinebrief.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

A 35B-parameter model trained with a three-state autonomous search harness (Rubric, Answer, Verify) and a Seal Memory tool scored 72.83 on BrowseComp, outperforming comparable open-source systems, according to the arXiv paper 2609.37082v1. The authors report that reinforcement learning on this behavior can induce "Seal Collapse," causing unstable training, and solve it by training only the final segment after context management. The model also improves over its base model on BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search, with ablations showing autonomous compression beats automatic compaction.

by read1 min views1 publishedSep 30, 2026

arXiv:2609.37082v1 Announce Type: new Abstract: Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.

── more in #ai-agents 4 stories · sorted by recency
── more on @browsecomp 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/traverse-learning-wh…] indexed:0 read:1min 2026-09-30 · —