cd /news/artificial-intelligence/search-g1-grounded-search-agents-via… · home topics artificial-intelligence article
[ARTICLE · art-91390] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

Researchers propose Search-G1, a representation-based intrinsic reward framework that improves the grounding-search-cost trade-off for search-augmented language agents by using two intervention-calibrated readouts to measure answer grounding. Experiments across multiple search-based question-answering benchmarks and two model scales show that Search-G1 produces shorter response-side trajectories at competitive task accuracy without requiring process annotations or LLM-as-judge inference during policy optimization. Code is available at https://github.com/Rosy0912/Search-G1.

read1 min views1 publishedAug 11, 2026

arXiv:2608.07531v1 Announce Type: new Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training. Internal rewards based on policy-side signals such as entropy, likelihood, or information gain are graded and inexpensive to evaluate, yet mainly reflect model confidence rather than evidence grounding. We propose Search-G1, a representation-based intrinsic reward framework that measures the operational grounding of an agent's answers through two intervention-calibrated readouts. A prompt-state readout predicts closed-book sufficiency, whose complement defines policy-relative retrieval necessity; an answer-commit readout estimates evidence reliance from answer-stage sensitivity to evidence deletion. Together, they provide additional credit to correct searched trajectories when retrieval is estimated necessary and the answer is evidence-sensitive, favor correct direct answers when closed-book knowledge suffices, and penalize repeated search. After calibration, reward scoring requires neither process annotations nor LLM-as-judge inference during policy optimization. Because reinforcement learning changes policy representations, Search-G1 periodically refits both readouts on trajectories from the latest checkpoint, allowing the reward to co-evolve with the policy. Experiments across multiple search-based question-answering benchmarks and two model scales show that Search-G1 improves the grounding--search-cost trade-off, producing shorter response-side trajectories at competitive task accuracy. Code is available at https://github.com/Rosy0912/Search-G1.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @search-g1 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/search-g1-grounded-s…] indexed:0 read:1min 2026-08-11 ·