cd /news/ai-agents/lookahead-r-budget-aware-tool-retrie… · home › topics › ai-agents › article
[ARTICLE · art-142247] src=arxiv.org ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning

Lookahead-R, a planning-based tool retrieval framework from an arXiv paper (arXiv:2609.35811v1), reached an NDCG@5 of 91.40% on the I3 split of the ToolBench benchmark, beating the state-of-the-art ToolGen (90.16%) by 1.24%. The framework uses a lightweight execution-aware surrogate world model that predicts tool execution success, latency cost, and semantic utility without invoking real APIs, driving a cost-sensitive, uncertainty-guided Monte Carlo Tree Search under strict budget constraints. Ablation studies found explicit latency modeling is the key discriminative signal for identifying high-quality tools under resource constraints.

by read1 min views2 publishedSep 30, 2026

arXiv:2609.35811v1 Announce Type: new Abstract: Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based validation improves precision at the cost of prohibitive latency. We propose Lookahead-R, a planning-based framework that reformulates tool retrieval as a resource-constrained sequential decision-making problem. At its core, Lookahead-R introduces a lightweight execution-aware surrogate world model that jointly predicts tool execution success, latency cost, and semantic utility---without invoking real APIs. This world model drives a cost-sensitive, uncertainty-guided Monte Carlo Tree Search that navigates the tool space under strict budget constraints. Evaluated on the large-scale ToolBench benchmark, Lookahead-R achieves a superior accuracy-efficiency trade-off across all test scenarios. On the most challenging I3 split, it attains an NDCG@5 of 91.40%, outperforming the state-of-the-art ToolGen (90.16%) by 1.24%. Ablation studies confirm that explicit latency modeling is the key discriminative signal for identifying high-quality tools under resource constraints.

── more in #ai-agents 4 stories · sorted by recency
── more on @lookahead-r 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/lookahead-r-budget-a…] indexed:0 read:1min 2026-09-30 · —