# CITA improves Tool F1 and task success in long-horizon tool-use agents

> Source: <https://www.snipvote.com/story/cmuuxe1w10006140xoe44itn3>
> Published: 2026-10-05 08:16:25.545512+00:00

[arXiv](https://arxiv.org/abs/2610.02330)

### CITA improves Tool F1 and task success in long-horizon tool-use agents

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Evaluating and ranking alternative tool calls using a comparative inference model trained on a Bayesian tool-graph simulator allows production agents to predict downstream task success before executing an action. This shifts agent architectures from reactive trial-and-error to proactive trajectory selection, directly increasing Tool F1 and task success rates while preventing expensive API spend on dead-end tool executions.

A new method, CITA, improves tool-use agents by estimating the likelihood of a tool invocation leading to task success, and it consistently improves Tool F1 and task success across three benchmarks. This enables more accurate decision-making in long-horizon tool use for LLMs, directly impacting production agents' ability to choose the right tool invocations. This improvement matters for shipping reliable LLM-based applications that rely on complex tool invocation sequences.
