CITA improves Tool F1 and task success in long-horizon tool-use agents A new method called CITA improves tool-use agents by estimating the likelihood that a tool invocation will lead to task success, consistently raising Tool F1 and task success across three benchmarks, according to an arXiv paper (2610.02330). CITA uses a comparative inference model trained on a Bayesian tool-graph simulator so production agents can predict downstream task success before executing an action. The approach shifts agent architectures from reactive trial-and-error to proactive trajectory selection, which the paper says prevents expensive API spend on dead-end tool executions. arXiv https://arxiv.org/abs/2610.02330 CITA improves Tool F1 and task success in long-horizon tool-use agents Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated. Evaluating and ranking alternative tool calls using a comparative inference model trained on a Bayesian tool-graph simulator allows production agents to predict downstream task success before executing an action. This shifts agent architectures from reactive trial-and-error to proactive trajectory selection, directly increasing Tool F1 and task success rates while preventing expensive API spend on dead-end tool executions. A new method, CITA, improves tool-use agents by estimating the likelihood of a tool invocation leading to task success, and it consistently improves Tool F1 and task success across three benchmarks. This enables more accurate decision-making in long-horizon tool use for LLMs, directly impacting production agents' ability to choose the right tool invocations. This improvement matters for shipping reliable LLM-based applications that rely on complex tool invocation sequences.