cd /news/artificial-intelligence/beyond-single-turn-confidence-trajec… · home topics artificial-intelligence article
[ARTICLE · art-94717] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

A new arXiv study (2608.11552v1) finds that uncertainty quantification (UQ) methods for LLM agents do not transfer uniformly from single-turn outputs to interactive trajectories. Across five LLMs and four multi-turn tool-use datasets from BFCL-v4 and τ²-bench, the researchers evaluated white-box token-probability scorers, black-box consistency scorers, and reflexive self-assessment scorers, finding that black-box self-consistency, especially trajectory-equivalence and action-set consistency, is often the strongest UQ family, while token-probability scores are highly sensitive to aggregator choice. The authors conclude that UQ methods should be revalidated at the trajectory level, with attention to consistency measurement, aggregator choice, and computational budget.

read1 min views1 publishedAug 13, 2026

arXiv:2608.11552v1 Announce Type: new Abstract: Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome. We study whether three common families of single-turn UQ methods transfer to this setting. Across five LLMs and four multi-turn tool-use datasets from BFCL-v4 and $\tau^2$-bench, we evaluate white-box scorers based on action-token probabilities, black-box consistency scorers based on resampled trajectories, and reflexive scorers based on model self-assessment of the trajectory. We find that transfer is often useful but uneven. Token-probability scores are highly sensitive to the choice of aggregator used across turns, reflexive scores provide the strongest low-cost baseline in most evaluated settings, and black-box self-consistency is often the strongest UQ family, with trajectory-equivalence and action-set consistency typically ranking highest among its variants. These results suggest that UQ methods developed for single generations should be revalidated at the trajectory level, with careful attention to the consistency measurement, aggregator choice, and computational budget.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/beyond-single-turn-c…] indexed:0 read:1min 2026-08-13 ·