{"slug": "not-all-tokens-are-equal-inflation-aware-routing-for-agentic-llm-systems", "title": "Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems", "summary": "A new arXiv paper introduces InflationAgent, a four-stage router that accounts for token inflation—the ratio of true workflow cost to single-call cost—which can exceed 2x on difficult tasks. The system, which uses a pre-execution difficulty signal called CoT Branching Entropy (CBE) with AUROC 0.887, achieves 94.7% accuracy on GSM8K under a fixed budget versus 91.0% for FrugalGPT while using 31% fewer tokens. The paper also finds that forwarding a failed reasoning chain to GPT-4o reduces its accuracy by up to 34.8 percentage points, validating the fresh-escalation design.", "body_md": "arXiv:2608.13571v1 Announce Type: new\nAbstract: When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overhead creates a gap between what a model's per-token price implies and what a full workflow actually costs. We call this gap \\emph{token inflation} and define it as the ratio of true workflow cost to single-call cost. Systems like FrugalGPT route based on the latter, which can underestimate real cost by more than $2\\times$ on difficult tasks. We address this with InflationAgent, a four-stage router that (1) measures token inflation systematically across model tiers and task types, finding inflation as high as $4.25\\times$ for a 7B model on multi-hop question answering; (2) introduces CoT Branching Entropy (CBE), a pre-execution difficulty signal computed entirely from local inference, which predicts high inflation with AUROC 0.887; and (3) selects models by maximizing a Semantic Exchange Rate (SER) that divides expected accuracy by predicted true cost, with a fresh-escalation policy that discards failed chains before routing to a stronger model. On GSM8K under a fixed budget, InflationAgent achieves 94.7\\% accuracy versus 91.0\\% for FrugalGPT while using 31\\% fewer tokens, and we show that forwarding a failed reasoning chain to GPT-4o reduces its accuracy by up to 34.8 percentage points, validating the fresh-escalation design.", "url": "https://wpnews.pro/news/not-all-tokens-are-equal-inflation-aware-routing-for-agentic-llm-systems", "canonical_source": "https://arxiv.org/abs/2608.13571", "published_at": "2026-08-17 04:00:00+00:00", "updated_at": "2026-08-17 04:13:53.739347+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research"], "entities": ["InflationAgent", "FrugalGPT", "GPT-4o", "GSM8K", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/not-all-tokens-are-equal-inflation-aware-routing-for-agentic-llm-systems", "markdown": "https://wpnews.pro/news/not-all-tokens-are-equal-inflation-aware-routing-for-agentic-llm-systems.md", "text": "https://wpnews.pro/news/not-all-tokens-are-equal-inflation-aware-routing-for-agentic-llm-systems.txt", "jsonld": "https://wpnews.pro/news/not-all-tokens-are-equal-inflation-aware-routing-for-agentic-llm-systems.jsonld"}}