{"slug": "pyrodash-cost-efficient-token-level-small-large-language-model-collaborative", "title": "PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference", "summary": "Researchers introduce PyroDash, a cost-aware framework for token-level small-large language model collaborative inference that reduces large language model (LLM) usage while preserving reasoning performance. Across five mathematical reasoning benchmarks, PyroDash achieves 64.04% average accuracy with a 20.4% cost reduction at one operating point, and at another reduces total inference cost from USD 49.36 to USD 1.78 with a 1.90% LLM token ratio. The framework trains a small language model to emit control tokens for selective LLM assistance, requiring no separate router or LLM retraining.", "body_md": "arXiv:2607.20327v1 Announce Type: new\nAbstract: Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware framework for token-level SLM-LLM collaborative inference. During generation, the SLM decides whether to request assistance by emitting a control token. A Collaborate Engine then sends the query and partial reasoning trace to a frozen LLM for completion through a single handoff. The policy is internalized in the SLM, requiring neither a separate router, LLM retraining, nor access to LLM logits. PyroDash trains the SLM in three stages: control-token embedding learning, offloading-oriented supervised fine-tuning, and cost-aware alignment with Group Relative Policy Optimization. Its reward balances answer accuracy against inference cost normalized by LLM-only inference. Across five mathematical reasoning benchmarks, PyroDash supports different accuracy-cost operating points. With $\\lambda=0.05$, it achieves 64.04 percent average accuracy, 6.36 percentage points above the LLM-only baseline, while reducing cost by 20.4 percent. With $\\lambda=0.6$, it achieves 54.55 percent accuracy with a 1.90 percent LLM token ratio and 0.012 LLM calls per example, reducing total cost from USD 49.36 to USD 1.78. These results show that learned token-level handoffs can reduce LLM use while preserving strong reasoning performance.", "url": "https://wpnews.pro/news/pyrodash-cost-efficient-token-level-small-large-language-model-collaborative", "canonical_source": "https://www.machinebrief.com/news/pyrodash-cost-efficient-token-level-small-large-language-mod-if7r", "published_at": "2026-07-23 04:00:00+00:00", "updated_at": "2026-07-23 05:34:27.058780+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-infrastructure"], "entities": ["PyroDash", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/pyrodash-cost-efficient-token-level-small-large-language-model-collaborative", "markdown": "https://wpnews.pro/news/pyrodash-cost-efficient-token-level-small-large-language-model-collaborative.md", "text": "https://wpnews.pro/news/pyrodash-cost-efficient-token-level-small-large-language-model-collaborative.txt", "jsonld": "https://wpnews.pro/news/pyrodash-cost-efficient-token-level-small-large-language-model-collaborative.jsonld"}}