{"slug": "routing-without-training-controllable-ratio-llm-offloading-via-reliability", "title": "Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating", "summary": "Researchers propose CARGO, a training-free routing framework for local-cloud LLM collaboration that uses the local model's own inference-time agreement across sampled responses to decide when to offload to a stronger cloud model, eliminating the need for trained routers or collaboration-aware finetuning. In tests across diverse reasoning and QA tasks with multiple LLM families and scales, CARGO outperforms other training-free baselines and in several settings surpasses supervised learned routers, supporting arbitrary target collaboration ratios through lightweight calibration.", "body_md": "arXiv:2607.20481v1 Announce Type: new\nAbstract: Local-cloud collaboration is a practical way to deploy large language models under resource constraints, but existing methods often rely on trained routers or collaboration-aware finetuning that tie routing behavior to a particular operating regime. In this work, we show that such training may be unnecessary: the local model's own inference-time agreement across sampled responses already provides a strong signal for deciding when to trust local execution and when to offload to a stronger cloud model. We propose CARGO, a training-free routing framework that estimates this agreement through prompt-varied sampling, applies Bayesian early stopping for sample-efficient uncertainty control, and supports arbitrary target collaboration ratios through lightweight deployment-time calibration. Across diverse reasoning and question-answering tasks, multiple local LLM families and scales, and both pretrained and finetuned local models, CARGO consistently outperforms other training-free baselines and in several settings surpasses supervised learned routers. These results suggest that effective and adaptable local-cloud collaboration can emerge directly from the local model's intrinsic response behavior, without requiring an additional trained router.", "url": "https://wpnews.pro/news/routing-without-training-controllable-ratio-llm-offloading-via-reliability", "canonical_source": "https://arxiv.org/abs/2607.20481", "published_at": "2026-07-24 04:00:00+00:00", "updated_at": "2026-07-24 04:07:52.949427+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "ai-infrastructure"], "entities": ["CARGO", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/routing-without-training-controllable-ratio-llm-offloading-via-reliability", "markdown": "https://wpnews.pro/news/routing-without-training-controllable-ratio-llm-offloading-via-reliability.md", "text": "https://wpnews.pro/news/routing-without-training-controllable-ratio-llm-offloading-via-reliability.txt", "jsonld": "https://wpnews.pro/news/routing-without-training-controllable-ratio-llm-offloading-via-reliability.jsonld"}}