{"slug": "sage-slo-aware-adaptive-retrieval-for-production-rag-systems", "title": "SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems", "summary": "A new paper on arXiv (2608.08237v1) proposes SAGE, a learned SLO-aware adaptive retrieval policy for production RAG systems that dynamically selects the number of passages per query. On Natural Questions, under a 5s P95 latency SLO, SAGE achieves 95% SLO compliance versus 30% for the best static baseline (k=20), reduces P95 latency by 36% and retrieval cost by 51% with only 2 percentage points Exact Match (EM) loss. A single policy trained on Natural Questions generalizes across HotpotQA, UnSeenTimeQA, and four LLM families (Llama, Qwen, Mistral, Gemma), consistently yielding +45-52 point SLO improvements without quality degradation.", "body_md": "arXiv:2608.08237v1 Announce Type: new\nAbstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cost. However, standard retrieval pipelines rely on fixed retrieval budgets that ignore query difficulty, over-retrieving for easy queries and under-serving hard ones, forcing operators to trade answer quality against SLO compliance. This paper proposes SAGE, a learned SLO-aware adaptive retrieval policy that dynamically selects the number of passages k per query. SAGE uses lightweight features derived from initial retrieval (e.g., score distributions, rank gaps, lexical signals) and is trained offline via imitation learning from an oracle that approximates optimal latency-quality trade-offs. At inference, it adds no LLM calls and minimal overhead. On Natural Questions, under a 5s P95 latency SLO, SAGE achieves 95% SLO compliance versus 30% for the best static baseline (k=20), reduces P95 latency by 36% and retrieval cost by 51% with only 2 percentage points Exact Match (EM) loss. A single policy trained on Natural Questions generalizes across HotpotQA, UnSeenTimeQA, and four LLM families (Llama, Qwen, Mistral, Gemma), consistently yielding +45-52 point SLO improvements without quality degradation.", "url": "https://wpnews.pro/news/sage-slo-aware-adaptive-retrieval-for-production-rag-systems", "canonical_source": "https://www.machinebrief.com/news/sage-slo-aware-adaptive-retrieval-for-production-rag-systems-yeg2", "published_at": "2026-08-11 04:00:00+00:00", "updated_at": "2026-08-11 05:13:26.701712+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-infrastructure"], "entities": ["SAGE", "arXiv", "Natural Questions", "HotpotQA", "UnSeenTimeQA", "Llama", "Qwen", "Mistral"], "alternates": {"html": "https://wpnews.pro/news/sage-slo-aware-adaptive-retrieval-for-production-rag-systems", "markdown": "https://wpnews.pro/news/sage-slo-aware-adaptive-retrieval-for-production-rag-systems.md", "text": "https://wpnews.pro/news/sage-slo-aware-adaptive-retrieval-for-production-rag-systems.txt", "jsonld": "https://wpnews.pro/news/sage-slo-aware-adaptive-retrieval-for-production-rag-systems.jsonld"}}