{"slug": "layerrag-bench-a-cross-layer-reliability-benchmark-for-agentic-retrieval", "title": "LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation", "summary": "Researchers introduced LayerRAG-Bench, a cross-layer reliability benchmark for agentic retrieval-augmented generation systems, covering 8 enterprise domains, 240 tasks, 9 fault scenarios, 2 contract modes, and 38,880 live task-level records across nine models from OpenAI, Anthropic, and Gemini. The benchmark found that schema normalization raises schema-drift success from 0.000 to 0.913, but fails to recover from stale evidence, missing tool output, denied permissions, and wrong-session context, and that groundedness-only evaluation produces substantial false positives under stale and wrong-session evidence. The findings support a layer-specific evaluation principle: reliability interventions should be credited for repairing their target layer without being mistaken for universal fixes.", "body_md": "arXiv:2607.27353v1 Announce Type: new\nAbstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer. We introduce LayerRAG-Bench, a controlled cross-layer reliability benchmark with 8 enterprise domains, 240 tasks, 9 fault scenarios, 2 contract modes, and 38,880 live task-level records across nine models from OpenAI, Anthropic, and Gemini. Schema normalization raises schema-drift success from 0.000 to 0.913, but stale evidence, missing tool output, denied permissions, and wrong-session context are not recovered by schema normalization. Groundedness-only evaluation also produces substantial false positives under stale and wrong-session evidence. These results support a layer-specific evaluation principle: a reliability intervention should be credited for repairing its target layer without being mistaken for a universal fix.", "url": "https://wpnews.pro/news/layerrag-bench-a-cross-layer-reliability-benchmark-for-agentic-retrieval", "canonical_source": "https://arxiv.org/abs/2607.27353", "published_at": "2026-07-31 04:00:00+00:00", "updated_at": "2026-07-31 04:40:26.885459+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-safety"], "entities": ["LayerRAG-Bench", "OpenAI", "Anthropic", "Gemini"], "alternates": {"html": "https://wpnews.pro/news/layerrag-bench-a-cross-layer-reliability-benchmark-for-agentic-retrieval", "markdown": "https://wpnews.pro/news/layerrag-bench-a-cross-layer-reliability-benchmark-for-agentic-retrieval.md", "text": "https://wpnews.pro/news/layerrag-bench-a-cross-layer-reliability-benchmark-for-agentic-retrieval.txt", "jsonld": "https://wpnews.pro/news/layerrag-bench-a-cross-layer-reliability-benchmark-for-agentic-retrieval.jsonld"}}