{"slug": "raguard-a-layered-defense-framework-for-retrieval-augmented-generation-systems", "title": "RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning", "summary": "Researchers introduced RAGuard, a layered defense framework for retrieval-augmented generation (RAG) systems against data poisoning, achieving a 0.000% attack success rate in all defended configurations on poisoned Natural Questions at 5–30% poison ratios while keeping Recall@5 within 0.03 of the clean baseline. The framework combines adversarial retriever training with a zero-knowledge inference patch (ZKIP) that requires no poison labels or model internals, costing k+1 generator passes per query.", "body_md": "arXiv:2607.26339v1 Announce Type: new\nAbstract: Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously injected passages that manipulate retrieved evidence. We introduce RAGuard, a layered defense against \\emph{factual} corpus-poisoning attacks on RAG pipelines. The first layer adversarially fine-tunes a dense retriever on synthetic poisoned documents (fabricated facts, contradictions, and reasoning traps), teaching it to downrank malicious passages before generation. The second layer, the Zero-Knowledge Inference Patch ZKIP, is a label-free, black-box filter: for each retrieved document, it performs a leave-one-out decode and scores the document by the semantic shift and output-entropy change that its removal induces. ZKIP requires no poison labels, no ground-truth answers, and no access to model internals; it compares the model's own answers under counterfactual contexts. On poisoned Natural Questions at 5--30\\% poison ratios, adversarial retriever training alone reduces but does not eliminate attack success, while ZKIP drives the measured attack success rate to 0.000 in every defended configuration, keeping Recall@5 within 0.03 of the clean-corpus baseline. Supervised analyses on both Natural Questions and BEIR (NFCorpus) confirm that the counterfactual signals ZKIP relies on carry learnable poison structure. The defense costs $k{+}1$ generator passes per query ($6\\times$ for $k{=}5$); we analyze batching and early-stopping approximations that reduce this overhead. We also show that keyword-preserving poisons leave lexical retrievers such as BM25 essentially unaffected, an observation that delineates the boundary of the threat model. Code, datasets, and evaluation harnesses are released for reproducibility.", "url": "https://wpnews.pro/news/raguard-a-layered-defense-framework-for-retrieval-augmented-generation-systems", "canonical_source": "https://arxiv.org/abs/2607.26339", "published_at": "2026-07-30 04:00:00+00:00", "updated_at": "2026-07-30 04:30:44.592229+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-research"], "entities": ["RAGuard", "arXiv", "Natural Questions", "BEIR", "NFCorpus", "BM25"], "alternates": {"html": "https://wpnews.pro/news/raguard-a-layered-defense-framework-for-retrieval-augmented-generation-systems", "markdown": "https://wpnews.pro/news/raguard-a-layered-defense-framework-for-retrieval-augmented-generation-systems.md", "text": "https://wpnews.pro/news/raguard-a-layered-defense-framework-for-retrieval-augmented-generation-systems.txt", "jsonld": "https://wpnews.pro/news/raguard-a-layered-defense-framework-for-retrieval-augmented-generation-systems.jsonld"}}