5 RAG Mistakes That Leak Private Docs Into Chat Answers A technical post identifies five common retrieval-augmented generation (RAG) mistakes that leak private documents into chatbot answers, including storing all documents in one vector store without tenant or audience metadata and applying permission filters after retrieval instead of inside VectorSearchOptions.Filter. The post recommends tagging every record with indexed TenantId, Audience, and Quarantined fields, deriving CallerContext from authenticated API key or JWT claims rather than request JSON, and setting a ScoreThreshold so zero hits return NotFound and skip the model. The full working project is built on ASP.NET Core / .NET 10 with Microsoft.Extensions.VectorData and Microsoft.Extensions.AI. 5 RAG Mistakes That Leak Private Docs Into Chat Answers Someone asks, "What does a Senior Engineer earn here?" Your chatbot answers. With citations. Nobody hacked anything. Similarity search found the HR salary chunk because that chunk lived in the same index as the travel policy. I keep seeing the same leaks in "docs chatbot" demos. The full working pro Someone asks, "What does a Senior Engineer earn here?" Your chatbot answers. With citations. Nobody hacked anything. Similarity search found the HR salary chunk because that chunk lived in the same index as the travel policy. I keep seeing the same leaks in "docs chatbot" demos. The full working project ASP.NET Core / .NET 10, Microsoft.Extensions.VectorData + Microsoft.Extensions.AI, tests, offline mode is in Tech Skill Builder. This post is the short version: five mistakes that turn RAG into an accidental data-access path, and the shape that closes them. If a chunk does not know who may read it, no query can enforce that. All documents share one vector store; cosine similarity does not care about roles. Fix: put tenant and audience on every record, and mark them indexed so providers can filter. public sealed class KnowledgeChunk { VectorStoreKey public required string Key { get; set; } VectorStoreData IsIndexed = true public required string TenantId { get; set; } VectorStoreData IsIndexed = true public required string Audience { get; set; } // everyone | hr VectorStoreData IsIndexed = true public bool Quarantined { get; set; } // Title, Section, Text, Fingerprint, Embedding … // full implementation in the complete project } Admin rights to re-index are not the same as rights to read salaries. Keep those roles separate. Retrieve five chunks, then drop the ones the caller should not see. Ranking already saw the private text. You also end up with empty contexts more often than you expect. Fix: put the permission predicate inside VectorSearchOptions.Filter so the store applies it during search, not in your app afterward. var options = new VectorSearchOptions