{"slug": "your-rag-filter-runs-too-late-build-a-tenant-safe-retriever-in-typescript", "title": "Your RAG Filter Runs Too Late. Build a Tenant-Safe Retriever in TypeScript.", "summary": "A developer published a TypeScript teaching build demonstrating that post-retrieval tenant filtering in RAG pipelines is both a retrieval bug and a data-boundary bug, since globally top-ranked documents from other tenants can crowd out authorized results before any permission check runs. The fixture corpus shows the authorized top two documents (a1, a2) are never returned when the global top-k is filtered after ranking, and the code argues the correct order is verified principal, eligible documents, ranking, reranker, then authorized citations. The repository runs on Node.js 22.18+ using native TypeScript stripping with no API key or model call required.", "body_md": "Your RAG system finds the two highest-scoring documents.\n\nThen it checks whether the user can read them.\n\nBoth belong to another tenant. You drop them and tell the user there is no answer.\n\nBut there were three useful documents the user could read. They never made the candidate list.\n\nThat is a retrieval bug. If those first two documents reached an external reranker before you dropped them, it is also a data-boundary bug.\n\nLet's build the smallest version that makes both mistakes visible.\n\nA tenant filter is not a prompt instruction.\n\nIt is a predicate on the documents the application is allowed to retrieve for this authenticated user. Tenant membership is only one part of it. Groups, document ACLs and permission changes matter too.\n\nThe ordering we want is:\n\n``` php\nverified principal\n  -> eligible documents\n  -> ranking and top-k\n  -> reranker and model\n  -> authorized citations\n```\n\nNot every database implements filtering in the same way. Microsoft's [Azure AI Search documentation](https://learn.microsoft.com/en-us/azure/search/vector-search-filters) distinguishes prefiltering during graph traversal from postfiltering after search. It also documents a separate strict postfilter mode that filters the unfiltered global top-k, which can return zero results even when eligible matches exist. Selective prefilters can increase traversal cost and latency.\n\nThe array example below represents the global-top-k-then-filter mistake. It is not an implementation of Azure's sharded `postFilter` algorithm.\n\nThe query is `refund policy`. Our authenticated fixture user belongs to tenant `acme` and group `support`.\n\n| Document | Tenant | Allowed group | Synthetic score | \n|---|---|---|---|\n| b1 | beta | support | 0.99 | \n| b2 | beta | support | 0.98 | \n| a-hr | acme | hr | 0.95 | \n| a1 | acme | support | 0.91 | \n| a2 | acme | support | 0.89 | \n| a3 | acme | support | 0.80 | \n\nThese scores are invented inputs. There is no embedding model hiding behind them.\n\nThe exact authorized top two are `a1` and `a2`. Our fixture recall is how many of those two the pipeline returns, divided by two.\n\nGlobally, `b1` and `b2` rank first. Filter that list and you get nothing. Increasing k might recover eligible hits in this small corpus, but it does not establish an authorization boundary.\n\nThe [public GitHub repository](https://github.com/bobbyhalljr/tenant-safe-rag) contains the source, fixture data, README, diagrams and expected output.\n\nYou need Node.js 22.18 or newer. [Node documents native TypeScript stripping](https://nodejs.org/api/typescript.html); it executes this erasable syntax without a separate compiler. It does not type-check the file.\n\n```\ngit clone https://github.com/bobbyhalljr/tenant-safe-rag.git\ncd tenant-safe-rag\nnode --experimental-strip-types rag.ts\n```\n\nNo install, API key or model call is required.\n\nHere is the complete teaching build:\n\n``` python\nimport assert from 'node:assert/strict';\n\ntype Principal = { tenantId: string; groups: readonly string[] };\ntype Request = { query: string };\ntype Doc = { id: string; tenantId: string; groups: readonly string[]; text: string };\ntype Hit = Doc & { score: number };\n\n// Synthetic ranking scores for ONE query. This is not an embedding model.\nconst query = 'refund policy';\nconst docs: Doc[] = [\n  { id: 'b1', tenantId: 'beta', groups: ['support'], text: 'Beta private refund policy' },\n  { id: 'b2', tenantId: 'beta', groups: ['support'], text: 'Beta private refund exception' },\n  { id: 'a-hr', tenantId: 'acme', groups: ['hr'], text: 'Acme private employee refund' },\n  { id: 'a1', tenantId: 'acme', groups: ['support'], text: 'Acme refund policy' },\n  { id: 'a2', tenantId: 'acme', groups: ['support'], text: 'Acme refund exception' },\n  { id: 'a3', tenantId: 'acme', groups: ['support'], text: 'Acme refund FAQ' }\n];\nconst scores: Record<string, number> = { b1: 0.99, b2: 0.98, 'a-hr': 0.95, a1: 0.91, a2: 0.89, a3: 0.80 };\nconst principal: Principal = { tenantId: 'acme', groups: ['support'] };\n\n// Production: derive this scope from verified server-side authentication and ACLs.\n// Never trust a tenantId/group supplied by the model or request body.\nfunction allowed(doc: Doc, scope: Principal | null): boolean {\n  return scope !== null && doc.tenantId === scope.tenantId\n    && doc.groups.some(group => scope.groups.includes(group));\n}\nfunction rank(candidates: readonly Doc[], request: Request, k: number): Hit[] {\n  if (!Number.isInteger(k) || k < 1) throw new Error('invalid k');\n  if (request.query !== query) throw new Error('demo supports one fixed query');\n  return candidates.map(doc => ({ ...doc, score: scores[doc.id]! }))\n    .sort((a, b) => b.score - a.score || a.id.localeCompare(b.id)).slice(0, k);\n}\nfunction retrieve(request: Request, scope: Principal | null, k: number, corpus = docs): Hit[] {\n  return rank(corpus.filter(doc => allowed(doc, scope)), request, k);\n}\nconst request: Request = { query };\nconst globalTop = rank(docs, request, 2);\nconst postFiltered = globalTop.filter(doc => allowed(doc, principal));\nassert.deepEqual(postFiltered, []);\nconsole.log('post-filter: returned=0, authorized top-2 recall=0/2');\n\nconst safe = retrieve(request, principal, 2);\nassert.deepEqual(safe.map(doc => doc.id), ['a1', 'a2']);\nassert(safe.every(doc => allowed(doc, principal)));\nconsole.log('pre-filter: ids=a1,a2, authorized top-2 recall=2/2');\n\n// An intentionally unsafe reranker boundary. Both texts leave the trusted store.\nconst rerankerInput = globalTop.map(doc => doc.text);\nassert.equal(rerankerInput.filter(text => text.startsWith('Beta')).length, 2);\nconsole.log('late access check: foreign texts sent to reranker=2');\n\nassert(!safe.some(doc => doc.id === 'a-hr'));\nconsole.log('same-tenant HR document: excluded');\n\nconst forgedRequest = { query, tenantId: 'beta', groups: ['hr'] };\nassert.deepEqual(retrieve(forgedRequest, principal, 2).map(doc => doc.id), ['a1', 'a2']);\nconsole.log('forged request scope: ignored, ids=a1,a2');\n\nassert.deepEqual(retrieve(request, null, 2), []);\nconsole.log('missing authentication: returned=0');\n\nconst revoked = docs.map(doc => doc.id === 'a1' ? { ...doc, groups: ['hr'] } : doc);\nassert.deepEqual(retrieve(request, principal, 2, revoked).map(doc => doc.id), ['a2', 'a3']);\nconsole.log('revoked a1: ids=a2,a3');\n\nassert.throws(() => retrieve(request, principal, 0), /invalid k/);\nconsole.log('invalid k: rejected');\nconsole.log('All 8 scenario checks passed.');\n```\n\nThe important line is deliberately boring:\n\n``` js\nreturn rank(corpus.filter(doc => allowed(doc, scope)), request, k);\n```\n\nThe request contains a query. It does not choose the trusted scope. The fixture even sends forged `tenantId` and `groups` fields; the retrieval function ignores them and continues using the server's principal.\n\nThat only helps if the server principal is trustworthy. This example constructs it locally. It does not implement authentication.\n\nRun under Node.js 22.20.0:\n\n```\npost-filter: returned=0, authorized top-2 recall=0/2\npre-filter: ids=a1,a2, authorized top-2 recall=2/2\nlate access check: foreign texts sent to reranker=2\nsame-tenant HR document: excluded\nforged request scope: ignored, ids=a1,a2\nmissing authentication: returned=0\nrevoked a1: ids=a2,a3\ninvalid k: rejected\nAll 8 scenario checks passed.\n```\n\nThe late-check fixture records two Beta document texts at the reranker boundary. Removing them from the final answer does not remove them from a service that already received them.\n\nThe revocation fixture changes `a1` from support access to HR access. A fresh retrieval returns `a2` and `a3`. A cached answer would need its own invalidation or fresh authorization check.\n\nAn in-memory filter demonstrates ordering. It is not the production design.\n\nBuild the access predicate from verified session context and current permission records. Pass it through the retrieval layer before unauthorized text leaves the trusted store. Apply it consistently to keyword and vector branches, reranking, citations and cached answers.\n\nIf you use PostgreSQL, [row security policies](https://www.postgresql.org/docs/current/ddl-rowsecurity.html) can restrict which rows a role sees. Once row security is enabled, absence of an applicable policy is default deny. Owners normally bypass row security unless forced to obey it; superusers and roles with `BYPASSRLS` also bypass it. Run access tests under the actual application role.\n\nDo not claim that adding an application filter or enabling one database feature proves the whole pipeline is isolated.\n\nFor retrieval evals, keep two measurements separate:\n\nA pipeline that returns zero documents can have zero exposed documents and terrible retrieval quality. A pipeline that recovers relevant text from another tenant can have impressive relevance and a broken access boundary.\n\nEight assertions passed on six documents with hand-set scores.\n\nThat does not measure ANN recall, latency, embedding quality or model correctness. There is no live reranker, ACL database, concurrent permission change, distributed cache or authentication service. It also does not prevent prompt injection inside documents the user is allowed to read.\n\nThe useful result is narrower: filter placement and trusted scope are testable parts of retrieval, before you evaluate the answer.\n\nI'm building Roster around AI employees that do real work. Those employees need useful context, and they need to get it within the access boundary of the person they are working for.", "url": "https://wpnews.pro/news/your-rag-filter-runs-too-late-build-a-tenant-safe-retriever-in-typescript", "canonical_source": "https://dev.to/bobbyhalljr/your-rag-filter-runs-too-late-build-a-tenant-safe-retriever-in-typescript-5elk", "published_at": "2026-10-08 13:17:02+00:00", "updated_at": "2026-10-08 13:20:44.284988+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "developer-tools"], "entities": ["Azure AI Search", "Microsoft", "Node.js", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-rag-filter-runs-too-late-build-a-tenant-safe-retriever-in-typescript", "markdown": "https://wpnews.pro/news/your-rag-filter-runs-too-late-build-a-tenant-safe-retriever-in-typescript.md", "text": "https://wpnews.pro/news/your-rag-filter-runs-too-late-build-a-tenant-safe-retriever-in-typescript.txt", "jsonld": "https://wpnews.pro/news/your-rag-filter-runs-too-late-build-a-tenant-safe-retriever-in-typescript.jsonld"}}