Your RAG system finds the two highest-scoring documents.
Then it checks whether the user can read them.
Both belong to another tenant. You drop them and tell the user there is no answer.
But there were three useful documents the user could read. They never made the candidate list.
That is a retrieval bug. If those first two documents reached an external reranker before you dropped them, it is also a data-boundary bug.
Let's build the smallest version that makes both mistakes visible.
A tenant filter is not a prompt instruction.
It is a predicate on the documents the application is allowed to retrieve for this authenticated user. Tenant membership is only one part of it. Groups, document ACLs and permission changes matter too.
The ordering we want is:
verified principal
-> eligible documents
-> ranking and top-k
-> reranker and model
-> authorized citations
Not every database implements filtering in the same way. Microsoft's Azure AI Search documentation distinguishes prefiltering during graph traversal from postfiltering after search. It also documents a separate strict postfilter mode that filters the unfiltered global top-k, which can return zero results even when eligible matches exist. Selective prefilters can increase traversal cost and latency.
The array example below represents the global-top-k-then-filter mistake. It is not an implementation of Azure's sharded postFilter algorithm.
The query is refund policy. Our authenticated fixture user belongs to tenant acme and group support.
| Document | Tenant | Allowed group | Synthetic score |
|---|---|---|---|
| b1 | beta | support | 0.99 |
| b2 | beta | support | 0.98 |
| a-hr | acme | hr | 0.95 |
| a1 | acme | support | 0.91 |
| a2 | acme | support | 0.89 |
| a3 | acme | support | 0.80 |
These scores are invented inputs. There is no embedding model hiding behind them.
The exact authorized top two are a1 and a2. Our fixture recall is how many of those two the pipeline returns, divided by two.
Globally, b1 and b2 rank first. Filter that list and you get nothing. Increasing k might recover eligible hits in this small corpus, but it does not establish an authorization boundary.
The public GitHub repository contains the source, fixture data, README, diagrams and expected output.
You need Node.js 22.18 or newer. Node documents native TypeScript stripping; it executes this erasable syntax without a separate compiler. It does not type-check the file.
git clone https://github.com/bobbyhalljr/tenant-safe-rag.git
cd tenant-safe-rag
node --experimental-strip-types rag.ts
No install, API key or model call is required.
Here is the complete teaching build:
import assert from 'node:assert/strict';
type Principal = { tenantId: string; groups: readonly string[] };
type Request = { query: string };
type Doc = { id: string; tenantId: string; groups: readonly string[]; text: string };
type Hit = Doc & { score: number };
// Synthetic ranking scores for ONE query. This is not an embedding model.
const query = 'refund policy';
const docs: Doc[] = [
{ id: 'b1', tenantId: 'beta', groups: ['support'], text: 'Beta private refund policy' },
{ id: 'b2', tenantId: 'beta', groups: ['support'], text: 'Beta private refund exception' },
{ id: 'a-hr', tenantId: 'acme', groups: ['hr'], text: 'Acme private employee refund' },
{ id: 'a1', tenantId: 'acme', groups: ['support'], text: 'Acme refund policy' },
{ id: 'a2', tenantId: 'acme', groups: ['support'], text: 'Acme refund exception' },
{ id: 'a3', tenantId: 'acme', groups: ['support'], text: 'Acme refund FAQ' }
];
const scores: Record<string, number> = { b1: 0.99, b2: 0.98, 'a-hr': 0.95, a1: 0.91, a2: 0.89, a3: 0.80 };
const principal: Principal = { tenantId: 'acme', groups: ['support'] };
// Production: derive this scope from verified server-side authentication and ACLs.
// Never trust a tenantId/group supplied by the model or request body.
function allowed(doc: Doc, scope: Principal | null): boolean {
return scope !== null && doc.tenantId === scope.tenantId
&& doc.groups.some(group => scope.groups.includes(group));
}
function rank(candidates: readonly Doc[], request: Request, k: number): Hit[] {
if (!Number.isInteger(k) || k < 1) throw new Error('invalid k');
if (request.query !== query) throw new Error('demo supports one fixed query');
return candidates.map(doc => ({ ...doc, score: scores[doc.id]! }))
.sort((a, b) => b.score - a.score || a.id.localeCompare(b.id)).slice(0, k);
}
function retrieve(request: Request, scope: Principal | null, k: number, corpus = docs): Hit[] {
return rank(corpus.filter(doc => allowed(doc, scope)), request, k);
}
const request: Request = { query };
const globalTop = rank(docs, request, 2);
const postFiltered = globalTop.filter(doc => allowed(doc, principal));
assert.deepEqual(postFiltered, []);
console.log('post-filter: returned=0, authorized top-2 recall=0/2');
const safe = retrieve(request, principal, 2);
assert.deepEqual(safe.map(doc => doc.id), ['a1', 'a2']);
assert(safe.every(doc => allowed(doc, principal)));
console.log('pre-filter: ids=a1,a2, authorized top-2 recall=2/2');
// An intentionally unsafe reranker boundary. Both texts leave the trusted store.
const rerankerInput = globalTop.map(doc => doc.text);
assert.equal(rerankerInput.filter(text => text.startsWith('Beta')).length, 2);
console.log('late access check: foreign texts sent to reranker=2');
assert(!safe.some(doc => doc.id === 'a-hr'));
console.log('same-tenant HR document: excluded');
const forgedRequest = { query, tenantId: 'beta', groups: ['hr'] };
assert.deepEqual(retrieve(forgedRequest, principal, 2).map(doc => doc.id), ['a1', 'a2']);
console.log('forged request scope: ignored, ids=a1,a2');
assert.deepEqual(retrieve(request, null, 2), []);
console.log('missing authentication: returned=0');
const revoked = docs.map(doc => doc.id === 'a1' ? { ...doc, groups: ['hr'] } : doc);
assert.deepEqual(retrieve(request, principal, 2, revoked).map(doc => doc.id), ['a2', 'a3']);
console.log('revoked a1: ids=a2,a3');
assert.throws(() => retrieve(request, principal, 0), /invalid k/);
console.log('invalid k: rejected');
console.log('All 8 scenario checks passed.');
The important line is deliberately boring:
return rank(corpus.filter(doc => allowed(doc, scope)), request, k);
The request contains a query. It does not choose the trusted scope. The fixture even sends forged tenantId and groups fields; the retrieval function ignores them and continues using the server's principal.
That only helps if the server principal is trustworthy. This example constructs it locally. It does not implement authentication.
Run under Node.js 22.20.0:
post-filter: returned=0, authorized top-2 recall=0/2
pre-filter: ids=a1,a2, authorized top-2 recall=2/2
late access check: foreign texts sent to reranker=2
same-tenant HR document: excluded
forged request scope: ignored, ids=a1,a2
missing authentication: returned=0
revoked a1: ids=a2,a3
invalid k: rejected
All 8 scenario checks passed.
The late-check fixture records two Beta document texts at the reranker boundary. Removing them from the final answer does not remove them from a service that already received them.
The revocation fixture changes a1 from support access to HR access. A fresh retrieval returns a2 and a3. A cached answer would need its own invalidation or fresh authorization check.
An in-memory filter demonstrates ordering. It is not the production design.
Build the access predicate from verified session context and current permission records. Pass it through the retrieval layer before unauthorized text leaves the trusted store. Apply it consistently to keyword and vector branches, reranking, citations and cached answers.
If you use PostgreSQL, row security policies can restrict which rows a role sees. Once row security is enabled, absence of an applicable policy is default deny. Owners normally bypass row security unless forced to obey it; superusers and roles with BYPASSRLS also bypass it. Run access tests under the actual application role.
Do not claim that adding an application filter or enabling one database feature proves the whole pipeline is isolated.
For retrieval evals, keep two measurements separate:
A pipeline that returns zero documents can have zero exposed documents and terrible retrieval quality. A pipeline that recovers relevant text from another tenant can have impressive relevance and a broken access boundary.
Eight assertions passed on six documents with hand-set scores.
That does not measure ANN recall, latency, embedding quality or model correctness. There is no live reranker, ACL database, concurrent permission change, distributed cache or authentication service. It also does not prevent prompt injection inside documents the user is allowed to read.
The useful result is narrower: filter placement and trusted scope are testable parts of retrieval, before you evaluate the answer.
I'm building Roster around AI employees that do real work. Those employees need useful context, and they need to get it within the access boundary of the person they are working for.