# Can Your AI Agent Search Files the User Cannot Open?

> Source: <https://www.digitalapplied.com/blog/ai-agent-search-document-permissions>
> Published: 2026-10-11 00:00:00+00:00

An AI agent should not search on a user’s behalf with the unrestricted visibility of its ingestion account. Document permissions must survive the journey from the source system into the index, retrieval results, model context, citations and caches. If any stage loses that boundary, an answer can disclose restricted information without ever opening the original file in the user’s browser.

1. 01Map the requester correctlyA connector service account and an end user are different identities.
2. 02Enforce access before contextDo not send unauthorized passages to the model and ask it to hide them.
3. 03Cover derived surfacesSnippets, citations, counts and cached answers can disclose information too.
4. 04Test revocationAn access change must affect indexed and cached copies within a defined policy.

## 01 — Identity boundarySeparate ingestion access from answer *access*

A connector often needs broad read access to collect an organization's documents. That does not mean every employee using the assistant may read everything the connector ingests. Preserve the distinction between the service identity that synchronizes content and the end-user identity whose question is being answered.

In a hypothetical business, the connector can ingest both a public staff handbook and a restricted acquisition folder. A general employee asks about upcoming organizational changes. A relevant passage in the restricted folder remains unauthorized even if it would produce a better answer. Relevance and permission are independent tests, and permission must win before the passage reaches the answer model.

Our [agent data-access guide](https://www.digitalapplied.com/blog/ai-agent-data-access-permissions-guide) covers row-level access in business databases. Document search adds a different problem: permissions must be copied or resolved across connectors, indexes and derived outputs. A secure database policy elsewhere in the application does not automatically protect a separate search index.

Include the user interface in the identity trace. A chat may show the correct signed-in user while a background search job runs under a shared administrative credential. The visible account badge is not evidence about the server-side request context. Record how the authenticated principal is passed to the retrieval service, and test that a model-generated tool argument cannot substitute another user or omit the tenant boundary to broaden the result set.

##### Connector account

Reads approved source material for synchronization under its own access scope.

##### End user and groups

Determines which indexed documents may contribute to this answer.

##### Snippets and answers

Must preserve the authorization boundary after retrieval and caching.

## 02 — Integration scopeVerify the connector security model *explicitly*

Check the exact connector, version, deployment and subscription configuration. A successful content sync proves that documents were copied; it does not prove that user and group permissions were copied correctly or enforced on every search. Some integrations support document-level security only under specific conditions, while others require a separate implementation.

Elastic's document-level security documentation for connectors, checked October 11, 2026, describes a connector-specific permission mechanism and notes availability and subscription constraints. It also labels that implementation beta. Those qualifications matter: the documentation is not a claim that every Elastic product, connector or query path applies the same access controls automatically.

Write down where the permission decision happens and what happens when permission data is missing. A recommended design is to exclude uncertain documents from user-facing results until the mapping is resolved. Do not present that recommendation as a vendor default unless the exact integration documents it. Unknown authorization should not silently become public visibility.

Treat connector upgrades as changes to the access path. A new connector version can alter identity mapping, supported source permissions or synchronization behavior even when document ingestion continues to succeed. Keep a small permission test suite attached to the integration and run it after relevant changes. The absence of sync errors is not an authorization test, because an over-broad result can look like a perfectly healthy search response.

Record both the connector’s capability and your configured behavior. “Supports document security” does not prove that a particular index and query path are enforcing it.

## 03 — Principal mappingMap identities without trusting the *prompt*

The application should establish the requester and relevant group memberships through its trusted authentication and authorization path. The model may describe the user's question, but it should not invent the user identifier or choose a broader group to make the search succeed. Keep those values outside prompt-controlled arguments where possible.

Map source-system identities to the identities used by the search application deliberately. A matching email string may be insufficient when guests, aliases, renamed accounts or multiple organizations are involved. Preserve tenant scope and use stable identifiers provided by the relevant identity systems. Test accounts with similar display names to confirm that presentation fields do not control authorization.

Group membership changes are part of the data flow. Decide how current membership is resolved and how stale mappings are invalidated. A user removed from a confidential group should not keep access indefinitely because the connector last synchronized the group weeks ago. The required revocation window is an operational requirement that the implementation must demonstrate.

Explicit sharing links and inherited folder permissions deserve their own cases. Source systems can distinguish individual grants, group grants, organization-wide visibility and link-based access. A connector that flattens these into one broad field may lose an important condition. Verify the exact source semantics your integration supports, and exclude unsupported permission patterns from the initial scope instead of pretending that every access mechanism maps cleanly into a single list of users.

- Take user and tenant context from trusted application state.
- Map stable source principals to search principals explicitly.
- Test guest accounts, aliases and changed group membership.

## 04 — Context boundaryFilter before retrieval results reach the *model*

Apply authorization at the retrieval boundary so unauthorized passages cannot enter the model context. Telling a model to ignore private material after supplying it is not equivalent. The content may influence the answer, appear in a summary or be retained in logs even if the final response does not quote it directly.

Keep authorization enforcement on every query path, including semantic search, lexical search and any direct document fetch used after ranking. A system that filters the first search but fetches a citation through a broad service credential can reintroduce the same disclosure later. The reranker must also receive only material the request is permitted to use.

A hypothetical assistant may search a broad internal index and then ask a second tool for a full document. Both calls need the same trusted access context. Our [retrieval diagnosis guide](https://www.digitalapplied.com/blog/reranker-vs-reembedding-search-diagnosis) separates search stages for quality testing; use that same stage-by-stage trace to verify the access boundary rather than assuming one filter protects the whole chain.

Check direct lookup tools as well as search tools. An agent may receive a document identifier in conversation and call a fetch endpoint without performing search first. If that endpoint relies only on possession of the identifier, it can bypass a well-protected result list. Every route that returns document content needs an authorization decision for the requester, including retries, citation expansion and background summarization jobs that operate after the initial response.

Inspect the actual documents and passages delivered to the model under a restricted test user. A clean final answer alone does not prove that unauthorized context was excluded.

## 05 — Derived disclosureProtect snippets, citations and aggregate *clues*

A result can reveal information without showing the full document. Titles, snippets, file paths, authors, timestamps and counts may disclose a confidential project or relationship. Decide which metadata the user may see, and apply that rule consistently to search results and the assistant's citations.

An answer that says it found several restricted documents about an unannounced project can disclose the project's existence even if it refuses to summarize them. A safer response describes the evidence available to the requester without exposing unauthorized matches. Treat this as an application design requirement rather than relying on the model to recognize every sensitive implication.

Citation links need a second check at access time. A user who could read a file when the answer was generated may lose permission later. The stored answer and the destination link are separate surfaces: the link should enforce current source access, while the application needs a policy for retaining or invalidating previously generated protected text.

Aggregations require an explicit policy because they can summarize information the user cannot inspect individually. A count of confidential documents or a list of hidden project tags may reveal sensitive activity without exposing a passage. Compute user-visible facets and counts over the authorized set, or omit them when the system cannot safely provide them. Do not let an agent infer and report hidden totals from internal diagnostic fields returned by a broad search service.

| Illustrative disclosure surfaces to inspect. Actual controls depend on the connector and application architecture. |  |  | 
|---|---|---|
| Surface | Possible disclosure | Required check | 
|---|---|---|
| Snippet | Restricted passage appears in preview | Filter before preview generation | 
| Title or path | Confidential project becomes visible | Authorize metadata as well as content | 
| Count | Presence of hidden records is inferred | Use permission-scoped aggregation | 
| Citation | Broad fetch bypasses search restriction | Recheck current access on retrieval | 
| Stored answer | Old permissions survive in cached prose | Apply scoped storage and invalidation policy | 

## 06 — Cache isolationTreat caches as protected *copies*

A cache keyed only by the question can return one user's authorized answer to another user who asks the same thing. Include the effective authorization scope or a validated equivalent in the cache design, and confirm that the stored result is still valid under current permissions before serving it. A user identifier alone may not capture a changed group membership or policy version.

The same issue applies to retrieved passages, summaries and conversation memory. If a summary was built from a restricted document, removing the original file from a later search does not remove the protected facts already copied into that summary. Track provenance and define how permission changes affect derived material.

Our [agent-memory deletion guide](https://www.digitalapplied.com/blog/deleting-ai-agent-memory-copies) explains why removing one copy is not necessarily complete removal. For search authorization, the practical question is which copies can still answer a request after access changes. Inventory those copies and test them rather than assuming the source system's revocation automatically reaches every cache.

Conversation history is another cache surface. If a user loses access after an earlier answer, a later prompt may ask the assistant to repeat or elaborate on the protected material already in context. Decide how the application handles that history under the organization's policy and technical capabilities. Revoking a source permission alone cannot make text already disclosed disappear, but it should not be mistaken for complete control of future derived responses.

- Keep authorization scope in retrieval and answer-cache decisions.
- Track which documents contributed to stored summaries.
- Invalidate or reauthorize derived results when access changes.

## 07 — Negative testsTest two users and a permission *change*

Create synthetic documents with public, team-only and individual-only access, then use test users whose permissions overlap without being identical. Ask the same questions under each identity. Inspect candidate results, model context, final answers, citations and logs. The expected differences should follow the permission policy, not the wording of the prompt.

Next, remove a user's access and repeat the query through warm caches and an existing conversation. Test a newly added document whose permissions have not yet synchronized and an identity-mapping failure. State the expected safe behavior in advance. These are proposed tests; this article does not report a completed security audit or certify any connector.

A useful failure record identifies the earliest stage that admitted unauthorized information and every downstream copy affected. Fix that boundary, then repeat the case with both allowed and denied users. Checking only the denied case can accidentally produce a system that blocks everyone, which is secure in a narrow sense but does not satisfy the business task.

Test failures that occur between synchronization steps. A document may arrive before its access metadata, or a group update may arrive before a document update. Define whether the integration stages these changes atomically or temporarily excludes uncertain material. The important invariant is that an incomplete synchronization does not become a window of broader access. A test that starts only after every job has finished may miss this transient but consequential state.

A permission test needs a known allowed result and a known denied result. Verify both so the repair preserves legitimate access while removing disclosure.

## 08 — Ongoing controlMake access freshness an operating *responsibility*

Assign responsibility for connector failures, group-sync delays and permission mismatches. A search service that continues operating while its authorization data is stale needs a defined degraded behavior, not a quiet fallback to broad access. The appropriate response may restrict results, pause a connector or ask the user to open the source directly.

Record the configured revocation behavior and test it when connectors or identity systems change. Keep audit evidence about denied requests without unnecessarily retaining the protected content itself. Operators need enough context to diagnose the boundary, but debugging should not create another ungoverned copy of confidential documents.

Our [AI transformation service](https://www.digitalapplied.com/services/ai-transformation) helps teams connect useful knowledge search with explicit data boundaries. The desired outcome is straightforward: an agent can find and explain the information the requester is entitled to use, while the connector's broader ingestion privileges remain invisible to the answer path.

Keep the operating record understandable to the security and content owners. They should be able to identify which source systems are connected, which permission patterns are supported and how quickly revocation is expected to take effect. If a requirement exceeds what the integration can demonstrate, narrow the scope or change the design. A clearly limited knowledge assistant is more useful than an apparently universal one whose authorization boundary nobody can explain.

- Name an owner for permission-sync failures and stale mappings.
- Define degraded behavior before the connector loses authorization freshness.
- Retest access paths when caches, connectors or identity mappings change.

### Test what reaches the model

Trace a restricted document from the source through retrieval, citations and stored answers. Confirm that the requester’s current permissions govern every surface.

A working connector is only the start. The access boundary is complete when unauthorized content stays out of the answer path and revocation reaches its derived copies.
