# The Token Vault: Secure OAuth Delegation and Identity for AI Agents

> Source: <https://dev.to/dks/the-token-vault-secure-oauth-delegation-and-identity-for-ai-agents-m68>
> Published: 2026-10-10 14:14:20+00:00

When building an enterprise AI agent, you will quickly hit a critical security dilemma:

A user asks an agent to inspect a document in their cloud storage drive and update an issue in their project management tracker. To perform these tasks, the agent must call external APIs on behalf of that specific user.

How do you give the agent the necessary authorization tokens to execute these actions?

In early prototypes, developers often inject the user's raw Bearer token directly into the LLM system prompt or into the tool arguments.

**Putting raw credentials into an LLM prompt or tool argument is an unacceptable security hazard.**

If an attacker executes a prompt injection attack (such as hiding instructions inside a shared document that direct the model to output its system prompt), the model can leak the user's secret keys into the chat window. Furthermore, prompts and completions are logged across monitoring tools and vector caches, exposing sensitive credentials in plain text.

To solve this, we architected the **Token Vault**. It enables secure On-Behalf-Of (OBO) delegation so agents can run tools without ever seeing the user's raw secrets.

## 
  
  
  The Idea: Opaque Identity References

The core idea of the Token Vault is to decouple the **agent's reasoning engine** from the **credential layer**.

Instead of giving credentials to the language model, all secrets live in an isolated, encrypted vault service:

1. 
**User Consent & Vault Storage** : When a user authorizes an integration, standard OAuth authorization code flows occur between the user's browser and the Token Vault. The vault encrypts and stores the access and refresh tokens.
2. 
**Opaque Token References** : The vault generates a randomized, opaque reference identifier (a non-sensitive handle) representing that active connection.
3. 
**Prompt-Free Execution** : The agent only ever sees and passes the opaque handle. When the agent calls a tool, our backend tool broker intercepts the call, validates the user's session, exchanges the opaque handle for a short-lived token inside the secure network perimeter, executes the API call, and returns the result.

The language model, the prompt context, and the observability logs never touch a real secret key.

## 
  
  
  How It Worked Well

1. 
**Zero Credential Leakage via Prompt Injections** : Even if an attacker completely compromises an agent's reasoning chain through indirect prompt injection, the model cannot leak the token because it does not have it. The model only knows an opaque reference handle that is useless outside of the authenticated session.
2. 
**Automated Refresh Management** : Third-party access tokens typically expire in one hour. Because the Token Vault manages the OAuth lifecycle, it automatically uses the stored refresh token to obtain a fresh access token without interrupting the user's conversation or requiring re-authentication.
3. 
**Strict Caller-Identity Binding** : The Token Vault enforces cryptographic ownership checks. If User B inspects an agent's session and attempts to invoke a tool using User A's reference handle, the vault rejects the exchange with a permission error.
4. 
**Clean Observability Compliance** : OpenTelemetry traces and audit logs capture full tool execution details without triggering data-loss-prevention alarms or exposing user credentials in telemetry dashboards.

## 
  
  
  What to Watch Out For

1. 
**Granular Scope Containment** : When users connect external accounts, request the minimum necessary OAuth scopes. If an agent only needs to read documents, never request write or admin permissions. If an agent needs write access, implement step-up confirmation so users explicitly approve the elevated scope.
2. 
**Session and Token Revocation** : If a user logs out of your enterprise platform or disconnects a third-party service, your Token Vault must immediately invalidate the cached credentials and revoke the upstream OAuth grant. Disconnecting one tool should never leave orphaned active tokens in storage.
3. 
**Rate Limits on Refresh Flows** : Third-party identity providers place rate limits on token refresh endpoints. If hundreds of concurrent agent runs attempt to refresh tokens simultaneously for the same user, you can trigger provider throttling. Implement distributed caching and locking around refresh token exchanges.
4. 
**Handling Disconnected State Gracefully** : When an upstream token expires and cannot be refreshed (for example, if the user changed their enterprise password), the tool broker must return a clean, structured re-authentication prompt to the chat client rather than crashing the agent with an unhandled exception.
