Exposing one internal tool over the Model Context Protocol is a weekend project. Exposing ninety of them to humans and agents, without turning the MCP endpoint into a credential dump and the audit log into a fire hazard, is a platform problem.
This post is the short version of how I'd design that platform. The one-line answer: treat the MCP server as a tool broker in the same family as your LLM gateway. Same identity rules (no static keys), same metering shape (every invocation is an event), same audit discipline. Plus two things the LLM path doesn't need: an envelope for stored third-party credentials, and a PII masking step for what flows back.
#
- Organize the catalog by category, not by tool
At 90+ tools, the first failures are organizational, not performance. A few rules keep the catalog sane:
Naming is a contract:{category}.{verb}_{object} , e.g.issues.create_ticket ,mail.send_draft ,ci.get_pipeline_logs . Renaming a tool is a breaking change with a deprecation window, not a silent edit. #
One tool = one unit of permission, audit, and test. A "do everything in the issue tracker" tool is a permission you can't grant safely, an audit row you can't read, and a test you can't write. Split until each tool is all three. #
Version in the schema, not the name (issues.create_ticket@v2 ), and serve all live versions during the deprecation window. #
Grant tool groups, not the universe. The server barely notices 94 schemas. Theclient does: an agent that loads every schema can burn 30k+ tokens of context on tool definitions alone. Clients request a group (ci ,mail ), and the group is the grant unit.
#
- Tool design patterns that keep agents (and auditors) sane
A tool definition worth reviewing looks roughly like this:
The patterns behind it:
Idempotency keys on retryable writes. A retried create with the same key returns the same ticket, not a duplicate. Make it a registration-time lint for write tools. #
Errors are part of the contract. Normalize upstream failures into a small taxonomy with aretryable flag (andRetry-After when you have it). Agents retry on the flag; humans read the code. #
Reads are cheap, writes are loud. Writes need a minimum role and a scope (project, space, folder). The audit event for a write carries the full (masked) input; for a read, just the shape. #
Bound every list. Defaultlimit 50, server cap 200, cursor pagination. An unbounded list tool is a context-window bomb. #
Return what was asked, not what you know. Enrichment ("related tickets") is a second tool the agent can choose to call. Tools that over-return are tools agents can't plan around.
#
- Auth: OAuth 2.1 + PKCE, with a few non-negotiables
The happy path is simple: the client runs the authorization-code flow with PKCE against your identity platform, gets a short-lived access token, and presents it on the MCP endpoint. The server verifies it, resolves the principal, and checks tool grants. What makes it actually safe:
PKCE for every client type , including IDEs and agent runtimes. MCP clients usually can't hold a secret. #
Resource indicators. The token's audience is the MCP server. A valid token minted for some other service gets rejected. #
Dynamic client registration is not a grant. DCR keeps "bring your own agent runtime" possible (redirect URI allowlist, auth method checks), but tool access still flows through an explicit grant. #
One token-time posture for the org. For example: 15-minute access tokens for interactive clients, 1 hour for workloads, no refresh tokens for agents (they re-authenticate through their workload identity). #
Scopes per tool group (mcp:ci:read ,mcp:issues:write ), so the consent screen is readable instead of a wall of 94 strings. #
The client's token never travels to the SaaS. The broker calls downstream systems with its own credential and carries the principal in-band for audit and metering. Where a downstream system has a clean per-user delegation model, use that instead; it's the stronger design.
#
- Stored secrets: envelope encryption, and why it isn't the control
The broker has to store credentials for the systems behind it. Do it with a standard envelope:
Per-category DEKs keep the blast radius at one category. Per-user tokens get a per-user DEK under the category DEK. #
Rotation is cheap by design: the KMS rotates the KEK; a scheduled job re-wraps DEKs. Neither touches the secret ciphertexts. #
But encryption at rest only answers "a DB leak is not a credential leak." The real control is that no role (except break-glass) canread a secret. The brokeruses it in-process for one call, zeroes the buffer, and the audit event records the call and the principal, never the credential. A portal should show "credential present, last rotated N days ago", nothing more.
#
- Mask on the way back
The LLM gateway can stay payload-blind. The MCP path can't: the tool result is the product. So results pass through a PII masking step before they reach the client and before anything is stored, and the per-tool pii field in the catalog says which inputs and outputs need it. This is also why write audits store masked inputs, not raw ones.
#
The checklist version
- Broker, not proxy: no static keys, every invocation metered and audited
- Category naming, one-tool-one-permission, schema versioning, group grants
- Idempotency, error taxonomy, bounded lists, no over-returning
- OAuth 2.1 + PKCE, audience-bound tokens, DCR ≠ grant, group scopes
- Envelope encryption with per-category DEKs, secrets used and never read
- PII masking on results and on stored audit inputs
This is a condensed version of Chapter 5 of my book The AI Gateway Playbook, which also covers the LLM gateway itself: model registry, keyless auth and RBAC, quotas and budgets, a RAG team assistant, and operations runbooks. Free sample on the Leanpub page.
Disclosure: this article and the book were drafted with AI assistance from my own experience designing and running this kind of platform. All examples are generic.