Uber Ships MCP Gateway at Production Scale — 800 Servers and 5,000 Tools Behind a Single Control Surface Uber is running more than 800 Model Context Protocol servers and 5,000 individual tools in production behind a single execution-layer gateway, the company detailed in an October 2026 engineering blog by a team of eight engineers. The dual-plane design pairs an MCP Registry control plane with a Proxy Gateway data plane that routes MCP calls as HTTP, gRPC, or TChannel traffic through Uber's Muttley service mesh, with every discovered tool disabled until its owning team reviews and enables it. A Cadence-powered AutoCrawler scans more than 10,000 service IDLs to auto-generate agent-friendly tool descriptions and register them disabled by default, while Omni MCP and Response Projection cut context-window and response-token bloat. Uber is currently operating more than 800 Model Context Protocol servers and 5,000 individual tools in production. The details, published in an October 2026 engineering blog https://www.uber.com/us/en/blog/designing-mcp-gateway/ by a team of eight engineers including a Distinguished Engineer and Principal Engineer, represent the first publicly documented Fortune 500 MCP deployment at this scale. The execution-layer gateway — the architectural pattern we have tracked from thesis https://forkast.news/the-harness-pattern-is-no-longer-a-thesis-it-is-the-infrastructure-architecture/ to security control surface https://forkast.news/the-execution-layer-gateway-is-where-enterprise-ai-security-actually-lives/ to commodity SKU https://forkast.news/agent-infrastructure-is-becoming-a-commodity-sku-and-the-moat-is-moving-upstream/ — now has a production reference architecture from one of the world’s largest microservice environments. The gateway follows a dual-plane design: an MCP Registry serves as the control plane, maintaining a catalog of servers, tools, ownership, and enablement state; a Proxy Gateway forms the data plane, translating MCP protocol calls into HTTP, gRPC, or TChannel requests and routing them through Muttley, Uber’s service mesh. Every discovered tool starts in a disabled state and must be explicitly reviewed and enabled by the owning team before it becomes accessible to agents. This design treats discovery as a starting point, not an authorization — a distinction that matters when thousands of services are involved. The most notable engineering decision is AutoCrawler, a Cadence-powered distributed workflow that continuously scans Uber’s IDL registry for new services, APIs, and schema changes. For each discovered entity, AutoCrawler uses LLMs to generate agent-friendly tool descriptions from the extracted Protobuf or Thrift definitions, translates schemas into MCP-compatible JSON-RPC format, and registers the tools in a disabled-by-default state. This automation handles the discovery problem that manual registration cannot solve at Uber’s scale — more than 10,000 service IDLs generating thousands of tool definitions without any service team needing to author an MCP server. Two extensions address the operational friction of scaling beyond a handful of servers. Omni MCP acts as a single proxy that enables incremental discovery across the entire fleet, exposing four tools — discover server, discover tools, get tool schema, and invoke tool — that let agents find and use tools without loading every server’s full schema into context. Response Projection applies a GraphQL-like field selection pattern to MCP tool responses, trimming the payload to only the fields the agent actually needs. Together, these mechanisms attack the two scaling constraints that break enterprise MCP: context-window exhaustion and response-token bloat. For developers, Uber has standardized on Code Mode through its internal CLI, aifx https://www.uber.com/us/en/blog/designing-mcp-gateway/ , which lets coding agents invoke MCP tools through the gateway without installing any local MCP server. The CLI exposes list, search, and call commands that route through the gateway, and agents can chain these to write output to files for selective grep — loading only what they need into context. Code Mode is now the company default for MCP tool use in coding agents, including integrations with Claude Code and Cursor. The security architecture, detailed in a separate May 2026 blog post https://www.uber.com/hn/en/blog/solving-the-agent-identity-crisis/page/22/ , connects the gateway to an AI Agent Mesh and a Security Token Service that issues short-lived, single-hop JWT tokens with full actor chain propagation. Every agent action carries a verifiable lineage — from the originating user through each intermediate agent to the tool invocation — with P99 latency under 40 milliseconds for token exchange. The gateway enforces charter policies at the tool level, applies PII redaction automatically, and tiers access between first-party and third-party MCP servers. Uber’s 2026 roadmap adds Dynamic Access Control and a Unified Policy Enforcement Plane, signaling a move toward more granular, session-aware authorization. What this deployment demonstrates is structural, not just technical. Uber built this gateway because its existing microservice architecture — thousands of services across HTTP, gRPC, and TChannel — made ad-hoc agent integrations impossible to govern. The gateway is the answer to a question every large enterprise will face: how do you let agents interact with production systems at scale without losing control of what they can do? For infrastructure builders evaluating the incumbent API providers https://forkast.news/kong-ships-volcano-api-infrastructure-company-becomes-agent-infrastructure/ now entering the agent space, Uber’s architecture sets a high bar for what production-grade MCP governance actually requires.