How we built the Redis Docs MCP for agents Redis launched Redis Docs MCP, a public, unauthenticated MCP server at redis.io/mcp that exposes three tools — search, fetch, and ask — over the Redis documentation, built by four engineers over one quarter from April to July 2026. The server is designed to stop coding agents from answering Redis 8 configuration questions with stale Redis 7 defaults pulled from pretraining or scraped HTML, and it was preceded by a three-week grounded docs assistant for the Context Retriever product page that went live at the start of June 2026. Redis chose a keyless design because the server serves public pages, unlike Supabase's authenticated docs MCP, and cites Microsoft Learn, Cloudflare, and Upstash's Context7 as prior docs MCP servers. Blog How we built the Redis Docs MCP for agents Ask a coding agent how to configure eviction in Redis 8 and you'll usually get an answer. Less often than we'd like, that answer comes from authoritative docs. The agent either recalls something from pretraining, or it scrapes the rendered HTML off redis.io at whatever freshness, marketing copy and old version notes included, and paraphrases. Redis 8 changed several defaults, so a model trained before it will hand you the Redis 7 answer with full confidence. The existing options each fall short of closing that gap. Web search returns pages without stable identifiers, so an agent can't cite what it read or come back to it. Client-side retrieval works, but it makes every framework reimplement chunking, embedding, and freshness against a corpus it doesn't own. So we built Redis Docs MCP, a public MCP server at redis.io/mcp https://redis.io/mcp/ . It serves three tools over the Redis docs search , fetch , and ask , unauthenticated, to any MCP client. Four of us built it over one quarter, from April to July 2026. In this post we cover: - The MCP tool surface, and why search / fetch and ask take different paths to the same corpus - Why the docs index has one writer and two readers, and what that constraint costs - How we reviewed designs before code, and what changed once coding agents were doing much of the work - A worked example: making ask return citations the model can't fabricate - How we measured retrieval quality, and what the measurements pointed at that we weren't looking for If you're building an agent-facing service on a small team, the middle sections are the transferable part. To connect a client to the docs, skip to the end. It started as a demo The project didn't begin as an MCP server. In late April 2026 we had three weeks to put a grounded docs assistant on the product page for Context Retriever, the Redis product this whole project leans on. You declare an entity schema, and it provisions a managed retrieval surface over a Redis database, then exposes generated retrieval tools to an agent over MCP. That assistant, a chat endpoint on the marketing page, was the first thing this project shipped. Three weeks produced a FastAPI service streaming server-sent events, an agent that decomposes a question into retrieval calls, and roughly a hundred curated docs pages embedded into a Context Retriever surface. It put the retrieval on screen: a thinking step, a tool call, a tool result, then the answer streaming in. It went live at the start of June 2026. Those three weeks proved the pipeline worked, and told us nothing about the interface. A demo has one caller, a human reading the output, and a launch date standing in for a specification. A public endpoint has none of those: callers are anonymous MCP clients, nobody is watching the stream, and no date forces the design. We wrote the risk down in the quarterly plan before starting the next phase. The failure mode to avoid was "spending the quarter polishing the web demo without producing a reusable agent-facing interface." One contract, many clients MCP is already how coding agents expect to reach an external tool, so the protocol was never the open question. What needed deciding was the tool surface: what it should expose, and how much of it there should be. One MCP contract reaches Claude, Codex, Cursor, and ChatGPT Deep Research without us shipping four integrations or asking four vendors to ship one. The protocol also carries tool descriptions in the handshake, which turns out to be the main control we have over agent behaviour. Other providers had already made that call in different ways. Microsoft Learn and Cloudflare both run docs MCP servers. Upstash's Context7 established that a public, keyless docs MCP is a reasonable thing to operate. Supabase authenticates every caller, which is right for a server acting under a developer's permissions. Ours serves public pages, so we didn't. Keyless was a deliberate choice, and it follows from who we want calling this. The Redis docs serve our open-source community and our Enterprise customers alike, and putting a signup in front of a docs lookup would shut out most of the first group. That decision then constrains everything downstream. Any caller can reach the server and we can't tell who they are, so per-caller request text never reaches logs, telemetry is aggregate-only, and neither internal ranking scores nor non-fetchable identifiers cross the wire. An Enterprise customer's questions are as anonymous to us as anyone else's, which is the property we wanted. And because all three tools arrive as one JSON-RPC POST /mcp , per-tool rate limits have to live in the application: the edge can't tell search from ask without parsing the body. What we shipped Redis Docs MCP exposes three tools and nothing else. We held that count deliberately: every tool description is folded into the client's context for the whole connection, so a fourth tool costs every caller tokens on every call whether they use it or not. | Tool | Returns | When a client should reach for it | |---|---|---| | search query | Up to 8 matching pages, each with a snippet, a title, a URL, and a stable id | The default for any Redis question | | fetch id | Full text plus topic/version/section metadata for one page | A snippet wasn't enough, or the agent wants to quote exactly | | ask question | A prose answer plus the source documents behind it | A question spanning several areas where no single page answers | The names collide, so one disambiguation. This is not the official Redis MCP Server, which is self-hosted and connects an agent to your own Redis instance to run commands against your data. Redis Docs MCP is hosted by us, read-only, and touches no Redis instance but our own docs index. The direct path: search and fetch search and fetch match the shape OpenAI established for docs MCP servers https://developers.openai.com/mcp , down to naming the snippet field text instead of content . ChatGPT Deep Research connectors require that exact contract, so following it means those clients work against Redis Docs MCP with no integration work on either side. Both tools query the Redis index directly through RedisVL, without going through Context Retriever. That was the more consequential decision, because owning the query is what lets us tune it. Going direct means search issues its own FT.HYBRID against the Redis Query Engine, so the text field, the scorer, the fusion method, and the candidate count are all ours to sweep, and we swept all four. Through generated tools, the query belongs to the gateway. The gateway path: ask ask is the tool outside the convention, and the only one that touches Context Retriever and a language model. It runs the same retrieval-and-synthesis pipeline as the marketing demo's chat endpoint, exposed over a second transport, as a shared library rather than an HTTP call into that service. The ordering inside it is worth being precise about: a model holds Context Retriever's generated tools and decides which to call, so the model drives the retrieval rather than summarising it afterwards. Keeping search and fetch off that path buys availability: the two cheap, high-volume tools keep serving when the gateway or the model is unavailable. The cost is that the two paths have to be debugged separately, and that they can and do disagree about the best match for a query. One writer, two readers Two services read the docs index and exactly one process writes it, which is what keeps them from drifting apart. The ingest pipeline reads Hugo markdown from the redis/docs https://github.com/redis/docs repository, strips front matter and shortcodes, chunks long pages by heading, embeds each chunk, and imports the result through Context Retriever. Neither the demo chat endpoint nor Redis Docs MCP owns the index lifecycle, the schema, or embedding generation. The MCP server doesn't even hold the index name as configuration in production; it discovers the index at startup. Because Context Retriever generates its agent-facing tools from the declared schema, the schema is the API for anything reaching the corpus through the gateway. Declaring a field as a tag also produces a filter by