Blog
Ask a coding agent how to configure eviction in Redis 8 and you'll usually get an answer. Less often than we'd like, that answer comes from authoritative docs. The agent either recalls something from pretraining, or it scrapes the rendered HTML off redis.io at whatever freshness, marketing copy and old version notes included, and paraphrases. Redis 8 changed several defaults, so a model trained before it will hand you the Redis 7 answer with full confidence.
The existing options each fall short of closing that gap. Web search returns pages without stable identifiers, so an agent can't cite what it read or come back to it. Client-side retrieval works, but it makes every framework reimplement chunking, embedding, and freshness against a corpus it doesn't own.
So we built Redis Docs MCP, a public MCP server at redis.io/mcp. It serves three tools over the Redis docs (search, fetch, and ask), unauthenticated, to any MCP client. Four of us built it over one quarter, from April to July 2026.
In this post we cover:
- The MCP tool surface, and why
search/fetchandasktake different paths to the same corpus - Why the docs index has one writer and two readers, and what that constraint costs
- How we reviewed designs before code, and what changed once coding agents were doing much of the work
- A worked example: making
askreturn citations the model can't fabricate - How we measured retrieval quality, and what the measurements pointed at that we weren't looking for
If you're building an agent-facing service on a small team, the middle sections are the transferable part. To connect a client to the docs, skip to the end.
It started as a demo #
The project didn't begin as an MCP server. In late April 2026 we had three weeks to put a grounded docs assistant on the product page for Context Retriever, the Redis product this whole project leans on. You declare an entity schema, and it provisions a managed retrieval surface over a Redis database, then exposes generated retrieval tools to an agent over MCP. That assistant, a chat endpoint on the marketing page, was the first thing this project shipped.
Three weeks produced a FastAPI service streaming server-sent events, an agent that decomposes a question into retrieval calls, and roughly a hundred curated docs pages embedded into a Context Retriever surface. It put the retrieval on screen: a thinking step, a tool call, a tool result, then the answer streaming in. It went live at the start of June 2026.
Those three weeks proved the pipeline worked, and told us nothing about the interface. A demo has one caller, a human reading the output, and a launch date standing in for a specification. A public endpoint has none of those: callers are anonymous MCP clients, nobody is watching the stream, and no date forces the design. We wrote the risk down in the quarterly plan before starting the next phase. The failure mode to avoid was "spending the quarter polishing the web demo without producing a reusable agent-facing interface."
One contract, many clients #
MCP is already how coding agents expect to reach an external tool, so the protocol was never the open question. What needed deciding was the tool surface: what it should expose, and how much of it there should be.
One MCP contract reaches Claude, Codex, Cursor, and ChatGPT Deep Research without us shipping four integrations or asking four vendors to ship one. The protocol also carries tool descriptions in the handshake, which turns out to be the main control we have over agent behaviour.
Other providers had already made that call in different ways. Microsoft Learn and Cloudflare both run docs MCP servers. Upstash's Context7 established that a public, keyless docs MCP is a reasonable thing to operate. Supabase authenticates every caller, which is right for a server acting under a developer's permissions. Ours serves public pages, so we didn't.
Keyless was a deliberate choice, and it follows from who we want calling this. The Redis docs serve our open-source community and our Enterprise customers alike, and putting a signup in front of a docs lookup would shut out most of the first group.
That decision then constrains everything downstream. Any caller can reach the server and we can't tell who they are, so per-caller request text never reaches logs, telemetry is aggregate-only, and neither internal ranking scores nor non-fetchable identifiers cross the wire. An Enterprise customer's questions are as anonymous to us as anyone else's, which is the property we wanted. And because all three tools arrive as one JSON-RPC POST /mcp, per-tool rate limits have to live in the application: the edge can't tell search from ask without parsing the body.
What we shipped #
Redis Docs MCP exposes three tools and nothing else. We held that count deliberately: every tool description is folded into the client's context for the whole connection, so a fourth tool costs every caller tokens on every call whether they use it or not.
| Tool | Returns | When a client should reach for it |
|---|---|---|
| search(query) | Up to 8 matching pages, each with a snippet, a title, a URL, and a stable id | The default for any Redis question |
| fetch(id) | Full text plus topic/version/section metadata for one page | A snippet wasn't enough, or the agent wants to quote exactly |
| ask(question) | A prose answer plus the source documents behind it | A question spanning several areas where no single page answers |
The names collide, so one disambiguation. This is not the official Redis MCP Server, which is self-hosted and connects an agent to your own Redis instance to run commands against your data. Redis Docs MCP is hosted by us, read-only, and touches no Redis instance but our own docs index.
The direct path: search and fetch
search and fetch match the shape OpenAI established for docs MCP servers, down to naming the snippet field text instead of content. ChatGPT Deep Research connectors require that exact contract, so following it means those clients work against Redis Docs MCP with no integration work on either side.
Both tools query the Redis index directly through RedisVL, without going through Context Retriever. That was the more consequential decision, because owning the query is what lets us tune it. Going direct means search issues its own FT.HYBRID against the Redis Query Engine, so the text field, the scorer, the fusion method, and the candidate count are all ours to sweep, and we swept all four. Through generated tools, the query belongs to the gateway.
The gateway path: ask
ask is the tool outside the convention, and the only one that touches Context Retriever and a language model. It runs the same retrieval-and-synthesis pipeline as the marketing demo's chat endpoint, exposed over a second transport, as a shared library rather than an HTTP call into that service. The ordering inside it is worth being precise about: a model holds Context Retriever's generated tools and decides which to call, so the model drives the retrieval rather than summarising it afterwards. Keeping search and fetch off that path buys availability: the two cheap, high-volume tools keep serving when the gateway or the model is unavailable. The cost is that the two paths have to be debugged separately, and that they can and do disagree about the best match for a query.
One writer, two readers #
Two services read the docs index and exactly one process writes it, which is what keeps them from drifting apart.
The ingest pipeline reads Hugo markdown from the redis/docs repository, strips front matter and shortcodes, chunks long pages by ## heading, embeds each chunk, and imports the result through Context Retriever. Neither the demo chat endpoint nor Redis Docs MCP owns the index lifecycle, the schema, or embedding generation. The MCP server doesn't even hold the index name as configuration in production; it discovers the index at startup.
Because Context Retriever generates its agent-facing tools from the declared schema, the schema is the API for anything reaching the corpus through the gateway. Declaring a field as a tag also produces a filter_by_<field> tool that agents can call.
That has teeth. search returns an id that fetch resolves, so both fields have to come back from a query, and in a JSON-mode Redis index only declared attributes resolve by bare name. The obvious fix was to declare doc_id and url as tags. It worked, and it added filter_redisiodoc_by_doc_id and filter_redisiodoc_by_url to the agent's tool surface: two tools nobody should call, competing for the model's attention every turn. We switched to reaching both fields by JSON path and promoting them back in the adapter, then wrote the decision down so nobody re-derives it.
The constraint is cheap to state and not cheap to hold. EMBEDDING_MODEL and EMBEDDING_DIM are exported constants imported by both ingest and the query path, so those two can't diverge. A schema change, though, has to move ingest and both readers together, and re-provision the surface so the gateway regenerates its tools.
We reviewed specs harder than code #
Reviewing the design before any code exists is what made the work delegable, which mattered more than we expected.
Spec-driven development means writing a change's design down in the repository and reviewing that document first, then checking the implementation against it. Ours live in a spec/ directory of about thirty documents, grouped into proposals, interface contracts, reviews, and notes. Each carries a status line rather than moving between folders as it progresses, so its URL never changes.
A wrong paragraph costs an hour to fix and a wrong abstraction costs a week. Neither figure is measured, and it is the ratio that moved us, so we put the friction earlier. Specs got slow, careful, multi-reviewer attention from the whole team, and code review ran behind that at a lighter cadence. Code review still happened, but a review whose contract, interfaces, and acceptance criteria are already agreed is checking conformance rather than arguing about the design. Those reviews are faster, and much easier to hand to someone else.
That last part is the payoff, because coding agents wrote a lot of the code here, with 135 of 374 commits carrying an agent as their author. An agent handed an agreed contract and a set of failable acceptance criteria is doing bounded work. An agent handed a goal is doing exploratory work that someone will have to redo.
We put agents on the review side too, in cycles. That practice has a name now: loop engineering, the design of the feedback loop around an agent rather than the prompt going into it. What paid off was running one implementation past several reviewer personas in sequence, each looking for something different. One reads for conformance to the contract, one for what the change lets escape to a caller (an internal score, an unfetchable id, a request body reaching a log), one tries adversarially to break the change, one checks whether the docs still describe the code.
Those passes changed designs rather than polishing them. Five of them ran over the tool descriptions, and one killed our first escalation rule for ask, which asked the model to judge in advance whether a question would need ask at all. Two reviewers independently made the same objection: that asks for a forecast at the moment the model knows least about the question. We replaced it with something observable, escalating only after a search comes back empty or off-target.
So a spec has to carry what the codebase can't, because an agent reading only the code will re-derive every dead end. We standardised on five things a spec has to contain, in rough order of what they returned:
- Record the alternatives you rejected, and why, so nobody re-opens a settled argument.
- State the non-goals, because given only goals both agents and engineers will helpfully build past them.
- Write the acceptance criteria before the implementation, and make them able to fail.
- Cite the code by path, so a reader can check the spec against the repository instead of trusting it.
- Correct the spec in place when it turns out to be wrong, rather than leaving the error for the next reader.
The sources work in the next section shipped behind one such criterion, and several of our specs now carry a section listing the claims their author got wrong, verified against run logs.
Agents also write too much. A recurring class of commit here does nothing but strip banner comments and over-explanation out of code, and out of specs, that an agent produced.
Citations a model cannot author #
A model will write a citation that looks right and resolves to nothing, so we stopped letting it write them at all.
For a while ask cited docs the way a language model does by default, writing the URLs itself. Then it cited a page that doesn't exist: develop/ai/context-engineering/langcache/api-examples, which 404s, where the real page is develop/ai/context-engine/langcache/api-examples#sec:w0. One path segment off, context-engineering for context-engine, and exactly the kind of error a reader won't catch by looking. The #sec:w0 suffix is ours: ingest splits a long page into windowed sections and gives each its own id, so a citation names a section rather than a whole page.
ask now returns a sources array alongside the answer, derived from what the retrieval pipeline actually fetched during the turn. Every id came out of a real tool call, so a fabricated one can't appear, and every id resolves through fetch. Context Retriever's gateway returns a structured record of what it retrieved, so the citation list was already sitting there to be read off instead of inferred from the answer text. We hadn't planned for that.
That left an ordering problem. Each source gets a weight, calculated from the distinctive facts the answer asserts (command names, configuration directives, numbers with units) and how rare each is across the retrieved set. We reorder the full set by weight and never prune it, because the risk is asymmetric: dropping a document that grounded a claim is undetectable from the output, while ranking one three positions too low costs a reader an extra click. What matters is the top of the list, since a caller who reads only the first few entries should get the documents that grounded the answer.
Round-robin looks like the way to protect it, and it fails in a way that only shows up when you run it. Long pages contribute one record per section, so a single page can occupy six of fifteen slots, and interleaving by page looks like the fix. But round-robin pushes a page's k-th hit to position (k−1) × P + 1, where P is the number of distinct pages retrieved, so displacement tracks how many other pages came back rather than how redundant the page is. Measured on one real payload, where a page contributed three well-grounded sections weighted 8, 7, and 6 alongside five weak documents weighted 0.2 each, interleaving sent the second and third sections to positions 9 and 11, behind documents carrying a thirtieth of their weight. The top five ended up holding 42% less total weight. It also failed on the case that motivated it, leaving the top five unchanged and making the longest same-page run longer, because once the shallow pages are exhausted the deep page's remainder emerges consecutively anyway.
A weight penalty worked instead. A page's n-th hit has to out-score another page's best by n−1 units, where a unit is what a token earns by appearing in exactly one of the documents retrieved (log(total_docs)). So a page with three independently strong sections keeps all three near the top, while a page repeating itself drops away after the first hit. The measured improvement is small: two repeats of one page leave the top six, and two distinct pages take their slots.
The feature shipped behind an explicit gate, requiring 100% of returned ids to resolve through fetch before it went out.
Checking the work #
Grade the evidence, not the answer. A strong model turns incomplete evidence into convincing prose, so grading answers tells you little about whether the retrieval was right. The rule has an edge worth naming: evidence that scores well can still be synthesised badly, so these numbers bound what the retriever offered and never what the caller read.
That meant two harnesses, kept separate because they measure different layers:
| Harness | What it measures | Determinism | Cost and time |
|---|---|---|---|
| Pure retrieval | Query → embedding → retrieval tool. No model in the loop. | Deterministic | Free, ~30 s |
| Agentic | Query → the full agent, capturing every tool call it makes | Stochastic even at temperature 0 | ~$0.20–1, ~10 min |
Both run against the same labelled dataset with graded relevance, where a document that directly answers the question is worth more than one that provides context, and both use the same scoring function. Splitting them isolated the agent's contribution from the retriever's. Measured over the 11 simple-lookup questions in that set, the agent's query reformulation lifted recall@5 from 0.545 to 0.818, because it rewrites a bad query and searches again. Over the 8 multi-hop questions it moved top-5 chunk precision the other way, from 0.275 to 0.225, because it issues more searches and dilutes the window.
Agents built most of the surrounding harness: the Locust load-testing scenarios, the collectors that capture raw event frames for later review, and the report generators.
A load test measures whatever your configuration lets it reach. Our first serious run put 50 concurrent users against the demo chat endpoint and came back with a 98.2% failure rate, every failure an HTTP 429. We had measured the rate limiter, not the service, and the run said nothing about where the service breaks or whether autoscaling responds in time, which were the two questions we'd set out to answer.
Making a change earn itself is cheaper than arguing about it. The prototype ran on pure vector search, and we expected to need hybrid search for anything production-grade, because exact tokens are where embedding similarity is weakest: XADD and SINTERCARD carry almost no semantic signal as vectors and are unambiguous as strings. Rather than assert that, the pull request implementing it shipped with an evaluation attached, measured over 73 queries against the vector-only baseline it replaced, on the chunked corpus with an approximate vector index.
Hybrid search runs as FT.HYBRID, with Redis doing the text scoring, the score normalisation, and the rank fusion server-side. Nothing is fused in application code.
| Configuration | MRR | Hit@1 | Hit@3 | recall@8 |
|---|---|---|---|---|
| Vector only | 0.816 | 0.75 | 0.88 | 0.93 |
| Hybrid, title field, RRF fusion | 0.836 | 0.71 | 0.96 | 0.99 |
Hit@3 gains 0.08 and recall@8 gains 0.06, and retrieval latency stays where it was, at 0.5 to 0.7 ms p50 and under 1.3 ms p95 measured locally. Hit@1 loses 0.04, because rank fusion sometimes swaps the exact top result: an acceptable trade for a tool whose caller reads all eight and can fetch any of them, and a bad one for a tool that returns a single answer. Ranking against title beat content on every metric, because a title match fires when the query names a page and rarely otherwise. Reciprocal rank fusion beat weighted linear combination, which was both slower and worse on every ranking metric we measured.
The same evaluation told us something we hadn't gone looking for. Most of hybrid search's gain comes from backstopping approximate-nearest-neighbour misses rather than from adding lexical matching: on an exact-search index, the very exact-token queries we'd been worried about already ranked first with vector search alone. Retrieval over whole pages also edged out retrieval over chunks on the lexical subset of those queries, 0.776 against 0.762 MRR. Treat that margin as a direction rather than a measurement, because the two runs used different vector indexes, and on a matched index the chunked corpus wins overall. It was still enough to send us looking at the corpus rather than the algorithm: the tokens weren't hard to match, and our chunking was splitting up the pages that contained them.
What we'd change #
Design the identifier scheme before there's a corpus. search returns an id that fetch has to resolve, and that one contract reaches into ingest, chunking, the URL scheme, and both readers. We widened validators and re-derived what makes an id fetchable more than once before it settled.
Measure the corpus before tuning the retriever. Hybrid search was a real improvement and not the improvement that mattered most: our numbers kept pointing at the shape of the data underneath. We had declared one flat document entity, chunked by heading, with nothing distinguishing a command reference page from a conceptual guide and no way to express that the two relate, so the retriever had no path to a command as an entity and no edge to follow from a concept to the commands that implement it. No amount of tuning the retrieval algorithm fixes that. Declaring entities and the relationships between them turns the corpus into something an agent can navigate rather than only sample, and it is the difference between using Context Retriever as a vector index with filters and using what it actually does.
That navigable layer, and the ontology under it, is the larger piece of work this project turned into, and it's the subject of a companion post by Hillary Toh.
Try it #
Redis Docs MCP is public and needs no key. For an MCP client that reads a JSON config:
Start with search, fetch what you need, and reserve ask for questions that span several areas of the docs.
Acknowledgements #
Built along with Robert Shelton, Hillary Toh, and Joo Bin Lim, who each worked across most of what this post describes.
Get started with Redis today #
Speak to a Redis expert and learn more about enterprise-grade Redis today.