{"slug": "caching-hints-and-pagination-in-mcp", "title": "Caching hints and pagination in MCP", "summary": "The Model Context Protocol's 2026-07-28 revision requires complete results from six methods — server/discover, tools/list, prompts/list, resources/list, resources/templates/list and resources/read — to include the caching hints ttlMs and cacheScope, while input_required results carry none. The revision also directs servers to return tools in a deterministic order to keep client caches reliable and improve prompt-cache hit rates, and to paginate with opaque cursors via nextCursor and params.cursor. The author advises marking a list private when it depends on the caller's authorization scopes, warning that a public list could reach the wrong user through a shared cache.", "body_md": "# Caching hints and pagination in MCP\n\nOnce the Model Context Protocol (MCP) dropped sessions, a small question got bigger: how often should a client ask your server for its tool list? With no long-lived connection in which to learn about a server once, a client would otherwise fetch the same lists over and over. The 2026-07-28 revision answers with explicit caching hints. The part that surprised me sits right next to them: the order your tools come back in can quietly cost your users money.\n\nThis is part 6 of my series on what an MCP server does under the 2026-07-28 specification. [Part 5](https://pournasserian.com/writing/mcp-2026-stateless-requests) covered `_meta`, `resultType` and explicit handles; this part covers the caching hints on a server’s results, a stable order and pagination. The facts are as I read them in October 2026.\n\n## In brief\n\n1. **The results of six methods carry caching hints.** Complete results of`server/discover` and the five list and read methods MUST include`ttlMs` and`cacheScope` ; an`input_required` result carries none.\n2. **A time to live (TTL) caps staleness, and a notification ends a cached copy early.**`ttlMs` tells a client how long it may reuse a result; a`list_changed` notification tells a subscribed client the moment something changes.\n3. **My advice: mark a list `private` when it depends on who’s asking.** If you filter tools by the caller’s authorization scopes, a`public` list could reach the wrong user through a shared cache.\n4. **Keep the order the same on every call.** Servers SHOULD return tools in a deterministic order, which keeps client caches reliable and improves prompt-cache hit rates.\n5. **Pagination uses opaque cursors.** The server returns`nextCursor` , the client sends it back as`params.cursor` , and a stable order keeps the page boundaries stable.\n\n## What the hints say\n\nThe [specification’s caching page](https://modelcontextprotocol.io/specification/2026-07-28/server/utilities/caching) names six methods whose complete results, those with `resultType: \"complete\"`, MUST include `ttlMs` and `cacheScope`:\n\n- `server/discover` , the method[part 3](https://pournasserian.com/writing/mcp-2026-server-discover) covers\n- `tools/list`\n- `prompts/list`\n- `resources/list`\n- `resources/templates/list`\n- `resources/read`\n\nAn interim result with `resultType: \"input_required\"` is not cacheable and carries no hints. That matters for `resources/read`: the [Resources page](https://modelcontextprotocol.io/specification/2026-07-28/server/resources) says a server MAY answer it with an `input_required` result. Extensions follow the same convention; `skills/list` in the Skills extension is one example.\n\n| Field | Type | Meaning | \n|---|---|---|\n| `ttlMs` | integer, in milliseconds | A freshness hint: how long the client may reuse this result without asking again | \n| `cacheScope` | `\"public\"` or`\"private\"` | Whether shared intermediaries (proxies, gateways, content delivery networks) may cache the response | \n\nThe hints add to the `list_changed` notifications rather than replace them. A TTL puts a ceiling on staleness for a client that isn’t subscribed, and a notification invalidates the cache at once for a client that is. A client subscribes with [`subscriptions/listen`](https://modelcontextprotocol.io/specification/2026-07-28/basic/patterns/subscriptions), which gets its own part later in this series.\n\n## Choosing the scope and the TTL\n\nThe spec says which results carry the hints; the values are yours. This is how I choose them.\n\n### `cacheScope`\n\n| Situation | Scope | \n|---|---|\n| The list is identical for every caller | `public` | \n| The list depends on the caller’s scopes or identity | `private` | \n| A resource’s content is user-specific, such as a mailbox or personal files | `private` | \n| Public documentation, static user interface (UI) templates, skill manifests that are the same for everyone | `public` | \n\nThe spec allows a list to vary by authorization, so my advice is this: if your server filters tools by the caller’s authorization scopes, mark its list results `private`. Otherwise a shared gateway cache could serve one user’s tool list to another. That leak isn’t catastrophic on its own, but it reveals which capabilities exist, and it confuses clients.\n\n### `ttlMs`\n\nThe spec sets no ranges, so this table is my rule of thumb. The spec’s own examples sit inside it: they use 5 minutes for lists and 1 hour for [`server/discover`](https://modelcontextprotocol.io/specification/2026-07-28/server/discover).\n\n| Data | My typical TTL | \n|---|---|\n| `server/discover` | 1 hour or more | \n| Tool, prompt and template lists for a stable deployment | 5 to 60 minutes, shorter if you rely on feature flags | \n| Static resources, such as docs and UI templates | Hours | \n| Live resources, such as status and metrics | Seconds, or 0 with subscriptions | \n\n### In Python\n\nWith the official Python software development kit (SDK), `mcp` 2.3.0, you set the hints once per method with `cache_hints` on `MCPServer`. The keys are the six methods above, and each value is a `CacheHint`.\n\n``` python\nimport anyio\nfrom mcp import Client\nfrom mcp.server import CacheHint, MCPServer\n\nmcp = MCPServer(\n    \"catalog\",\n    cache_hints={\n        \"server/discover\": CacheHint(ttl_ms=3_600_000, scope=\"public\"),\n        \"tools/list\": CacheHint(ttl_ms=300_000, scope=\"private\"),\n    },\n)\n\n@mcp.tool()\ndef search_catalog(query: str) -> str:\n    \"\"\"Search the product catalog.\"\"\"\n    return f\"No results for {query!r}\"\n\nasync def main():\n    async with Client(mcp) as client:\n        tools = await client.list_tools()\n        prompts = await client.list_prompts()  # no hint set for prompts/list\n        print(\"tools/list  \", tools.ttl_ms, tools.cache_scope)\n        print(\"prompts/list\", prompts.ttl_ms, prompts.cache_scope)\n\nif __name__ == \"__main__\":\n    anyio.run(main)\n```\n\nRun in memory, it prints:\n\n```\ntools/list   300000 private\nprompts/list 0 private\n```\n\nThe last line is what the SDK does when you set nothing: `ttlMs` 0, which means stale at once, and `cacheScope` `private` (the same default [part 3](https://pournasserian.com/writing/mcp-2026-server-discover) showed for `server/discover`). The result stays valid on the wire and is never shared by accident. The SDK also leaves an `input_required` result without hints, as the spec asks.\n\nOn the client side, the SDK’s [`Client`](https://py.sdk.modelcontextprotocol.io/client/caching/) follows the hints by default:\n\n- It keeps an in-memory cache per client and reuses a result until its `ttlMs` runs out, capped at 24 hours.\n- It drops the result when a list-changed notification arrives.\n- Passing `cache_mode=\"refresh\"` or`\"bypass\"` to a call sends it to the server anyway.\n\n## A stable order, and pagination\n\nServers SHOULD return tools from [`tools/list`](https://modelcontextprotocol.io/specification/2026-07-28/server/tools) in a deterministic order: the same order across requests whenever the set of tools hasn’t changed. The spec gives two reasons. Clients can cache the list reliably, and prompt-cache hit rates improve when tool definitions sit in the same position of the model’s context on every turn.\n\nThe spec asks this of `tools/list`. My advice goes further: sort by a stable key (the name, or an explicit display order) before you serialize, and do the same for prompts, resources and `skills/list`. Never build the list from a dictionary or a reflection call whose order can change between processes. An unstable order silently costs your users money through prompt-cache misses. In the Python SDK, `MCPServer` lists tools in the order you register them, so with it the advice becomes: register your tools in a fixed order in your code.\n\nThe list methods (`tools/list`, `prompts/list`, `resources/list`, `resources/templates/list`, and extension lists such as `skills/list`) page with cursors:\n\n- The server returns `nextCursor` when more results exist.\n- The client passes it back as `params.cursor` .\n- Cursors are opaque, so a client must not parse or construct them.\n\nPagination and a stable order go together: page boundaries are only stable if the order underneath them is. My advice: if your catalog changes often, encode a snapshot version in the cursor, so a page fetched mid-change doesn’t skip or repeat items.\n\n`MCPServer` returns each list in one page, with no `nextCursor`. To page a large catalog in Python, write the list handler on the SDK’s [low-level `Server`](https://py.sdk.modelcontextprotocol.io/advanced/low-level-server/): it receives the cursor in `params.cursor` and returns `next_cursor` with each page.\n\n## Method and caveats\n\n- Built from my guide, written against the 2026-07-28 specification, with the facts as I read them in October 2026.\n- The Python sample was run against `mcp` 2.3.0 on 8 October 2026 with the SDK’s in-memory client, and the output above is what it printed. The other SDK notes (the client cache, the tool order, single-page lists, low-level paging) come from the installed package’s source and a second run.\n- Where the specification is more precise than my notes (six methods carry the hints, `server/discover` among them, and only complete results do), this article follows the specification.\n- The TTL table is my advice, not the spec’s; the spec’s examples quoted beside it were checked on 8 October 2026.", "url": "https://wpnews.pro/news/caching-hints-and-pagination-in-mcp", "canonical_source": "https://pournasserian.com/writing/mcp-2026-caching-and-pagination", "published_at": "2026-10-09 00:00:00+00:00", "updated_at": "2026-10-10 16:48:33.763420+00:00", "lang": "en", "topics": ["agent-protocols", "ai-agents", "developer-tools", "ai-infrastructure"], "entities": ["Model Context Protocol", "server/discover", "tools/list", "prompts/list", "resources/list", "resources/templates/list", "resources/read", "skills/list"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/caching-hints-and-pagination-in-mcp", "markdown": "https://wpnews.pro/news/caching-hints-and-pagination-in-mcp.md", "text": "https://wpnews.pro/news/caching-hints-and-pagination-in-mcp.txt", "jsonld": "https://wpnews.pro/news/caching-hints-and-pagination-in-mcp.jsonld"}}