You build a support workflow in two steps. At 06:00 a scheduled job connects to your MCP server, reads the night's tickets and works out that 4127 and 4133 are the same bug. At 10:00 you open a chat client and ask the agent which tickets can be closed. That is a new client connection, perhaps to a different server instance, and the agent starts from nothing. What the first step worked out lived with the first connection, and that connection closed four hours ago.
The short answer: do not key the context by the connection. Store it outside the server process, under the stable identity it belongs to, such as the user, the project or the agent, and have every connection authenticate as that identity and read the context before it acts. Inside one multi-step task, carry an explicit handle as an ordinary tool argument. Since the 2026-07-28 revision of MCP there is no protocol session left to key context by, and the specification tells servers not to require the same connection for related operations.
Disclosure: I work on Mnemoverse, the memory layer in the example near the end.
Earlier revisions of MCP gave you a key to keep context under. The 2025-06-18 transport says a server "MAY assign a session ID at initialization time" and that the server may "terminate the session at any time" (2025-06-18 transports). The current transport page sums up that model as "a connection-scoped session with an initialize handshake" (2026-07-28 transports). Context kept under that ID could be found only by requests that carried the same ID, which means only inside the same session.
The 2026-07-28 revision drops the session: the first major change in its changelog is "Remove protocol-level sessions and the Mcp-Session-Id header from the Streamable HTTP transport". The Statelessness section of the specification says what follows for everyone who builds on the protocol:
A server processes each request independently; no state should be inferred from previous requests, even those on the same connection or stream.
Two of the rules under that sentence describe the opening example. Servers "SHOULD NOT require that a client reuse the same connection or process to perform related operations", and state that spans requests "MUST be referenced by an explicit identifier the client passes on each request". Behind a load balancer the next request can land on any instance, so the store has to be one that every instance reaches.
A server that also serves older clients still serves them by the legacy rules, "scoped to the stdio process (stdio) or the session (HTTP)" (versioning and compatibility), and context kept there ends with that process or session.
What a piece of context is keyed by decides which future connection can find it.
| Context keyed by | Example | Ends when | A new connection reaches it if |
|---|---|---|---|
| The connection or session | state in server memory under an Mcp-Session-Id , or per stdio process |
the connection closes, the session is terminated, or the process restarts | only with the same session ID, which an independent client lacks; 2026-07-28 requests carry none |
| A workflow handle | a job_id returned by the tool that started a long job |
the handle expires or the job is done | the model passes the handle back as a tool argument |
| A stable identity | the user, the team, the project, the agent | it is deliberately removed | it authenticates as that identity and reads before acting |
The first row is what broke in the opening example. The second row is the protocol's answer for work that spans calls: the changelog says "Servers that need cross-call state use explicit, server-minted handles passed as ordinary tool arguments". A handle reaches a new connection only if the model still holds it, which suits one job and not a conversation that starts four hours later. The third row is the only one that answers the question as asked.
Put as a definition: connection-keyed context ends with the connection that produced it; identity-keyed context outlives every connection, and any connection that authenticates as its owner can read it.
The identity comes from the credential, not from the model. The connection presents an API key or an OAuth token, and the server derives whose context it is from that. If the model chooses the key, any user can read another user's context by naming it.
Read at the start of every connection, before acting. A new connection brings nothing with it, so the workflow has to ask. Put the read in the agent's standing instructions or in the first step of the workflow.
Write when the context is established, not when the connection ends. Connections end without warning, and on the 2026-07-28 revision not even the request in flight survives a break: "A broken response stream loses the in-flight request; clients MUST re-issue it as a new request with a new request ID" (changelog). A summary planned for the end of a session is the first thing lost.
Scope inside the identity. One user runs several projects. Keep each project's context under its own namespace, so a connection working on one project does not read the conclusions of another.
In Mnemoverse, memory belongs to the Mnemoverse account, not to the connection. A connection shows which account it is through an API key or a one-click OAuth sign-in. Our MCP server page says "Memory is scoped to the account behind the key" and that "two keys from the same console account share one memory", and the hosted connector, which signs in with OAuth instead of a key, reaches the same memory. So a 10:00 client on the same account can read what the 06:00 job wrote.
One catch, stated on the connector page and true of OAuth connectors in general: the host app holds the connection, and signing out or switching accounts there usually drops it. The memory is untouched: "Sign in again with the same Mnemoverse account and everything is back, unchanged". For unattended workflows the docs recommend the local package with an API key, which lives in your config, not in the host app's session. Connecting gives the agent the tools, not the habit of reading first; Make Your Agent Use Memory has the instructions for that.
The same rule holds outside MCP. With the Python SDK (mnemoverse 0.4.0 from PyPI), which calls the REST API, the two steps of the opening example are two processes with two connections, sharing only the account behind the key and the domain:
from mnemoverse import MnemoClient
with MnemoClient() as memory: # reads MNEMOVERSE_API_KEY
result = memory.write(
"Tickets 4127 and 4133 are the same bug: both fail when an OAuth refresh "
"token has expired. Keep 4127, close 4133 as its duplicate.",
concepts=["tickets", "duplicate", "oauth"],
domain="project:support",
)
if not result.stored:
print("Not stored:", result.reason)
python
from mnemoverse import MnemoClient
with MnemoClient() as memory:
found = memory.read("which tickets are duplicates", domain="project:support", top_k=5)
for item in found.items:
print(item.created_at, item.content)
domain is the namespace inside the account (the API reference gives forms such as user:X and project:Z), so the second script reads the support project and nothing else. A write that is too similar to the nearest memory in the same domain is not stored, which here would mean the conclusion is already there, so the first script prints the reason instead of failing silently. For everything written since a given time, recent() lists the newest entries first with nothing ranked away. Setup for scripts is on the Python SDK page.
The library page has the longer account: what the 2026-07-28 revision removes, how a handle differs from durable memory, and why A2A also leaves memory outside the protocol. Which step of your workflow still assumes the next call arrives on the same connection?
The Python SDK is open source (MIT) and on PyPI as mnemoverse: github.com/mnemoverse/mnemoverse-sdk-python. The MCP server package on npm is open source too: github.com/mnemoverse/mcp-memory-server.