Build Your First MCP Server in Python (Stateless Spec Edition) The Model Context Protocol's 2026-07-28 specification makes the protocol core stateless, removing the need for clients to establish a session or servers to rely on Mcp-Session-Id for ordinary requests, and the official Python SDK has moved to v2 as its stable line with a higher-level MCPServer API. A tutorial demonstrates building a stateless developer knowledge-base MCP server in Python 3.10+ that exposes a search_kb tool, a kb://articles resource, and a draft_support_reply prompt over Streamable HTTP, installed via uv add "mcp[cli]" or pip install "mcp[cli]" and inspectable with MCP Inspector. Build Your First MCP Server in Python Stateless Spec Edition MCP just got simpler. Learn how to build a stateless MCP server in Python and expose tools, resources, and prompts over HTTP. The Model Context Protocol https://modelcontextprotocol.io/ changed in an important way this summer. With the 2026-07-28 MCP specification https://modelcontextprotocol.io/specification/2026-07-28/changelog , the protocol core is now stateless : modern clients no longer need to establish a protocol session before making requests, and servers no longer rely on Mcp-Session-Id for ordinary requests. That makes MCP servers much easier to scale behind normal HTTP infrastructure. At the same time, the official Python SDK https://github.com/modelcontextprotocol/python-sdk has moved to v2 as its current stable line and provides a higher-level MCPServer API for defining tools, resources, and prompts with regular Python functions. In this tutorial, we will build a small but complete MCP server in Python, run it over Streamable HTTP, inspect it locally, and connect to it with a Python MCP client. Let's get started. What Are We Building? We will create a small developer knowledge-base server . It will expose three MCP primitives: Tool: search kb query, limit Resource: kb://articles Prompt: draft support reply customer message The server will contain no user session state. Every request will contain everything required to process it, which makes it a good example of the new stateless MCP model. Conceptually: LLM Host | | MCP request v +-----------------------+ | Python MCP Server | | | | search kb | | kb://articles | | draft support reply | +-----------------------+ Step 1: Creating the Project The current Python SDK requires Python 3.10 or newer. The official documentation recommends installing the CLI extra because it gives us the development command and MCP Inspector workflow. Using uv https://docs.astral.sh/uv/ : mkdir first-mcp-server cd first-mcp-server uv init uv add "mcp cli " Or with pip : pip install "mcp cli " Your project can be as small as: first-mcp-server/ ├── server.py └── client.py No framework boilerplate is required. Step 2: Creating Your First MCP Server Create server.py : python from mcp.server import MCPServer mcp = MCPServer "Developer Support KB", instructions= "Use the knowledge-base tools to answer support questions. " "Prefer retrieved KB information over guessing." , MCPServer is the high-level server API in the current Python SDK. For most servers, this is the API you want. The SDK also exposes a lower-level Server class, but that is intended for cases where you need exact control over schemas, protocol metadata, or custom methods. Now let's give our server some data. ARTICLES = { "id": "python-env", "title": "Creating a Python virtual environment", "body": "Create a virtual environment with python -m venv .venv , " "then activate it before installing dependencies." , }, { "id": "reset-password", "title": "Resetting your password", "body": "Open Account Settings, choose Security, and select " "Reset Password. A verification email will be sent." , }, { "id": "api-rate-limit", "title": "Understanding API rate limits", "body": "API rate limits restrict the number of requests allowed " "within a time window. Clients should retry using " "exponential backoff after receiving a rate-limit response." , }, So far this is just Python. The interesting part begins when we expose functions through MCP. Step 3: Adding an MCP Tool An MCP tool is a function the model can decide to call. Add this to server.py : php @mcp.tool def search kb query: str, limit: int = 3 - list dict str, str : """Search the support knowledge base. Args: query: Words or phrases to search for. limit: Maximum number of articles to return. """ query = query.lower matches = for article in ARTICLES: searchable text = article "title" + " " + article "body" .lower if query in searchable text: matches.append article return matches :limit Notice what we did not write. There is no JSON Schema. There is no manually written tool manifest. There is no argument parser. The SDK derives the tool definition from the Python function itself. Its type hints become the MCP input schema, and defaults such as: limit: int = 3 make parameters optional in the generated schema. The official SDK documentation uses exactly this pattern. Conceptually, your function: python def search kb query: str, limit: int = 3 becomes something similar to: { "name": "search kb", "inputSchema": { "type": "object", "properties": { "query": { "type": "string" }, "limit": { "type": "integer", "default": 3 } }, "required": "query" } } This is one of the reasons MCP development in Python feels pleasantly ordinary: your function signature is effectively your interface definition. Step 4: Adding a Resource Tools are actions the model can call. Resources are different. They expose information that the host application can load into context. Add: php @mcp.resource "kb://articles" def list articles - str: """Return the available knowledge-base articles.""" lines = for article in ARTICLES: lines.append f"{article 'id' }: {article 'title' }" return "\n".join lines The resource has the URI: kb://articles A client can read it without invoking a tool. Resource ≈ data that can be read Tool ≈ function that can perform work The SDK documentation roughly compares resources with GET -like behavior and tools with action-oriented POST -like behavior. Step 5: Adding an MCP Prompt We can also expose a reusable prompt template. php @mcp.prompt def draft support reply customer message: str - str: """Create a prompt for drafting a concise support response.""" return f""" You are a technical support assistant. Write a concise and helpful response to this customer message: {customer message} Use the support knowledge base when relevant. Do not invent product policies. """.strip Again, this is just a Python function plus a decorator. Prompts are generally initiated by the user or host rather than autonomously invoked by the model. The current SDK supports tools, resources, and prompts through the same decorator-oriented server interface. At this point, server.py looks like this: python from mcp.server import MCPServer mcp = MCPServer "Developer Support KB", instructions= "Use the knowledge-base tools to answer support questions. " "Prefer retrieved KB information over guessing." , ARTICLES = { "id": "python-env", "title": "Creating a Python virtual environment", "body": "Create a virtual environment with python -m venv .venv , " "then activate it before installing dependencies." , }, { "id": "reset-password", "title": "Resetting your password", "body": "Open Account Settings, choose Security, and select " "Reset Password. A verification email will be sent." , }, { "id": "api-rate-limit", "title": "Understanding API rate limits", "body": "API rate limits restrict the number of requests allowed " "within a time window. Clients should retry using " "exponential backoff after receiving a rate-limit response." , }, @mcp.tool def search kb query: str, limit: int = 3, - list dict str, str : """Search the support knowledge base.""" query = query.lower matches = for article in ARTICLES: searchable text = article "title" + " " + article "body" .lower if query in searchable text: matches.append article return matches :limit @mcp.resource "kb://articles" def list articles - str: """Return the available knowledge-base articles.""" return "\n".join f"{article 'id' }: {article 'title' }" for article in ARTICLES @mcp.prompt def draft support reply customer message: str, - str: """Create a support-response prompt.""" return f""" You are a technical support assistant. Write a concise and helpful response to this customer message: {customer message} Use the support knowledge base when relevant. Do not invent product policies. """.strip if name == " main ": mcp.run "streamable-http" That is a complete network-accessible MCP application. Step 6: Running It in Development Mode For development, the SDK includes a convenient command: uv run mcp dev server.py The MCP development command launches the server with MCP Inspector support, giving you a UI for listing and invoking tools. The official SDK recommends this as the basic development loop. Open the Inspector URL printed in your terminal. You should see: search kb under Tools. Try calling it with: { "query": "rate limit" } The result should contain: { "id": "api-rate-limit", "title": "Understanding API rate limits", "body": "API rate limits restrict ..." } You now have a working MCP server. Step 7: Running It Over Streamable HTTP For an actual HTTP server, run: uv run python server.py By default, your MCP endpoint is exposed at: http://127.0.0.1:8000/mcp The SDK's Streamable HTTP server uses /mcp as its default endpoint. If you prefer running it as a normal ASGI application, replace the main block with: app = mcp.streamable http app Then launch it with Uvicorn https://www.uvicorn.org/ : uvicorn server:app This is particularly useful when MCP is one component inside a larger FastAPI https://fastapi.tiangolo.com/ or Starlette https://www.starlette.io/ deployment. streamable http app returns a standard Starlette-compatible ASGI app. What Changed in the Stateless MCP Update? This deserves special attention because a lot of MCP tutorials online now describe the older lifecycle. Under older versions of the protocol, an HTTP client effectively did this: Client | | initialize v Server | | Mcp-Session-Id v Client | | later request + session id v Same logical session This created a multi-instance deployment that often needed sticky routing or shared session infrastructure. The 2026-07-28 protocol changes that. A modern MCP request is designed to be self-contained: Request 1 | v Server A Request 2 | v Server C Request 3 | v Server B No protocol session needs to tie those calls together. The MCP team explicitly describes this as moving from a bidirectional, stateful protocol core to a stateless request/response model. This makes ordinary load balancing much easier. Writing a Python Client Let's verify the server without depending on a third-party AI application. Create client.py : php import asyncio from mcp import Client async def main - None: async with Client "http://127.0.0.1:8000/mcp" as client: print "Protocol:", client.protocol version, tools = await client.list tools print "\nAvailable tools:" for tool in tools.tools: print "-", tool.name result = await client.call tool "search kb", { "query": "rate limit", "limit": 2, }, print "\nTool result:" if result.structured content: print result.structured content else: print result.content if name == " main ": asyncio.run main Run the server in one terminal: uv run python server.py Then run the client in another: uv run python client.py The v2 Client accepts an HTTP URL directly and automatically uses Streamable HTTP. It also exposes the negotiated protocol version, so with a current client/server pair you should see the modern protocol version reported by the connection. Your output should be similar to mine: Protocol: 2026-07-28 Available tools: - search kb {'result': {'id': 'api-rate-limit', 'title': 'Understanding API rate limits', 'body': 'API rate limits restrict the number of requests allowed within a time window. Clients should retry using exponential backoff after receiving a rate-limit response.'} } But What If My Application Actually Needs State? "Stateless protocol" does not mean your application can never maintain state. It means MCP itself no longer hides application state inside a protocol session. Suppose you were building a shopping server. Instead of relying on: MCP session 42 owns this basket you could expose: php @mcp.tool def create basket - dict str, str : basket id = create new basket return { "basket id": basket id } Then later: php @mcp.tool def add item basket id: str, product id: str, - dict: return add product basket id, product id, Now the model sees and passes: basket id explicitly. The MCP maintainers specifically recommend this explicit-handle pattern for application-level state under the new stateless protocol. It is a subtle but useful architectural shift: Old idea: protocol remembers state New idea: application owns state and identifiers travel explicitly Scaling the Server The stateless core becomes particularly valuable when you deploy multiple workers. For example: uvicorn server:app --workers 4 A modern MCP request can be handled by any worker because the protocol no longer requires it to return to the worker that handled a previous request. Conceptually: php +-- Worker 1 Client -- LB +-- Worker 2 +-- Worker 3 +-- Worker 4 There are additional considerations for advanced features such as multi-round-trip interactions, shared subscription events, authorization, and distributed state, but those are application architecture concerns rather than a requirement of basic MCP tool execution. Wrapping Up The practical shift here is smaller than the spec diff makes it look: you still decorate functions with @mcp.tool , you still return dicts and let the framework build the result, and you still run mcp.run to serve it. What's different is what happens underneath. You don't need any handshake to negotiate, no session to keep warm, no sticky routing to configure. Get comfortable with MCPServer , keep stdout clean, and the rest of the stateless spec mostly stays out of your way. \ Kanwal Mehreen\ https://www.linkedin.com/in/kanwal-mehreen1/ https://www.linkedin.com/in/kanwal-mehreen1/ is a machine learning engineer and a technical writer with a profound passion for data science and the intersection of AI with medicine. She co-authored the ebook "Maximizing Productivity with ChatGPT". As a Google Generation Scholar 2022 for APAC, she champions diversity and academic excellence. She's also recognized as a Teradata Diversity in Tech Scholar, Mitacs Globalink Research Scholar, and Harvard WeCode Scholar. Kanwal is an ardent advocate for change, having founded FEMCodes to empower women in STEM fields.