How I built a remote MCP server so AI agents can build product demos A developer built a remote Model Context Protocol (MCP) server for Demo My Product, a tool that turns screenshots into clickable product walkthroughs, allowing AI agents like Claude, Cursor and ChatGPT to create and publish demos directly. The server runs at a single URL using Streamable HTTP with OAuth-based authentication, and on the Pro plan it can screenshot a public URL and return the page's buttons and links so agents place hotspots on real elements instead of guessing coordinates from pixels. I build Demo My Product https://demomyproduct.com . It makes clickable product walkthroughs out of screenshots. You can click through an example without signing up: https://demomyproduct.com/d/zs2pfbuwu2 https://demomyproduct.com/d/zs2pfbuwu2 Most of the work in building a demo is repetitive. You upload screenshots and put them in order. You write a title for each step, draw a hotspot on the button that matters, and write a tooltip. Then you blur any email addresses before anyone sees them. It's structured, tedious work, and I wanted an AI agent to do it. So the product ships with a remote MCP Model Context Protocol server, and Claude, Cursor or ChatGPT can build and publish a demo for you. This post covers how it's put together and the decisions I made along the way. First, the data model in plain terms. A demo is a list of steps. Each step is an image with hotspots on it. A hotspot has a tooltip and a click action: go to the next step, jump to any step, open a link, or end the demo. You can blur areas of a step. The blur is rendered into the published image, not laid over it, so the original pixels never reach the viewer. A published demo gets a public link, an iframe embed, and a per-step funnel showing views, completion and where people drop off. The model is small and explicit, so an agent can reason about it well. "Put a hotspot on the main button of each screenshot and publish it" maps cleanly onto a handful of operations. A lot of MCP servers are local processes you install with npm or pip. I went with a remote server at a single URL: https://demomyproduct.com/mcp The agent works against your account , not your filesystem. Demos, screenshots and published links all live in the app, so the server belongs next to the app. Going remote also means: The server uses Streamable HTTP: JSON-RPC over POST to a single URL. A request without a token gets a standard MCP auth challenge: POST https://demomyproduct.com/mcp no token → 401 {"jsonrpc":"2.0","error":{"code":-32000,"message":"missing authorization header"}} → WWW-Authenticate: Bearer resource metadata="https://demomyproduct.com/.well-known/oauth-protected-resource/mcp", scope="demos" The protected-resource metadata points the client at the authorization server, https://demomyproduct.com/api/auth . That server publishes: registration endpoint , so clients can use openid profile offline access demos Sticking to the spec is what lets "paste a URL and it works" hold in every client. The first time an agent connects, your browser opens Demo My Product. You sign in or create an account , check the agent's name, and press Allow. I care about that screen showing the agent's name. You should know which client you're granting access to. Connecting from Claude Code is one line: claude mcp add --transport http demomyproduct https://demomyproduct.com/mcp Then run /mcp , pick demomyproduct and authenticate. In claude.ai or Claude Desktop, go to Settings → Connectors → Add custom connector. In Cursor, it's Settings → MCP. ChatGPT works in workspaces that have custom connectors enabled. An agent can do anything you can do in the editor. It can list and create demos, add screenshots, write step titles, place hotspots and tooltips, blur sensitive areas, change a demo's settings, and publish. Screenshots are the awkward part. I don't push image data through tool arguments. The agent asks for a one-time upload link and sends the image there PNG or JPEG, up to 10 MB . The server turns it into a step. This works for any page, including ones behind a login, because the agent supplies the image itself. On the Pro plan there's a second route. The agent gives the server a public URL. The server opens it in a browser, takes the screenshot, and returns the buttons and links on the page . That last part matters most. Without it, the agent would be guessing hotspot coordinates from pixels. With the element list, hotspots land on the actual button. Pages that need a sign-in can't be captured this way. For those, the agent falls back to uploading its own screenshots. The agent acts as you, with your plan's limits. It can publish as many demos as your plan allows. It can only capture pages on Pro. Publishing needs a verified email address. There are also rate limits: 120 calls a minute, 60 uploads an hour and 30 captures an hour. Everything the agent creates shows up in the normal editor. You can review and change it before or after publishing. That was deliberate. I don't want the agent's output to be a black box. It's a first draft you can fix by hand. The app is built on Next.js and uses Stripe for billing. These are the prompts I start from: The guide to connecting an agent is at https://demomyproduct.com/help/build-demos-with-ai-agents https://demomyproduct.com/help/build-demos-with-ai-agents . The free plan publishes one demo with no card needed, and the agent works on it. The feedback I want most is on the agent workflow. Which tools did you expect that aren't there? Where does the agent get things wrong? Leave a comment and I'll answer here.