Running an AI Agent Inside the Browser with Pyodide and Ollama A developer demonstrated running an AI agent loop inside a browser tab using Pyodide (CPython compiled to WebAssembly) for the agent logic and a local Ollama model as the reasoning backend, rather than driving a separate headless Chrome instance via Playwright. The experiment implements an MCP-style tool layer — named tools like read_dom, click_element, evaluate_js, and done called with JSON arguments — while explicitly noting it is not a full MCP protocol implementation. The author warns that the evaluate_js tool deliberately exposes arbitrary JavaScript execution with access to the page's session and credentials, which is acceptable only for local experiments on pages you own. Browser automation is the default answer when an AI agent needs to use a web app: launch headless Chrome, drive it with Playwright, ferry state back and forth. It works, but the agent ends up living outside the browser, re-implementing sessions and rendering in parallel. There's a less obvious option: run the agent inside the browser tab. Pyodide gives us CPython in WebAssembly, the page's DOM and network are directly reachable, and a local model such as Ollama can act as the reasoning backend. This post walks through a small working experiment along those lines, inspired by the tool-call pattern of MCP Model Context Protocol . One clarification up front, because it matters: what we build here is an MCP-style tool layer, not a full MCP protocol implementation. The official MCP SDK defines tools, resources, and prompts over standard transports stdio, Streamable HTTP, SSE . We borrow the "model calls named tools with structured arguments" idea and implement it with a simple JSON message shape. If you want wire-level MCP in the browser, that's a follow-up project, not this one. Driving a browser from Python looks like this: python from playwright.sync api import sync playwright def get page text url : with sync playwright as p: browser = p.chromium.launch page = browser.new page page.goto url text = page.inner text "body" browser.close return text Every interaction is a round-trip between two processes. Session state logins, cookies, rendered state lives in the browser your user is actually looking at, while the agent operates on a separate, synthetic copy. Keeping those two worlds in sync is where the bugs and the latency come from. Flip the arrangement: the agent loop runs in the tab, and the page's own APIs become its tools. The architecture: read dom , click element as named tools the model can call with JSON arguments — the same shape MCP popularized. Model inference does not happen in the browser — the agent loop does. The division of labor is: reasoning in Ollama, perception and action in the tab. Create tools.py : python import json import js class ToolRegistry: """MCP-style named tools, callable via a simple JSON protocol.""" def init self : self.tools = { "read dom": self.read dom, "click element": self.click element, "evaluate js": self.evaluate js, "done": self.done, } def read dom self, selector="body" : element = js.document.querySelector selector if not element: return f"No element found for selector: {selector}" return element.innerText def click element self, selector : element = js.document.querySelector selector if not element: return f"No element found for selector: {selector}" element.click return f"Clicked {selector}" def done self, message="" : """Signal that the task is complete.""" return f"TASK DONE: {message}" def evaluate js self, code : try: return str js.eval code except Exception as e: return f"JS error: {e}" def handle request self, request json : request = json.loads request json name, args = request.get "tool" , request.get "args", {} if name not in self.tools: return json.dumps {"error": f"Unknown tool: {name}"} try: return json.dumps {"result": self.tools name args } except Exception as e: return json.dumps {"error": str e } registry = ToolRegistry js.window.toolRegistry = registry Security note: evaluate js deliberately exposes arbitrary JavaScript execution, including access to the page's session and credentials context. That's acceptable for a local experiment on pages you own. A real deployment should restrict the operation set whitelist specific actions and never expose unrestricted eval — the same way a production MCP server validates and sanitizes every tool input. Create agent.py : python import json import js from pyodide.http import pyfetch async def call ollama prompt : response = await pyfetch "http://localhost:11434/api/generate", method="POST", headers={"Content-Type": "application/json"}, body=json.dumps { "model": "llama3.2", "prompt": prompt, "stream": False, "format": "json", } , data = await response.json return data "response" async def agent step task : dom text = js.window.toolRegistry.read dom "body" prompt = f"""You are a browser agent. Task: {task} Current page content: {dom text :2000 } Respond with JSON: {{"tool": "read dom"|"click element"|"evaluate js"|"done", "args": {{...}}}} Use "done" when the task is complete. """ action = json.loads await call ollama prompt return js.window.toolRegistry.handle request json.dumps action async def run agent task, max steps=5 : for in range max steps : result = await agent step task print "Result:", result if "TASK DONE" in result: break Two things have to be true for this to work: http://localhost:11434 is a cross-origin request, and Ollama rejects those by default. Start it with an allowlist: OLLAMA ORIGINS="http://localhost:8000" ollama serve use your page's actual origin; works for local experiments only . http://localhost rather than opening it as file:// , otherwise some browsers block the call. In your HTML page: