{"slug": "running-an-ai-agent-inside-the-browser-with-pyodide-and-ollama", "title": "Running an AI Agent Inside the Browser with Pyodide and Ollama", "summary": "A developer demonstrated running an AI agent loop inside a browser tab using Pyodide (CPython compiled to WebAssembly) for the agent logic and a local Ollama model as the reasoning backend, rather than driving a separate headless Chrome instance via Playwright. The experiment implements an MCP-style tool layer — named tools like read_dom, click_element, evaluate_js, and done called with JSON arguments — while explicitly noting it is not a full MCP protocol implementation. The author warns that the evaluate_js tool deliberately exposes arbitrary JavaScript execution with access to the page's session and credentials, which is acceptable only for local experiments on pages you own.", "body_md": "Browser automation is the default answer when an AI agent needs to use a web app: launch headless Chrome, drive it with Playwright, ferry state back and forth. It works, but the agent ends up living *outside* the browser, re-implementing sessions and rendering in parallel.\n\nThere's a less obvious option: run the agent *inside* the browser tab. Pyodide gives us CPython in WebAssembly, the page's DOM and network are directly reachable, and a local model such as Ollama can act as the reasoning backend. This post walks through a small working experiment along those lines, inspired by the tool-call pattern of MCP (Model Context Protocol).\n\nOne clarification up front, because it matters: **what we build here is an MCP-style tool layer, not a full MCP protocol implementation.** The official MCP SDK defines tools, resources, and prompts over standard transports (stdio, Streamable HTTP, SSE). We borrow the \"model calls named tools with structured arguments\" idea and implement it with a simple JSON message shape. If you want wire-level MCP in the browser, that's a follow-up project, not this one.\n\nDriving a browser from Python looks like this:\n\n``` python\nfrom playwright.sync_api import sync_playwright\n\ndef get_page_text(url):\n    with sync_playwright() as p:\n        browser = p.chromium.launch()\n        page = browser.new_page()\n        page.goto(url)\n        text = page.inner_text(\"body\")\n        browser.close()\n        return text\n```\n\nEvery interaction is a round-trip between two processes. Session state (logins, cookies, rendered state) lives in the browser your user is actually looking at, while the agent operates on a separate, synthetic copy. Keeping those two worlds in sync is where the bugs and the latency come from.\n\nFlip the arrangement: the agent loop runs in the tab, and the page's own APIs become its tools. The architecture:\n\n`read_dom`, `click_element`) as named tools the model can call with JSON arguments — the same shape MCP popularized.\nModel inference does *not* happen in the browser — the agent loop does. The division of labor is: reasoning in Ollama, perception and action in the tab.\n\nCreate `tools.py`:\n\n``` python\nimport json\nimport js\n\nclass ToolRegistry:\n    \"\"\"MCP-style named tools, callable via a simple JSON protocol.\"\"\"\n\n    def __init__(self):\n        self.tools = {\n            \"read_dom\": self.read_dom,\n            \"click_element\": self.click_element,\n            \"evaluate_js\": self.evaluate_js,\n            \"done\": self.done,\n        }\n\n    def read_dom(self, selector=\"body\"):\n        element = js.document.querySelector(selector)\n        if not element:\n            return f\"No element found for selector: {selector}\"\n        return element.innerText\n\n    def click_element(self, selector):\n        element = js.document.querySelector(selector)\n        if not element:\n            return f\"No element found for selector: {selector}\"\n        element.click()\n        return f\"Clicked {selector}\"\n\n    def done(self, message=\"\"):\n        \"\"\"Signal that the task is complete.\"\"\"\n        return f\"TASK_DONE: {message}\"\n\n    def evaluate_js(self, code):\n        try:\n            return str(js.eval(code))\n        except Exception as e:\n            return f\"JS error: {e}\"\n\n    def handle_request(self, request_json):\n        request = json.loads(request_json)\n        name, args = request.get(\"tool\"), request.get(\"args\", {})\n        if name not in self.tools:\n            return json.dumps({\"error\": f\"Unknown tool: {name}\"})\n        try:\n            return json.dumps({\"result\": self.tools[name](**args)})\n        except Exception as e:\n            return json.dumps({\"error\": str(e)})\n\nregistry = ToolRegistry()\njs.window.toolRegistry = registry\n```\n\n**Security note:** `evaluate_js` deliberately exposes arbitrary JavaScript execution, including access to the page's session and credentials context. That's acceptable for a local experiment on pages you own. A real deployment should restrict the operation set (whitelist specific actions) and never expose unrestricted `eval` — the same way a production MCP server validates and sanitizes every tool input.\n\nCreate `agent.py`:\n\n``` python\nimport json\nimport js\nfrom pyodide.http import pyfetch\n\nasync def call_ollama(prompt):\n    response = await pyfetch(\n        \"http://localhost:11434/api/generate\",\n        method=\"POST\",\n        headers={\"Content-Type\": \"application/json\"},\n        body=json.dumps({\n            \"model\": \"llama3.2\",\n            \"prompt\": prompt,\n            \"stream\": False,\n            \"format\": \"json\",\n        }),\n    )\n    data = await response.json()\n    return data[\"response\"]\n\nasync def agent_step(task):\n    dom_text = js.window.toolRegistry.read_dom(\"body\")\n    prompt = f\"\"\"You are a browser agent. Task: {task}\nCurrent page content:\n{dom_text[:2000]}\n\nRespond with JSON: {{\"tool\": \"read_dom\"|\"click_element\"|\"evaluate_js\"|\"done\", \"args\": {{...}}}}\nUse \"done\" when the task is complete.\n\"\"\"\n    action = json.loads(await call_ollama(prompt))\n    return js.window.toolRegistry.handle_request(json.dumps(action))\n\nasync def run_agent(task, max_steps=5):\n    for _ in range(max_steps):\n        result = await agent_step(task)\n        print(\"Result:\", result)\n        if \"TASK_DONE\" in result:\n            break\n```\n\nTwo things have to be true for this to work:\n\n`http://localhost:11434` is a cross-origin request, and Ollama rejects those by default. Start it with an allowlist: `OLLAMA_ORIGINS=\"http://localhost:8000\" ollama serve` (use your page's actual origin; `*` works for local experiments only).`http://localhost` rather than opening it as `file://`, otherwise some browsers block the call.\nIn your HTML page:\n\n```\n<script src=\"https://cdn.jsdelivr.net/pyodide/v0.26.2/full/pyodide.js\"></script>\n<script>\nasync function main() {\n  const pyodide = await loadPyodide();\n  await pyodide.runPythonAsync(await (await fetch('tools.py')).text());\n  await pyodide.runPythonAsync(await (await fetch('agent.py')).text());\n  await pyodide.runPythonAsync(\"import asyncio; asyncio.ensure_future(run_agent('find the pricing link'))\");\n}\nmain();\n</script>\n```\n\nRun any static file server (`python -m http.server 8000`), start Ollama with the origins allowlist, open the page, and watch the tab operate on itself.\n\n`eval` as a model-callable tool is a local-experiment-only feature. Production versions whitelist operations and validate arguments.", "url": "https://wpnews.pro/news/running-an-ai-agent-inside-the-browser-with-pyodide-and-ollama", "canonical_source": "https://dev.to/gu_cci_f94bedb90083e6aab4/running-an-ai-agent-inside-the-browser-with-pyodide-and-ollama-g5g", "published_at": "2026-09-28 00:10:35+00:00", "updated_at": "2026-09-28 00:30:59.580841+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "ai-tools", "developer-tools"], "entities": ["Pyodide", "Ollama", "Playwright", "Model Context Protocol", "WebAssembly", "Chrome"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/running-an-ai-agent-inside-the-browser-with-pyodide-and-ollama", "markdown": "https://wpnews.pro/news/running-an-ai-agent-inside-the-browser-with-pyodide-and-ollama.md", "text": "https://wpnews.pro/news/running-an-ai-agent-inside-the-browser-with-pyodide-and-ollama.txt", "jsonld": "https://wpnews.pro/news/running-an-ai-agent-inside-the-browser-with-pyodide-and-ollama.jsonld"}}