cd /news/ai-agents/running-an-ai-agent-inside-the-brows… · home › topics › ai-agents › article
[ARTICLE · art-140674] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Running an AI Agent Inside the Browser with Pyodide and Ollama

A developer demonstrated running an AI agent loop inside a browser tab using Pyodide (CPython compiled to WebAssembly) for the agent logic and a local Ollama model as the reasoning backend, rather than driving a separate headless Chrome instance via Playwright. The experiment implements an MCP-style tool layer — named tools like read_dom, click_element, evaluate_js, and done called with JSON arguments — while explicitly noting it is not a full MCP protocol implementation. The author warns that the evaluate_js tool deliberately exposes arbitrary JavaScript execution with access to the page's session and credentials, which is acceptable only for local experiments on pages you own.

by read4 min views2 publishedSep 28, 2026

Browser automation is the default answer when an AI agent needs to use a web app: launch headless Chrome, drive it with Playwright, ferry state back and forth. It works, but the agent ends up living outside the browser, re-implementing sessions and rendering in parallel.

There's a less obvious option: run the agent inside the browser tab. Pyodide gives us CPython in WebAssembly, the page's DOM and network are directly reachable, and a local model such as Ollama can act as the reasoning backend. This post walks through a small working experiment along those lines, inspired by the tool-call pattern of MCP (Model Context Protocol).

One clarification up front, because it matters: what we build here is an MCP-style tool layer, not a full MCP protocol implementation. The official MCP SDK defines tools, resources, and prompts over standard transports (stdio, Streamable HTTP, SSE). We borrow the "model calls named tools with structured arguments" idea and implement it with a simple JSON message shape. If you want wire-level MCP in the browser, that's a follow-up project, not this one.

Driving a browser from Python looks like this:

from playwright.sync_api import sync_playwright

def get_page_text(url):
    with sync_playwright() as p:
        browser = p.chromium.launch()
        page = browser.new_page()
        page.goto(url)
        text = page.inner_text("body")
        browser.close()
        return text

Every interaction is a round-trip between two processes. Session state (logins, cookies, rendered state) lives in the browser your user is actually looking at, while the agent operates on a separate, synthetic copy. Keeping those two worlds in sync is where the bugs and the latency come from.

Flip the arrangement: the agent loop runs in the tab, and the page's own APIs become its tools. The architecture:

read_dom, click_element) as named tools the model can call with JSON arguments — the same shape MCP popularized. Model inference does not happen in the browser — the agent loop does. The division of labor is: reasoning in Ollama, perception and action in the tab.

Create tools.py:

import json
import js

class ToolRegistry:
    """MCP-style named tools, callable via a simple JSON protocol."""

    def __init__(self):
        self.tools = {
            "read_dom": self.read_dom,
            "click_element": self.click_element,
            "evaluate_js": self.evaluate_js,
            "done": self.done,
        }

    def read_dom(self, selector="body"):
        element = js.document.querySelector(selector)
        if not element:
            return f"No element found for selector: {selector}"
        return element.innerText

    def click_element(self, selector):
        element = js.document.querySelector(selector)
        if not element:
            return f"No element found for selector: {selector}"
        element.click()
        return f"Clicked {selector}"

    def done(self, message=""):
        """Signal that the task is complete."""
        return f"TASK_DONE: {message}"

    def evaluate_js(self, code):
        try:
            return str(js.eval(code))
        except Exception as e:
            return f"JS error: {e}"

    def handle_request(self, request_json):
        request = json.loads(request_json)
        name, args = request.get("tool"), request.get("args", {})
        if name not in self.tools:
            return json.dumps({"error": f"Unknown tool: {name}"})
        try:
            return json.dumps({"result": self.tools[name](**args)})
        except Exception as e:
            return json.dumps({"error": str(e)})

registry = ToolRegistry()
js.window.toolRegistry = registry

Security note: evaluate_js deliberately exposes arbitrary JavaScript execution, including access to the page's session and credentials context. That's acceptable for a local experiment on pages you own. A real deployment should restrict the operation set (whitelist specific actions) and never expose unrestricted eval — the same way a production MCP server validates and sanitizes every tool input.

Create agent.py:

import json
import js
from pyodide.http import pyfetch

async def call_ollama(prompt):
    response = await pyfetch(
        "http://localhost:11434/api/generate",
        method="POST",
        headers={"Content-Type": "application/json"},
        body=json.dumps({
            "model": "llama3.2",
            "prompt": prompt,
            "stream": False,
            "format": "json",
        }),
    )
    data = await response.json()
    return data["response"]

async def agent_step(task):
    dom_text = js.window.toolRegistry.read_dom("body")
    prompt = f"""You are a browser agent. Task: {task}
Current page content:
{dom_text[:2000]}

Respond with JSON: {{"tool": "read_dom"|"click_element"|"evaluate_js"|"done", "args": {{...}}}}
Use "done" when the task is complete.
"""
    action = json.loads(await call_ollama(prompt))
    return js.window.toolRegistry.handle_request(json.dumps(action))

async def run_agent(task, max_steps=5):
    for _ in range(max_steps):
        result = await agent_step(task)
        print("Result:", result)
        if "TASK_DONE" in result:
            break

Two things have to be true for this to work:

http://localhost:11434 is a cross-origin request, and Ollama rejects those by default. Start it with an allowlist: OLLAMA_ORIGINS="http://localhost:8000" ollama serve (use your page's actual origin; * works for local experiments only).http://localhost rather than opening it as file://, otherwise some browsers block the call. In your HTML page:

<script src="https://cdn.jsdelivr.net/pyodide/v0.26.2/full/pyodide.js"></script>
<script>
async function main() {
  const pyodide = await loadPyodide();
  await pyodide.runPythonAsync(await (await fetch('tools.py')).text());
  await pyodide.runPythonAsync(await (await fetch('agent.py')).text());
  await pyodide.runPythonAsync("import asyncio; asyncio.ensure_future(run_agent('find the pricing link'))");
}
main();
</script>

Run any static file server (python -m http.server 8000), start Ollama with the origins allowlist, open the page, and watch the tab operate on itself.

eval as a model-callable tool is a local-experiment-only feature. Production versions whitelist operations and validate arguments.

── more in #ai-agents 4 stories · sorted by recency
── more on @pyodide 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/running-an-ai-agent-…] indexed:0 read:4min 2026-09-28 · —