Browser automation is the default answer when an AI agent needs to use a web app: launch headless Chrome, drive it with Playwright, ferry state back and forth. It works, but the agent ends up living outside the browser, re-implementing sessions and rendering in parallel.
There's a less obvious option: run the agent inside the browser tab. Pyodide gives us CPython in WebAssembly, the page's DOM and network are directly reachable, and a local model such as Ollama can act as the reasoning backend. This post walks through a small working experiment along those lines, inspired by the tool-call pattern of MCP (Model Context Protocol).
One clarification up front, because it matters: what we build here is an MCP-style tool layer, not a full MCP protocol implementation. The official MCP SDK defines tools, resources, and prompts over standard transports (stdio, Streamable HTTP, SSE). We borrow the "model calls named tools with structured arguments" idea and implement it with a simple JSON message shape. If you want wire-level MCP in the browser, that's a follow-up project, not this one.
Driving a browser from Python looks like this:
from playwright.sync_api import sync_playwright
def get_page_text(url):
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto(url)
text = page.inner_text("body")
browser.close()
return text
Every interaction is a round-trip between two processes. Session state (logins, cookies, rendered state) lives in the browser your user is actually looking at, while the agent operates on a separate, synthetic copy. Keeping those two worlds in sync is where the bugs and the latency come from.
Flip the arrangement: the agent loop runs in the tab, and the page's own APIs become its tools. The architecture:
read_dom, click_element) as named tools the model can call with JSON arguments — the same shape MCP popularized.
Model inference does not happen in the browser — the agent loop does. The division of labor is: reasoning in Ollama, perception and action in the tab.
Create tools.py:
import json
import js
class ToolRegistry:
"""MCP-style named tools, callable via a simple JSON protocol."""
def __init__(self):
self.tools = {
"read_dom": self.read_dom,
"click_element": self.click_element,
"evaluate_js": self.evaluate_js,
"done": self.done,
}
def read_dom(self, selector="body"):
element = js.document.querySelector(selector)
if not element:
return f"No element found for selector: {selector}"
return element.innerText
def click_element(self, selector):
element = js.document.querySelector(selector)
if not element:
return f"No element found for selector: {selector}"
element.click()
return f"Clicked {selector}"
def done(self, message=""):
"""Signal that the task is complete."""
return f"TASK_DONE: {message}"
def evaluate_js(self, code):
try:
return str(js.eval(code))
except Exception as e:
return f"JS error: {e}"
def handle_request(self, request_json):
request = json.loads(request_json)
name, args = request.get("tool"), request.get("args", {})
if name not in self.tools:
return json.dumps({"error": f"Unknown tool: {name}"})
try:
return json.dumps({"result": self.tools[name](**args)})
except Exception as e:
return json.dumps({"error": str(e)})
registry = ToolRegistry()
js.window.toolRegistry = registry
Security note: evaluate_js deliberately exposes arbitrary JavaScript execution, including access to the page's session and credentials context. That's acceptable for a local experiment on pages you own. A real deployment should restrict the operation set (whitelist specific actions) and never expose unrestricted eval — the same way a production MCP server validates and sanitizes every tool input.
Create agent.py:
import json
import js
from pyodide.http import pyfetch
async def call_ollama(prompt):
response = await pyfetch(
"http://localhost:11434/api/generate",
method="POST",
headers={"Content-Type": "application/json"},
body=json.dumps({
"model": "llama3.2",
"prompt": prompt,
"stream": False,
"format": "json",
}),
)
data = await response.json()
return data["response"]
async def agent_step(task):
dom_text = js.window.toolRegistry.read_dom("body")
prompt = f"""You are a browser agent. Task: {task}
Current page content:
{dom_text[:2000]}
Respond with JSON: {{"tool": "read_dom"|"click_element"|"evaluate_js"|"done", "args": {{...}}}}
Use "done" when the task is complete.
"""
action = json.loads(await call_ollama(prompt))
return js.window.toolRegistry.handle_request(json.dumps(action))
async def run_agent(task, max_steps=5):
for _ in range(max_steps):
result = await agent_step(task)
print("Result:", result)
if "TASK_DONE" in result:
break
Two things have to be true for this to work:
http://localhost:11434 is a cross-origin request, and Ollama rejects those by default. Start it with an allowlist: OLLAMA_ORIGINS="http://localhost:8000" ollama serve (use your page's actual origin; * works for local experiments only).http://localhost rather than opening it as file://, otherwise some browsers block the call.
In your HTML page:
<script src="https://cdn.jsdelivr.net/pyodide/v0.26.2/full/pyodide.js"></script>
<script>
async function main() {
const pyodide = await loadPyodide();
await pyodide.runPythonAsync(await (await fetch('tools.py')).text());
await pyodide.runPythonAsync(await (await fetch('agent.py')).text());
await pyodide.runPythonAsync("import asyncio; asyncio.ensure_future(run_agent('find the pricing link'))");
}
main();
</script>
Run any static file server (python -m http.server 8000), start Ollama with the origins allowlist, open the page, and watch the tab operate on itself.
eval as a model-callable tool is a local-experiment-only feature. Production versions whitelist operations and validate arguments.