{"slug": "mini-web-agent-web-agents-from-scratch-in-under-400-lines", "title": "Mini-web-agent: Web agents from scratch in under 400 lines", "summary": "A developer released mini-web-agent, an open-source Python project that implements a browser-controlling web agent in 396 lines of code, published on GitHub and installable via PyPI as mini-web-agent. The project gives a model screenshots and browser controls through Playwright and the OpenAI SDK, and its demo uses Google's Gemini 3.8 Flash to book a Robotics Lab workshop and verify the confirmation. The code is split into imports (17 lines), 24 browser and conversation action functions (94 lines), helpers (72 lines), the WebAgent class (124 lines), run() (51 lines), and the CLI (38 lines).", "body_md": "We wanted to understand how web agents in tools like Dots, Muse, and Codex work: how frontier models see a page, choose actions, and decide when a task is done.\n\n**mini-web-agent** explores that loop in fewer than 400 lines of Python. This makes it easier\nto follow the code and see how changes affect the agent. Give a model screenshots and\nbrowser controls, then follow along as it works through a task.\n\n## Why keep it small?\n\n[mini-swe-agent](https://github.com/SWE-agent/mini-swe-agent#readme) asks how much work a\ncapable model can do with a small agent around it. We wanted to explore that idea in a\nbrowser, with screenshots and the same controls you use to navigate a website.\n\nA short implementation gives you room to experiment. Change the instructions, try another model, or add an action. You can follow the run and see how those changes affect it.\n\n| *Gemini 3.8 Flash uses screenshots and browser controls to book a Robotics Lab workshop and verify the confirmation. Pauses between actions are shortened. [Action log](https://github.com/xhluca/mini-web-agent/blob/main/demo/demo.json).* | \n|---|\n\nThe browser controls, model loop, and command line entry point all fit in this file:\n\n| Section | What it does | Lines | \n|---|---|---|\n| [Imports](https://github.com/xhluca/mini-web-agent/blob/main/agent.py#L1) | Standard library, OpenAI SDK, and Playwright | 17 | \n| [Actions](https://github.com/xhluca/mini-web-agent/blob/main/agent.py#L18) | 24 browser and conversation functions | 94 | \n| [Helpers](https://github.com/xhluca/mini-web-agent/blob/main/agent.py#L112) | Prompts, browser setup, tab lookup, tool schemas, and screenshot formatting | 72 | \n| [WebAgent](https://github.com/xhluca/mini-web-agent/blob/main/agent.py#L184) | Chrome lifecycle, tabs, actions, and screenshots | 124 | \n| [run()](https://github.com/xhluca/mini-web-agent/blob/main/agent.py#L308) | Model calls, results, callbacks, history, and recovery | 51 | \n| [CLI](https://github.com/xhluca/mini-web-agent/blob/main/agent.py#L359) | Options, browser setup, model run, and cleanup | 38 | \n| **Total** |  | **396** | \n\n## Record your own demo\n\nOnce you have completed the setup below, install [FFmpeg](https://ffmpeg.org/download.html)\nand record a run:\n\n```\nuv run demo/record.py --model google/gemini-3.8-flash\n```\n\nThe agent works through a booking on a local page while the recorder captures the browser\nand logs its actions. It checks the final confirmation, then saves `demo/demo.gif` and\n`demo/demo.json`. FFmpeg is only needed to make the GIF.\n\nStart with a task you can watch from beginning to end. With\n[uv](https://docs.astral.sh/uv/getting-started/installation/) installed:\n\n```\ngit clone https://github.com/xhluca/mini-web-agent.git\ncd mini-web-agent\nuv run playwright install chromium --no-shell\n\nexport OPENAI_API_KEY=\"your-openrouter-key\"\nexport OPENAI_BASE_URL=https://openrouter.ai/api/v1\n\nuv run agent.py --headed --cursor --model google/gemini-3.8-flash \\\n  'Open https://example.com and tell me the heading.'\n```\n\nuv handles the Python environment, dependencies, and Chromium installation. Chrome opens\nin a window, with an animated cursor showing where the agent moves and clicks. Actions,\nresults, and questions appear in your terminal. Press **Ctrl+C** to interrupt.\n\nReplace the task in quotes with something you want to try. Leave out `--headed` to run\nChrome headless and `--cursor` to hide the cursor. The project runs on Linux and macOS.\n\n## Install from PyPI\n\nTo run the agent without cloning the source:\n\n```\npython -m pip install mini-web-agent\npython -m playwright install chromium --no-shell\n\nexport OPENAI_API_KEY=\"your-openrouter-key\"\nexport OPENAI_BASE_URL=https://openrouter.ai/api/v1\n\nmini-web-agent --headed --cursor --model google/gemini-3.8-flash \\\n  'Open https://example.com and tell me the heading.'\n```\n\nThe installed command uses the same CLI as `python agent.py`. You can also run\n`python -u -m agent` with the same options.\n\n## Manual setup with venv\n\nIf you prefer to set up Python yourself, use Python 3.10+ and install the project this way:\n\n```\ngit clone https://github.com/xhluca/mini-web-agent.git\ncd mini-web-agent\npython3 -m venv .venv\nsource .venv/bin/activate\npython -m pip install .\npython -m playwright install chromium --no-shell\n\nexport OPENAI_API_KEY=\"your-openrouter-key\"\nexport OPENAI_BASE_URL=https://openrouter.ai/api/v1\n\npython -u agent.py --headed --cursor --model google/gemini-3.8-flash \\\n  'Open https://example.com and tell me the heading.'\n```\n\nYou will see the same browser window and terminal output as in the uv example.\n\n## CLI options and connecting to Chrome\n\nAs you try longer tasks, you can change the model, set a turn limit, or keep using a browser that is already open:\n\n| Option | Purpose | \n|---|---|\n| `--model` | Required model ID, e.g. `google/gemini-3.8-flash` | \n| `--max-steps` | Maximum model turns; defaults to 100 | \n| `--profile` | Persistent Chrome profile; defaults to `.chrome` | \n| `--port` | Launch port; a new profile defaults to a random localhost port | \n| `--connect` | Attach to Chrome already running with the selected profile | \n| `--headed` | Show Chrome; new launches are headless by default | \n| `--cursor` | Animate pointer actions before execution | \n\nTo run another task from the project directory, replace the text in quotes:\n\n```\nuv run agent.py --headed --cursor --model google/gemini-3.8-flash \\\n  'Open https://example.com and tell me the heading.'\n```\n\nTo try another provider, set `OPENAI_API_KEY` and `OPENAI_BASE_URL` for its endpoint.\nThe endpoint must support the Responses API, and the model needs image input and function calling.\n\nIf Chrome is still running from an earlier session, connect using the same profile:\n\n```\nuv run agent.py --connect --profile .chrome --model google/gemini-3.8-flash \\\n  'Tell me what is open in the current tab.'\n```\n\nYou pick up the browser with its existing tabs and display mode. Adding `--headed` prints\na warning and continues, since connecting cannot change how Chrome was launched.\nThe CLI closes browsers it launches. When you use `--connect`, it leaves Chrome running.\n\nWhen launching Chrome, the CLI replaces restored tabs with one blank tab and keeps profile data.\nUse `--port 9222` for a fixed launch port. With `--port 0`, Chrome chooses a port for a new\nprofile. Existing profiles reuse their recorded port. `--connect` reads the\nprofile's address and cannot be combined with a nonzero `--port`.\nThe CLI prints the CDP address after connecting.\n\nOn minimal Linux systems, Chromium may also need system dependencies. Install them with\n`uv run playwright install-deps chromium`.\n\nAt each turn, the model gets a screenshot and a list of open tabs. It chooses from the\nfunctions in `Actions`: click, type, scroll, open a tab, or talk to the user.\n\nPython calls those functions, sends back the results, and takes another screenshot.\nThe model decides what to try next. When it considers the task complete, it calls `finish`\nand the loop returns its answer.\n\nYou can follow this cycle in `run()`. The same function is used by the CLI, Python example,\nand demo recorder.\n\n## What goes to the model?\n\nThe default instructions include the source of `agent.py`, so the model can read the\nfunctions it may call. Their signatures and docstrings also provide the tool definitions\nsent to the API.\n\nThe implementation has two dependencies: Playwright for browser control and the OpenAI SDK for model requests. Requests use an OpenAI-compatible Responses API; the quick start uses OpenRouter. Models need to support both images and function calls.\n\nThe first screenshot is sent as a user message. After each batch of actions, the last tool result includes a new screenshot and the current tab information. Earlier screenshots and reasoning stay in history, so you can follow what the model received at each turn. Appending to that history also lets providers cache the unchanged prefix. Longer tasks use more context, and cache hits depend on the provider.\n\n## Available actions\n\nThe model has 24 actions to choose from. They are ordinary functions grouped in `Actions`,\nso you can read exactly what each one does:\n\n| Interaction | Actions | \n|---|---|\n| Navigation | `navigate` ,`back` ,`forward` ,`reload` | \n| Pointer | `click` ,`double_click` ,`right_click` ,`hover` ,`mouse_down` ,`mouse_up` ,`drag` | \n| Scroll and keyboard | `scroll` ,`type_text` ,`press_key` ,`key_down` ,`key_up` | \n| Tabs | `list_tabs` ,`new_tab` ,`switch_tab` ,`close_tab` | \n| Timing and conversation | `wait` ,`send_message` ,`wait_for_reply` ,`finish` | \n\nClicks and pointer movements use coordinates from the screenshot. The viewport defaults to\n1280×800; pass `w` and `h` to `WebAgent` to change it. `type_text` types into the focused field,\nwith 10 ms between characters. Tab indices come from the latest observation and can shift\nafter closing a tab.\n\n## When a step goes wrong\n\nA failed action becomes feedback for the next turn. Each action returns one of:\n\n```\n{\"state\": \"success\", \"output\": ...}\n{\"state\": \"error\", \"output\": \"...\"}\n```\n\nThe model sees the error and can try another action. If it returns no tool calls, the loop asks it to choose one. Incomplete responses are discarded and retried before any of their actions execute.\n\nEach model turn counts toward `max_steps`, including retries. At the limit, the loop stops\nwith an unfinished-task message. A successful `finish` returns the model's final answer.\n\n## What can it see and control?\n\nThe model sees screenshots and tab information. It uses the browser controls in the action dictionary to interact with pages. It has no tools for reading the DOM, parsing accessibility trees, running generated code, managing uploads or downloads, or controlling the OS.\n\nThe action dictionary limits which functions the model can call. The browser can still visit websites and interact with them normally.\n\nTo experiment with the loop from your own code, call it as a Python function. Keep the API environment variables from the quick start set, then:\n\n``` python\nfrom openai import OpenAI\nfrom agent import WebAgent, get_action_space, get_instructions, run\n\nagent = WebAgent(\".chrome\", action_space=get_action_space())\ntry:\n    agent.launch().connect()\n    with OpenAI(timeout=60, max_retries=1) as client:\n        answer = run(\n            agent, \"Find the heading on https://example.com.\", client,\n            model=\"google/gemini-3.8-flash\", instructions=get_instructions(),\n            max_steps=20, callbacks=[dict(type=\"after\", function=print)],\n        )\n        print(answer)\nfinally:\n    agent.shutdown()\n```\n\nUse `agent.launch(headed=True)` for a visible browser. To leave Chrome running for another task,\ncall `agent.disconnect()` instead of `agent.shutdown()`.\n\n## Browser lifecycle\n\nChrome runs as a separate process, so you can leave it open between tasks. Starting Chrome, connecting to it, and closing it are separate operations:\n\n| Function | What it does | \n|---|---|\n| `agent.launch()` | Start detached Chrome | \n| `agent.connect()` | Attach Playwright to Chrome | \n| `agent.ctx` | Browser context assigned when connecting | \n| `agent.get_page()` | Get the active tab and set its viewport | \n| `agent.observe()` | Return tab metadata as JSON and a screenshot data URL | \n| `agent.act(name, arguments)` | Execute an allowed action and return its result | \n| `run(agent, task, client, model, instructions, ...)` | Run the model loop | \n| `agent.disconnect()` | Detach Playwright, leaving Chrome running | \n| `agent.shutdown()` | Close Chrome and disconnect, even if Chrome was already running when you connected | \n\nUse `port=9222` for a fixed port, or leave it at `0` to let the agent choose. After launch,\n`agent.port` holds the port Chrome is using. To return to this browser later, create another\n`WebAgent` with the same profile and call `connect()`.\n\nPlaywright connects through Chrome's debugging interface, CDP. It finds the current WebSocket address over HTTP, then uses that connection to control the browser.\n\n## Customize actions, instructions, and callbacks\n\nStart by changing what you tell the model or what you let it do. `action_space` and\n`instructions` are required, so each run uses the actions and prompt you supply.\n`get_action_space()` and `get_instructions()` provide the defaults used in the example.\n\nThe action dictionary determines which functions the model can call and which tool definitions are sent to the API. Add or remove a function there to change its options.\n\nCallbacks let you watch an action before or after it happens. Each receives\n`(step, action, result)`: the model-turn index, the action's name and arguments, and its\nresult. Before execution, `result` is `None`. For example, show the cursor before an\naction and print the result afterward:\n\n``` python\nfrom functools import partial\nfrom callbacks.cursor import show_cursor\n\ncallbacks = [\n    dict(type=\"before\", function=partial(show_cursor, agent)),\n    dict(type=\"after\", function=print),\n]\n```\n\nThe [cursor.py](https://github.com/xhluca/mini-web-agent/blob/main/callbacks/cursor.py) overlay\nglides between targets, follows drags, and pulses on\nclicks. It respects reduced-motion preferences and lets clicks pass through to the page.\nThe CLI loads it when you pass `--cursor`.\n\nWhen the model needs to talk to you, `on_message` displays its message and `on_reply`\ncollects your answer. They default to `print` and `input`; replace them to use your own UI.\n\nIf you want to change the loop, start with\n[agent.py](https://github.com/xhluca/mini-web-agent/blob/main/agent.py). The animated cursor lives\nin [callbacks/cursor.py](https://github.com/xhluca/mini-web-agent/blob/main/callbacks/cursor.py), and the\n[demo recorder](https://github.com/xhluca/mini-web-agent/blob/main/demo/record.py) shows how to record a run.\nTests live in [tests/](https://github.com/xhluca/mini-web-agent/tree/main/tests):\n\n```\nuv run python -m unittest discover -s tests -v\nuv run python -m tests.test_live  # Optional paid OpenRouter test.\n```\n\nThe local tests exercise browser actions, tabs, cleanup, API results, and error recovery using Chromium and a test server. The live test asks a model to complete a signup form, then checks what it submitted.", "url": "https://wpnews.pro/news/mini-web-agent-web-agents-from-scratch-in-under-400-lines", "canonical_source": "https://github.com/xhluca/mini-web-agent", "published_at": "2026-10-05 17:30:02+00:00", "updated_at": "2026-10-05 17:50:25.295461+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models"], "entities": ["mini-web-agent", "mini-swe-agent", "Gemini 3.8 Flash", "Google", "OpenAI SDK", "Playwright", "GitHub", "PyPI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/mini-web-agent-web-agents-from-scratch-in-under-400-lines", "markdown": "https://wpnews.pro/news/mini-web-agent-web-agents-from-scratch-in-under-400-lines.md", "text": "https://wpnews.pro/news/mini-web-agent-web-agents-from-scratch-in-under-400-lines.txt", "jsonld": "https://wpnews.pro/news/mini-web-agent-web-agents-from-scratch-in-under-400-lines.jsonld"}}