{"slug": "how-to-make-playwright-autonomous-without-letting-an-ai-agent-run-wild", "title": "How to Make Playwright Autonomous (Without Letting an AI Agent Run Wild)", "summary": "A developer published a tutorial showing how to make Playwright autonomous by pairing it with an OpenAI-compatible LLM in a simple inspectable loop, using a local checkout page as a fixture. The agent is given a goal — add one notebook and one pen, apply code SAVE5, and check out for a $10 total — rather than a hardcoded list of steps, with the author stressing that the agent should find a path through the UI but not get the final vote on whether the application passes.", "body_md": "Playwright is very good at following instructions.\n\nTell it to click a button, it clicks the button. Tell it to fill an input, it fills the input. Give it a test with 40 steps and, assuming the page behaves, it'll run all 40.\n\nThe problem is that someone has to write those steps.\n\nWhat if you gave Playwright a goal instead?\n\nAdd two items to a shopping cart, apply the discount code, and check that the total is correct.\n\nThe browser would need to inspect the page, work out what to click, notice when something changes, and decide what to do next. That's a different kind of automation. And it's a fun engineering problem, provided you don't confuse *the agent completed its plan* with *the application passed a test*.\n\nLet's build a small version. No giant agent framework. Just Playwright, an LLM API, and a loop we can inspect.\n\nA normal Playwright script is a list of instructions:\n\n```\nawait page.getByRole('button', { name: 'Add notebook' }).click();\nawait page.getByRole('button', { name: 'Checkout' }).click();\n```\n\nOur agent will do something closer to this:\n\nThe important word is **separately**. An agent is useful for finding a path through the interface. It shouldn't get the final vote on whether the system works.\n\nWe'll use a local checkout page so this tutorial doesn't depend on somebody else's website or an account with real payment details.\n\nCreate a directory and install the dependencies:\n\n```\nmkdir autonomous-playwright\ncd autonomous-playwright\nnpm init -y\nnpm install playwright dotenv\nnpx playwright install chromium\n```\n\nYou'll also need an API key for an OpenAI-compatible chat completions endpoint. The example below uses OpenAI's endpoint and defaults to `gpt-4.1-mini`. You can select a different compatible model with an environment variable.\n\nCreate `index.html`:\n\n```\n<!doctype html>\n<html lang=\"en\">\n<head>\n  <meta charset=\"UTF-8\">\n  <title>Tiny Shop</title>\n</head>\n<body>\n  <main>\n    <h1>Tiny Shop</h1>\n    <p>Notebook: $12</p>\n    <p>Pen: $3</p>\n    <button id=\"notebook\">Add notebook</button>\n    <button id=\"pen\">Add pen</button>\n    <p id=\"cart\" aria-live=\"polite\">Cart: 0 items</p>\n    <label for=\"code\">Discount code</label>\n    <input id=\"code\" />\n    <button id=\"apply\">Apply discount</button>\n    <p id=\"discount\">Discount: $0</p>\n    <button id=\"checkout\">Checkout</button>\n    <p id=\"receipt\" role=\"status\"></p>\n  </main>\n  <script>\n    let notebookCount = 0;\n    let penCount = 0;\n    let discount = 0;\n    const $ = id => document.getElementById(id);\n    function renderCart() {\n      $('cart').textContent = `Cart: ${notebookCount + penCount} items`;\n    }\n    $('notebook').onclick = () => { notebookCount++; renderCart(); };\n    $('pen').onclick = () => { penCount++; renderCart(); };\n    $('apply').onclick = () => {\n      discount = $('code').value.trim() === 'SAVE5' ? 5 : 0;\n      $('discount').textContent = `Discount: $${discount}`;\n    };\n    $('checkout').onclick = () => {\n      const total = notebookCount * 12 + penCount * 3 - discount;\n      $('receipt').textContent = `Order total: $${total}`;\n    };\n  </script>\n</body>\n</html>\n```\n\nThis is deliberately boring. Boring fixtures are good. If the experiment goes sideways, we want to know whether the problem is in the agent, not the store.\n\nRun the page in one terminal:\n\n```\npython3 -m http.server 4173\n```\n\nYou should now have a checkout page at `http://127.0.0.1:4173`.\n\nPut your API key in `.env`:\n\n```\nOPENAI_API_KEY=your_api_key_here\nMODEL=gpt-4.1-mini\n```\n\nDon't commit this file to Git.\n\nNow create `agent.mjs`:\n\n``` python\nimport 'dotenv/config';\nimport { chromium } from 'playwright';\nimport fs from 'node:fs/promises';\n\nconst BASE_URL = 'http://127.0.0.1:4173';\nconst MAX_STEPS = 12;\nconst MODEL = process.env.MODEL || 'gpt-4.1-mini';\nconst GOAL =\n  'Add exactly one notebook and one pen, apply code SAVE5, ' +\n  'then checkout. The correct final total is $10.';\n\nif (!process.env.OPENAI_API_KEY) {\n  throw new Error('Set OPENAI_API_KEY in .env');\n}\n\n// The model can choose only one of these operations.\n// It cannot submit raw JS, shell commands, or arbitrary URLs.\nconst SYSTEM_PROMPT = `\nYou operate a local demo shopping page through a restricted browser API.\nReturn ONLY a JSON object with one of these shapes:\n{\"action\":\"click\",\"role\":\"button\",\"name\":\"visible accessible name\"}\n{\"action\":\"fill\",\"role\":\"textbox\",\"name\":\"visible accessible name\",\"value\":\"text\"}\n{\"action\":\"finish\",\"reason\":\"brief explanation\"}\n\nChoose ONE action per response. Base it on the current snapshot and history.\nDo not invent elements. Do not repeat an action that already succeeded.\nOnly use the role and accessible name shown on the page.\nWhen the checkout receipt shows the requested total, choose finish.\n`;\n\nasync function chooseAction(snapshot, history) {\n  const response = await fetch('https://api.openai.com/v1/chat/completions', {\n    method: 'POST',\n    headers: {\n      Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,\n      'Content-Type': 'application/json'\n    },\n    body: JSON.stringify({\n      model: MODEL,\n      temperature: 0,\n      response_format: { type: 'json_object' },\n      messages: [\n        { role: 'system', content: SYSTEM_PROMPT },\n        {\n          role: 'user',\n          content: JSON.stringify({ goal: GOAL, history, snapshot })\n        }\n      ]\n    })\n  });\n\n  if (!response.ok) {\n    throw new Error(`Model API returned ${response.status}: ${await response.text()}`);\n  }\n\n  const data = await response.json();\n  const content = data.choices?.[0]?.message?.content;\n  if (!content) throw new Error('Model returned no content');\n  return JSON.parse(content);\n}\n\nfunction validateAction(a) {\n  if (!a || typeof a !== 'object' || Array.isArray(a)) {\n    throw new Error('Invalid action');\n  }\n  if (a.action === 'finish') {\n    if (typeof a.reason !== 'string') throw new Error('Missing finish reason');\n    return;\n  }\n  if (a.action !== 'click' && a.action !== 'fill') {\n    throw new Error(`Unsupported action: ${a.action}`);\n  }\n  if (a.role !== (a.action === 'click' ? 'button' : 'textbox')) {\n    throw new Error('Role not permitted for this action');\n  }\n  if (typeof a.name !== 'string' || !a.name || a.name.length > 100) {\n    throw new Error('Invalid accessible name');\n  }\n  if (a.action === 'fill') {\n    if (typeof a.value !== 'string' || a.value.length > 100) {\n      throw new Error('Invalid input value');\n    }\n  }\n}\n\nasync function applyAction(page, action) {\n  validateAction(action);\n  if (action.action === 'finish') return;\n\n  const locator = page.getByRole(action.role, {\n    name: action.name,\n    exact: true\n  });\n\n  // Do not guess when a locator is ambiguous.\n  const count = await locator.count();\n  if (count !== 1) throw new Error(`Expected one match, found ${count}`);\n\n  if (action.action === 'click') await locator.click({ timeout: 3000 });\n  if (action.action === 'fill') await locator.fill(action.value, { timeout: 3000 });\n}\n\nconst browser = await chromium.launch({ headless: true });\nconst context = await browser.newContext();\nconst page = await context.newPage();\nconst history = [];\nlet agentFinished = false;\nlet failure;\n\ntry {\n  await context.tracing.start({ screenshots: true, snapshots: true });\n  await page.goto(BASE_URL, { waitUntil: 'domcontentloaded' });\n\n  for (let step = 1; step <= MAX_STEPS; step++) {\n    // The snapshot is a compact description of the accessible UI.\n    const snapshot = (await page.locator('body').ariaSnapshot()).slice(0, 12000);\n    const action = await chooseAction(snapshot, history);\n    validateAction(action);\n\n    console.log(`Step ${step}:`, action);\n    history.push(action);\n\n    if (action.action === 'finish') {\n      agentFinished = true;\n      break;\n    }\n\n    // Prevent the agent from navigating out of our local test target.\n    if (new URL(page.url()).origin !== new URL(BASE_URL).origin) {\n      throw new Error('Agent left the approved origin');\n    }\n    await applyAction(page, action);\n  }\n\n  if (!agentFinished) throw new Error(`Agent exceeded ${MAX_STEPS} steps`);\n\n  // Here is the actual test oracle, independent of the model's opinion.\n  const cart = await page.locator('#cart').textContent();\n  const discount = await page.locator('#discount').textContent();\n  const receipt = await page.locator('#receipt').textContent();\n\n  if (cart?.trim() !== 'Cart: 2 items') throw new Error(`Wrong cart: ${cart}`);\n  if (discount?.trim() !== 'Discount: $5') {\n    throw new Error(`Wrong discount: ${discount}`);\n  }\n  if (receipt?.trim() !== 'Order total: $10') {\n    throw new Error(`Wrong receipt: ${receipt}`);\n  }\n\n  console.log('PASS: correct cart, discount, and total');\n} catch (error) {\n  failure = error;\n  console.error('FAIL:', error.message);\n  await page.screenshot({ path: 'failure.png', fullPage: true }).catch(() => {});\n} finally {\n  await fs.writeFile('agent-history.json', JSON.stringify(history, null, 2));\n  await context.tracing.stop({ path: 'trace.zip' }).catch(() => {});\n  await browser.close();\n}\n\nif (failure) process.exitCode = 1;\n```\n\nRun it in a second terminal:\n\n```\nnode agent.mjs\n```\n\nOn a successful run, the model will choose actions corresponding to adding a notebook, adding a pen, entering `SAVE5`, applying it, checking out, and finishing. The exact order and number of model calls can vary. At the end, the script checks three concrete facts: two cart items, a $5 discount, and a $10 receipt.\n\nIf it fails, you'll still get `agent-history.json` and `trace.zip`; failures also attempt a screenshot. Open the trace with:\n\n```\nnpx playwright show-trace trace.zip\n```\n\nA note about the API: `response_format: { type: 'json_object' }` asks the model for JSON, not for a guaranteed valid browser command. That's why `validateAction()` still exists. The code uses the raw Playwright library rather than Playwright Test, so the trace contains browser activity but not Playwright Test's assertion metadata.\n\nNotice how little authority we gave the LLM.\n\nIt can click a **named button** or fill a **named textbox**. That's it. The model can't execute JavaScript, visit another website, delete files, call internal APIs, or decide that a failing assertion should be ignored.\n\nIt also doesn't receive the entire DOM. We use Playwright's `locator.ariaSnapshot()` to send a readable view of the accessible interface. This API is available in Playwright 1.49 and later. It's often enough for buttons, links, forms, and page text. It isn't magic: poorly labeled controls and visual-only interfaces remain difficult.\n\nThree design choices are doing most of the work:\n\n**One action at a time.** If the model proposes a 12-step plan in advance, step four may be wrong because step two opened a dialog. Observe again after every action.\n\n**A fixed budget.** Autonomous doesn't mean infinite. Twelve actions is plenty for our checkout. For a production app, give each scenario a time limit, a step limit, and a clear failure state.\n\n**An independent oracle.** We don't ask, \"Did you test checkout successfully?\" and accept \"Yes\". We inspect application state and compare it with expected values. In a real system you might check an order API, a database record, or a known fixture instead of trusting the page alone.\n\nThat last point is where a lot of impressive autonomous testing demos quietly stop being convincing.\n\nChange this line in `index.html`:\n\n``` js\nconst total = notebookCount * 12 + penCount * 3 - discount;\n```\n\nto:\n\n``` js\nconst total = notebookCount * 12 + penCount * 3;\n```\n\nNow the checkout ignores the discount. The agent may complete all the clicks and even report that it's finished. The script should still fail because the receipt says `$15`, not `$10`.\n\nThat's what you want from a test: a failure that doesn't depend on how confident the agent sounds.\n\nPut the original line back before continuing.\n\nFor our small page, a model call after every action is manageable. For a typical enterprise flow, the numbers look different.\n\nSay you have a signup journey with 25 browser actions. At one model call per action, that's 25 model requests for one attempt. Run it across four browsers and you've got up to 100 requests, before retries. Add 200 scenarios and you start caring about latency, context size, model availability, and cost.\n\nAnd the more interesting the app, the worse the edge cases become:\n\nThese aren't criticisms of Playwright. Playwright gives you APIs for popups, frames, uploads, and waiting. But now *your agent* needs policies for all of them, plus recovery logic, reporting, retries, and a way for a human to correct its decisions.\n\nYou can build it. The question is whether maintaining that platform is the best use of your team's time.\n\nI wouldn't point the script above at a production account. It's a learning example with intentionally narrow permissions, not a hardened security boundary. Its origin check happens between steps; real deployments should also restrict network access and use isolated test accounts.\n\nBefore a serious rollout, I'd make a few changes:\n\n`page` is always the right page or that accessibility snapshots expose every useful target.\nThe last one is especially useful. Autonomous discovery and deterministic regression testing solve slightly different problems. You don't have to choose only one.\n\nIf you're learning how browser agents work, this project is worth doing. You'll understand the observation/action loop, the failure modes, and why good test oracles matter.\n\nIf you're responsible for a team's regression suite, the calculation changes.\n\nThere are platforms that already package autonomous browser exploration with test creation and managed execution. [Endtest](https://endtest.io/) is one example. Its [Endtest Bot](https://endtest.io/docs/advanced/endtest-bot) is designed to explore a web application and generate test scenarios, while the broader platform handles running and managing web tests. That's much closer to the outcome most QA teams want than maintaining their own model prompts and browser orchestration code.\n\nI'd evaluate a product like that on the difficult flows, not on a login demo: iframes, multiple tabs, uploads, conditional paths, and what happens when an element changes. Also ask to inspect and edit the generated steps. If you can't understand what an autonomous test did, you have a maintenance problem disguised as an AI feature.\n\nA custom Playwright agent can still be the better choice if you need full control over models, infrastructure, or unusual browser behavior. You're trading platform fees for engineering time. Neither option is free.\n\nYou can make Playwright autonomous with surprisingly little code. The browser library was never the hard part.\n\nThe hard part is deciding what the agent is allowed to do, how to tell whether it did the right thing, and what evidence you have when it didn't.\n\nStart with one business-critical flow. Make the failure deterministic. Then add autonomy where it saves you work, rather than where it makes for the flashiest demo.\n\nThat's a much better test automation strategy than building a robot that can click anything and trusting it when it says everything is fine.\n\n**References:** [Playwright accessibility snapshots](https://playwright.dev/docs/aria-snapshots), [Playwright locator API](https://playwright.dev/docs/api/class-locator), [Playwright tracing](https://playwright.dev/docs/api/class-tracing), [Playwright browser contexts](https://playwright.dev/docs/api/class-browsercontext).", "url": "https://wpnews.pro/news/how-to-make-playwright-autonomous-without-letting-an-ai-agent-run-wild", "canonical_source": "https://dev.to/orbitpickle307/how-to-make-playwright-autonomous-without-letting-an-ai-agent-run-wild-3ck4", "published_at": "2026-10-10 21:14:46+00:00", "updated_at": "2026-10-10 21:16:02.701801+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models"], "entities": ["Playwright", "OpenAI", "gpt-4.1-mini", "Chromium", "Node.js"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-to-make-playwright-autonomous-without-letting-an-ai-agent-run-wild", "markdown": "https://wpnews.pro/news/how-to-make-playwright-autonomous-without-letting-an-ai-agent-run-wild.md", "text": "https://wpnews.pro/news/how-to-make-playwright-autonomous-without-letting-an-ai-agent-run-wild.txt", "jsonld": "https://wpnews.pro/news/how-to-make-playwright-autonomous-without-letting-an-ai-agent-run-wild.jsonld"}}