Playwright is very good at following instructions.
Tell it to click a button, it clicks the button. Tell it to fill an input, it fills the input. Give it a test with 40 steps and, assuming the page behaves, it'll run all 40.
The problem is that someone has to write those steps.
What if you gave Playwright a goal instead?
Add two items to a shopping cart, apply the discount code, and check that the total is correct.
The browser would need to inspect the page, work out what to click, notice when something changes, and decide what to do next. That's a different kind of automation. And it's a fun engineering problem, provided you don't confuse the agent completed its plan with the application passed a test.
Let's build a small version. No giant agent framework. Just Playwright, an LLM API, and a loop we can inspect.
A normal Playwright script is a list of instructions:
await page.getByRole('button', { name: 'Add notebook' }).click();
await page.getByRole('button', { name: 'Checkout' }).click();
Our agent will do something closer to this:
The important word is separately. An agent is useful for finding a path through the interface. It shouldn't get the final vote on whether the system works.
We'll use a local checkout page so this tutorial doesn't depend on somebody else's website or an account with real payment details.
Create a directory and install the dependencies:
mkdir autonomous-playwright
cd autonomous-playwright
npm init -y
npm install playwright dotenv
npx playwright install chromium
You'll also need an API key for an OpenAI-compatible chat completions endpoint. The example below uses OpenAI's endpoint and defaults to gpt-4.1-mini. You can select a different compatible model with an environment variable.
Create index.html:
<!doctype html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>Tiny Shop</title>
</head>
<body>
<main>
<h1>Tiny Shop</h1>
<p>Notebook: $12</p>
<p>Pen: $3</p>
<button id="notebook">Add notebook</button>
<button id="pen">Add pen</button>
<p id="cart" aria-live="polite">Cart: 0 items</p>
<label for="code">Discount code</label>
<input id="code" />
<button id="apply">Apply discount</button>
<p id="discount">Discount: $0</p>
<button id="checkout">Checkout</button>
<p id="receipt" role="status"></p>
</main>
<script>
let notebookCount = 0;
let penCount = 0;
let discount = 0;
const $ = id => document.getElementById(id);
function renderCart() {
$('cart').textContent = `Cart: ${notebookCount + penCount} items`;
}
$('notebook').onclick = () => { notebookCount++; renderCart(); };
$('pen').onclick = () => { penCount++; renderCart(); };
$('apply').onclick = () => {
discount = $('code').value.trim() === 'SAVE5' ? 5 : 0;
$('discount').textContent = `Discount: $${discount}`;
};
$('checkout').onclick = () => {
const total = notebookCount * 12 + penCount * 3 - discount;
$('receipt').textContent = `Order total: $${total}`;
};
</script>
</body>
</html>
This is deliberately boring. Boring fixtures are good. If the experiment goes sideways, we want to know whether the problem is in the agent, not the store.
Run the page in one terminal:
python3 -m http.server 4173
You should now have a checkout page at http://127.0.0.1:4173.
Put your API key in .env:
OPENAI_API_KEY=your_api_key_here
MODEL=gpt-4.1-mini
Don't commit this file to Git.
Now create agent.mjs:
import 'dotenv/config';
import { chromium } from 'playwright';
import fs from 'node:fs/promises';
const BASE_URL = 'http://127.0.0.1:4173';
const MAX_STEPS = 12;
const MODEL = process.env.MODEL || 'gpt-4.1-mini';
const GOAL =
'Add exactly one notebook and one pen, apply code SAVE5, ' +
'then checkout. The correct final total is $10.';
if (!process.env.OPENAI_API_KEY) {
throw new Error('Set OPENAI_API_KEY in .env');
}
// The model can choose only one of these operations.
// It cannot submit raw JS, shell commands, or arbitrary URLs.
const SYSTEM_PROMPT = `
You operate a local demo shopping page through a restricted browser API.
Return ONLY a JSON object with one of these shapes:
{"action":"click","role":"button","name":"visible accessible name"}
{"action":"fill","role":"textbox","name":"visible accessible name","value":"text"}
{"action":"finish","reason":"brief explanation"}
Choose ONE action per response. Base it on the current snapshot and history.
Do not invent elements. Do not repeat an action that already succeeded.
Only use the role and accessible name shown on the page.
When the checkout receipt shows the requested total, choose finish.
`;
async function chooseAction(snapshot, history) {
const response = await fetch('https://api.openai.com/v1/chat/completions', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({
model: MODEL,
temperature: 0,
response_format: { type: 'json_object' },
messages: [
{ role: 'system', content: SYSTEM_PROMPT },
{
role: 'user',
content: JSON.stringify({ goal: GOAL, history, snapshot })
}
]
})
});
if (!response.ok) {
throw new Error(`Model API returned ${response.status}: ${await response.text()}`);
}
const data = await response.json();
const content = data.choices?.[0]?.message?.content;
if (!content) throw new Error('Model returned no content');
return JSON.parse(content);
}
function validateAction(a) {
if (!a || typeof a !== 'object' || Array.isArray(a)) {
throw new Error('Invalid action');
}
if (a.action === 'finish') {
if (typeof a.reason !== 'string') throw new Error('Missing finish reason');
return;
}
if (a.action !== 'click' && a.action !== 'fill') {
throw new Error(`Unsupported action: ${a.action}`);
}
if (a.role !== (a.action === 'click' ? 'button' : 'textbox')) {
throw new Error('Role not permitted for this action');
}
if (typeof a.name !== 'string' || !a.name || a.name.length > 100) {
throw new Error('Invalid accessible name');
}
if (a.action === 'fill') {
if (typeof a.value !== 'string' || a.value.length > 100) {
throw new Error('Invalid input value');
}
}
}
async function applyAction(page, action) {
validateAction(action);
if (action.action === 'finish') return;
const locator = page.getByRole(action.role, {
name: action.name,
exact: true
});
// Do not guess when a locator is ambiguous.
const count = await locator.count();
if (count !== 1) throw new Error(`Expected one match, found ${count}`);
if (action.action === 'click') await locator.click({ timeout: 3000 });
if (action.action === 'fill') await locator.fill(action.value, { timeout: 3000 });
}
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();
const page = await context.newPage();
const history = [];
let agentFinished = false;
let failure;
try {
await context.tracing.start({ screenshots: true, snapshots: true });
await page.goto(BASE_URL, { waitUntil: 'domcontentloaded' });
for (let step = 1; step <= MAX_STEPS; step++) {
// The snapshot is a compact description of the accessible UI.
const snapshot = (await page.locator('body').ariaSnapshot()).slice(0, 12000);
const action = await chooseAction(snapshot, history);
validateAction(action);
console.log(`Step ${step}:`, action);
history.push(action);
if (action.action === 'finish') {
agentFinished = true;
break;
}
// Prevent the agent from navigating out of our local test target.
if (new URL(page.url()).origin !== new URL(BASE_URL).origin) {
throw new Error('Agent left the approved origin');
}
await applyAction(page, action);
}
if (!agentFinished) throw new Error(`Agent exceeded ${MAX_STEPS} steps`);
// Here is the actual test oracle, independent of the model's opinion.
const cart = await page.locator('#cart').textContent();
const discount = await page.locator('#discount').textContent();
const receipt = await page.locator('#receipt').textContent();
if (cart?.trim() !== 'Cart: 2 items') throw new Error(`Wrong cart: ${cart}`);
if (discount?.trim() !== 'Discount: $5') {
throw new Error(`Wrong discount: ${discount}`);
}
if (receipt?.trim() !== 'Order total: $10') {
throw new Error(`Wrong receipt: ${receipt}`);
}
console.log('PASS: correct cart, discount, and total');
} catch (error) {
failure = error;
console.error('FAIL:', error.message);
await page.screenshot({ path: 'failure.png', fullPage: true }).catch(() => {});
} finally {
await fs.writeFile('agent-history.json', JSON.stringify(history, null, 2));
await context.tracing.stop({ path: 'trace.zip' }).catch(() => {});
await browser.close();
}
if (failure) process.exitCode = 1;
Run it in a second terminal:
node agent.mjs
On a successful run, the model will choose actions corresponding to adding a notebook, adding a pen, entering SAVE5, applying it, checking out, and finishing. The exact order and number of model calls can vary. At the end, the script checks three concrete facts: two cart items, a $5 discount, and a $10 receipt.
If it fails, you'll still get agent-history.json and trace.zip; failures also attempt a screenshot. Open the trace with:
npx playwright show-trace trace.zip
A note about the API: response_format: { type: 'json_object' } asks the model for JSON, not for a guaranteed valid browser command. That's why validateAction() still exists. The code uses the raw Playwright library rather than Playwright Test, so the trace contains browser activity but not Playwright Test's assertion metadata.
Notice how little authority we gave the LLM.
It can click a named button or fill a named textbox. That's it. The model can't execute JavaScript, visit another website, delete files, call internal APIs, or decide that a failing assertion should be ignored.
It also doesn't receive the entire DOM. We use Playwright's locator.ariaSnapshot() to send a readable view of the accessible interface. This API is available in Playwright 1.49 and later. It's often enough for buttons, links, forms, and page text. It isn't magic: poorly labeled controls and visual-only interfaces remain difficult.
Three design choices are doing most of the work:
One action at a time. If the model proposes a 12-step plan in advance, step four may be wrong because step two opened a dialog. Observe again after every action.
A fixed budget. Autonomous doesn't mean infinite. Twelve actions is plenty for our checkout. For a production app, give each scenario a time limit, a step limit, and a clear failure state.
An independent oracle. We don't ask, "Did you test checkout successfully?" and accept "Yes". We inspect application state and compare it with expected values. In a real system you might check an order API, a database record, or a known fixture instead of trusting the page alone.
That last point is where a lot of impressive autonomous testing demos quietly stop being convincing.
Change this line in index.html:
const total = notebookCount * 12 + penCount * 3 - discount;
to:
const total = notebookCount * 12 + penCount * 3;
Now the checkout ignores the discount. The agent may complete all the clicks and even report that it's finished. The script should still fail because the receipt says $15, not $10.
That's what you want from a test: a failure that doesn't depend on how confident the agent sounds.
Put the original line back before continuing.
For our small page, a model call after every action is manageable. For a typical enterprise flow, the numbers look different.
Say you have a signup journey with 25 browser actions. At one model call per action, that's 25 model requests for one attempt. Run it across four browsers and you've got up to 100 requests, before retries. Add 200 scenarios and you start caring about latency, context size, model availability, and cost.
And the more interesting the app, the worse the edge cases become:
These aren't criticisms of Playwright. Playwright gives you APIs for popups, frames, uploads, and waiting. But now your agent needs policies for all of them, plus recovery logic, reporting, retries, and a way for a human to correct its decisions.
You can build it. The question is whether maintaining that platform is the best use of your team's time.
I wouldn't point the script above at a production account. It's a learning example with intentionally narrow permissions, not a hardened security boundary. Its origin check happens between steps; real deployments should also restrict network access and use isolated test accounts.
Before a serious rollout, I'd make a few changes:
page is always the right page or that accessibility snapshots expose every useful target.
The last one is especially useful. Autonomous discovery and deterministic regression testing solve slightly different problems. You don't have to choose only one.
If you're learning how browser agents work, this project is worth doing. You'll understand the observation/action loop, the failure modes, and why good test oracles matter.
If you're responsible for a team's regression suite, the calculation changes.
There are platforms that already package autonomous browser exploration with test creation and managed execution. Endtest is one example. Its Endtest Bot is designed to explore a web application and generate test scenarios, while the broader platform handles running and managing web tests. That's much closer to the outcome most QA teams want than maintaining their own model prompts and browser orchestration code.
I'd evaluate a product like that on the difficult flows, not on a login demo: iframes, multiple tabs, uploads, conditional paths, and what happens when an element changes. Also ask to inspect and edit the generated steps. If you can't understand what an autonomous test did, you have a maintenance problem disguised as an AI feature.
A custom Playwright agent can still be the better choice if you need full control over models, infrastructure, or unusual browser behavior. You're trading platform fees for engineering time. Neither option is free.
You can make Playwright autonomous with surprisingly little code. The browser library was never the hard part.
The hard part is deciding what the agent is allowed to do, how to tell whether it did the right thing, and what evidence you have when it didn't.
Start with one business-critical flow. Make the failure deterministic. Then add autonomy where it saves you work, rather than where it makes for the flashiest demo.
That's a much better test automation strategy than building a robot that can click anything and trusting it when it says everything is fine.
References: Playwright accessibility snapshots, Playwright locator API, Playwright tracing, Playwright browser contexts.