cd /news/ai-agents/how-to-write-playwright-tests-in-min… · home › topics › ai-agents › article
[ARTICLE · art-145305] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

How To Write Playwright tests in minutes with Playwright MCP and Claude Code

A developer walkthrough shows how to connect Anthropic's Playwright MCP server to Claude Code so the agent reads locators from a live browser's accessibility tree instead of guessing selectors from source code. The setup requires Node.js 20+, Claude Code, and a Playwright project, and is wired in with a single `claude mcp add playwright -- npx -y @playwright/mcp@latest` command, optionally scoped to a project via `.mcp.json`. The author argues grounded locators avoid the common failure where tests pass locally but break in CI because selectors were hallucinated from component code.

by read17 min views1 publishedOct 5, 2026

This article was originally published on the Endform blog*.*

By the end of this walkthrough you'll have a single Playwright test that Claude Code wrote against your live app, that you've reviewed in two passes, run repeatedly to check for flakiness, and committed next to the plain-English scenario it came from. The trick that makes it trustworthy: the agent reads locators off the running page through the Playwright MCP server instead of guessing them from your source code.

You need four things installed before step 1, so start there.

Confirm all four of these first. Skipping one is the most common reason step 1 fails.

Node.js 20 or newer. Check with node --version. The MCP server launches through npx, so Node is required no matter how you installed Claude Code.

Claude Code, installed and signed in. Check with claude --version.

A Playwright project. You need playwright.config.ts at the repo root. No project yet? Run npm init playwright@latest to scaffold one.

An app to test. Local dev server or staging URL, either works. Playwright can start it for you when the config says so.

Got all four? Connect Claude Code to a browser.

Here's a failure pattern you've probably lived through: Claude Code writes an end-to-end test, it passes locally, you merge, and CI goes red a day later on what looks like flakiness. The test was wrong from the start, it just had no way to show it.

The reason is where the locators came from. Reading your components without ever opening the app, an agent spots a button with a class like .btn-primary and writes a selector against it. Then a component library or some runtime logic rewrites that class on its way to the DOM, and by the time the browser renders, the selector points at a hashed string, a restructured node, or nothing. Nobody caught it because nobody ran the test against a real page. CI is the first thing that does.

Compare the two ways the same button gets targeted:

// Hallucinated: guessed from training data, does not exist on this page
await page.locator("#submit-btn").click();

// Grounded: read from the accessibility tree Claude Code can see
await page.getByRole("button", { name: "Place order" }).click();

Anthropic's Model Context Protocol is what closes the gap. The Playwright MCP server hands Claude Code a live browser: it can open the app, walk the accessibility tree, and build locators out of the roles and names the page actually exposes. It's working from a structured snapshot rather than a screenshot, which keeps the locators stable and the token cost low, because the model reasons over text it can already read. Same model, better inputs. Nothing here is a capability upgrade, only an access one.

A single command wires the server into Claude Code:

claude mcp add playwright -- npx -y @playwright/mcp@latest

That entry gets saved to your local config, and the server process spins up the next time you open a session. -y skips the install confirmation npx would otherwise wait on, and @latest grabs the newest published build so you're not stuck on a stale cache.

By default this is a personal, single-project entry. Working on a team? Add --scope project, which writes the same config to a .mcp.json at the repo root so everyone shares one server without redoing setup:

{
  "mcpServers": {
    "playwright": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@playwright/mcp@latest"]
    }
  }
}

Don't trust it until you've checked it. Run claude mcp list and look for ✔ Connected. Immediately after adding, you may catch a ✘ Failed to connect while npx is still pulling the package in the background; run the command again a few seconds later and it usually clears.

If retrying doesn't clear it, the problem is elsewhere. Claude Code runs the server as a subprocess in its own environment, and that subprocess can't always resolve npx the way your interactive shell does. Homebrew-installed Node on macOS is the usual culprit. Give it the absolute binary path instead of trusting PATH:

claude mcp remove playwright
claude mcp add playwright -- $(which npx) -y @playwright/mcp@latest

Run claude mcp list once more and wait for ✔ Connected before continuing.

Connected only means the process is alive. It doesn't prove Claude Code can actually drive a browser yet. Open a session and call the tool by name, otherwise Claude Code might reach for a Bash command instead of the MCP server:

Using Playwright MCP, open [your app's local URL] and tell me the page title and the first heading you see.

If it answers with something concrete off the page, not a hedge and not a generic description, the tool works. That answer is coming from a live snapshot rather than memory, which is the entire reason you connected it.

These commands are current as of August 2026. MCP tooling changes quickly, so if a command here stops matching what you see, check the official Playwright MCP repo. Server verified and driving a browser, the next hurdle is login.

Nearly every test worth writing sits behind authentication. A checkout, a settings screen, an admin panel, all of them dead ends if the agent can't clear the sign-in page.

The MCP server has no idea about your suite's existing auth. It either carries its own browser profile across sessions or, in isolated mode, boots logged out every time. Either way it's disconnected from however your tests currently authenticate.

Fix it by launching the server with a session already loaded:

npx @playwright/mcp@latest --isolated --storage-state .auth/user.json

--isolated holds the profile in memory rather than writing it to disk, so every session starts fresh. --storage-state reads a saved authenticated session from a file, and anything that changes mid-session gets thrown away at the end. But that file has to exist first.

Playwright's auth docs suggest a dedicated setup test that signs in and saves the session:

import { test as setup } from "@playwright/test";

const authFile = ".auth/user.json";

const password = process.env.TEST_USER_PASSWORD;
if (!password) {
  throw new Error("TEST_USER_PASSWORD is not set");
}

setup("authenticate", async ({ page }) => {
  await page.goto("https://your-app.example.com/login");
  await page.getByLabel("Email").fill("test-user@example.com");
  await page.getByLabel("Password").fill(password);
  await page.getByRole("button", { name: "Log in" }).click();
  await page.waitForURL("**/dashboard");
  await page.context().storageState({ path: authFile });
});

Wire it to run automatically by adding a setup project in playwright.config.ts that the browser projects depend on:

projects: [
  { name: 'setup', testMatch: /auth\.setup\.ts/ },
  {
    name: 'chromium',
    use: { storageState: '.auth/user.json' },
    dependencies: ['setup'],
  },
],

Run it once:

npx playwright test --project=setup

When it finishes, .auth/user.json holds a live authenticated session, and auth.setup.ts is now part of the suite, so login logic lives in exactly one place.

Worth doing if your app allows it: create a throwaway test user through an API or seed script before login and tear it down after. Claude Code pokes at the app while it explores, and a disposable account keeps it from mutating a shared login someone else is on.

With the file in place, re-register the server so it loads that session:

claude mcp remove playwright
claude mcp add playwright -- npx -y @playwright/mcp@latest --isolated --storage-state .auth/user.json

For teams, mirror it in .mcp.json:

{
  "mcpServers": {
    "playwright": {
      "type": "stdio",
      "command": "npx",
      "args": [
        "-y",
        "@playwright/mcp@latest",
        "--isolated",
        "--storage-state",
        ".auth/user.json"
      ]
    }
  }
}

For messier auth flows, Endform's Playwright MCP guide goes deeper on this.

Now that Claude Code browses as a logged-in user, it's time to write the prompt that turns a described flow into a committable test.

This step is where a shippable test and a throwaway one diverge. Rather than one rambling paragraph that asks for a test and hopes, split the prompt into three parts that each do one job.

Part one is environment setup: the app-specific facts the agent can't deduce on its own.

The website under test is at https://staging.yourapp.com.
You're already logged in as a test user through the storage state configured earlier.
The test user's cart currently holds one item, a placeholder t-shirt priced at $24.99.

Part two is the scenario, phrased the way you'd brief a teammate, not as pseudocode:

1. Open the cart page.
2. Proceed to checkout.
3. Confirm the shipping address shown is the default one.
4. Select the saved Visa card ending in 4242.
5. Place the order.
6. Confirm the order confirmation page shows an order number and the correct total.

Keep this as a markdown file in the repo next to the tests it drives, not buried in a chat log. Playwright Test Agents work the same way: a planner writes the scenario as markdown, and a later step compiles it into code. Splitting the two into separate, reviewable artifacts is worth carrying over here.

This part is also the one nobody on the team can write better than you. The environment facts are just facts, and the system prompt below is boilerplate you reuse everywhere. The scenario is the only part carrying judgment: does the address get confirmed before payment, does the saved card matter more than a fresh one, is the confirmation total worth asserting or is a loaded page enough? An agent can't rank those. Someone who knows the product has to.

Part three is the system prompt, the standing rules that ride along on every test and describe what "good" looks like, not just what to do:

You are a Playwright test generator.
Explore the app using the Playwright MCP tools before writing any code, don't generate steps from assumption alone.
Prefer getByRole, getByLabel, and getByTestId locators over CSS selectors or XPath.
Don't add manual waitForTimeout calls, rely on Playwright's built-in auto-waiting and retrying assertions instead.
Group related steps with test.step for readability in the trace viewer.
Assert on outcomes a user would actually see, not on incidental implementation details.
Save the finished test to the tests directory, run it, and keep iterating until it passes.

You don't have to write this cold. Debbie O'Brien, a longtime Playwright advocate and ex-member of Microsoft's Playwright team, keeps a public set of prompt files for exactly this, a better base than reinventing it.

Stacked together, this is the full message to Claude Code:

You are a Playwright test generator.
Explore the app using the Playwright MCP tools before writing any code, don't generate steps from assumption alone.
Prefer getByRole, getByLabel, and getByTestId locators over CSS selectors or XPath.
Don't add manual waitForTimeout calls, rely on Playwright's built-in auto-waiting and retrying assertions instead.
Group related steps with test.step for readability in the trace viewer.
Assert on outcomes a user would actually see, not on incidental implementation details.
Save the finished test to the tests directory, run it, and keep iterating until it passes.

The website under test is https://staging.yourapp.com.
You're already logged in as a test user through the storage state configured earlier.
The test user's cart currently holds one item, a placeholder t-shirt priced at $24.99.

1. Open the cart page.
2. Proceed to checkout.
3. Confirm the shipping address shown is the default one.
4. Select the saved Visa card ending in 4242.
5. Place the order.
6. Confirm the order confirmation page shows an order number and the correct total.

The proof this prompt works is watching the agent call the MCP tools to explore before it writes any test code. A "explore first" instruction only counts if the agent obeys it. Three parts, three jobs, one message. Next, the generation itself.

One look at the page isn't where Claude Code stops. It keeps navigating and reading through the MCP server until it has actually seen every element the scenario names. That's what separates this from feeding a model a plain-English description and hoping.

Watch a run and the sequence is plain: open the page, click through the scenario's steps, read back what each returned, and only then start writing assertions. It writes the spec last, runs it, and keeps tweaking until it passes instead of handing you untested code.

A real example: on a logout scenario, two headings both matched "Secure Area," so Playwright threw a strict mode violation, because a locator meant to act on one element matched several. Claude Code read the error, diagnosed it, and appended exact: true so the locator demanded the full heading text rather than a substring. That correction landed in the same pass, no nudge from the developer.

A checkout with a cart, saved card, and confirmation page runs the same loop, step by step, until green. Here's a second flow to show it's not a one-off.

Endform's guide to shipping quality end-to-end tests with Playwright MCP has a scenario worth reusing: check that a newly created team shows up in an activity log. Sign in, confirm a signup event is already logged, create a team, confirm the new event lands. Endform walks through the scenario and the review, not the code, so here's a plausible test around that flow, in the shape Claude Code would generate it against a live app:

import { test, expect } from "@playwright/test";

test("new team activity appears in the activity log", async ({ page }) => {
  await test.step("open the activity log", async () => {
    await page.goto("https://staging.yourapp.com/dashboard");
    await page.getByRole("link", { name: "Activity" }).click();
  });

  await test.step("confirm the signup event is already logged", async () => {
    await expect(page.getByText("you signed up")).toBeVisible();
  });

  await test.step("create a new team", async () => {
    await page.getByRole("button", { name: "Create a new team" }).click();
    await page.getByLabel("Team name").fill("QA Playground");
    await page.getByRole("button", { name: "Create team" }).click();
  });

  await test.step("confirm the new team event appears in the log", async () => {
    await expect(page.getByText("you created a new team")).toBeVisible();
  });
});

Notice how deliberate the file is. Every link and button name is text Claude Code read off the page. The test.step blocks track the scenario one to one, so a failure drops you straight on the broken step. No manual waits anywhere, because the retrying assertions cover timing.

You end up with a runnable file that clears its first run more often than not, precisely because each locator was verified against the page before a line got written. That still isn't the same as a good test, which is what step 5 is for.

Passing once doesn't earn a test a place in the suite. Run two review passes, because they catch different things: one on the code, one on the meaning.

Pass one is the code. Read it and ask if you'd have written it this way. Generated tests tend toward the baroque, redundant checks, extra steps, logic that folds down to less, and trimming that is usually the first win. Hold the locators to the priority you set (getByRole and getByTestId over CSS or XPath), swap out anything brittle, and make sure each assertion proves something a user would notice rather than just confirming an action didn't throw. The test.step groupings should read the way a person would narrate the flow.

Pass two is the one no linter will ever do for you, and it's the one that matters more. Put the scenario next to the finished test and ask whether the thing being verified still matches what the scenario meant, not just whether it's green. Tests drift here quietly. On that logout test, Claude Code noticed the login banner showed once and got eaten by the stored session, so it retargeted the assertion to the permanent heading before wrapping up. Useful that it caught it, but it happened during generation, not review, and that's the point: next time it might not, and this pass is your backstop.

The first draft gets much faster; the judgment moves into refinement. And that judgment, what the product does, what risk you're covering, what the scenario is really asserting, is yours to bring. Both passes done, one last check remains.

The commands below use this tutorial's logout test as the example. A test that passed while being generated still hasn't earned your trust. Run it again, clean, outside the generation loop:

npx playwright test tests/logout.spec.ts

A single green run tells you almost nothing about reliability. Fire it several times in a row:

npx playwright test tests/logout.spec.ts --repeat-each 5

--repeat-each reruns the same test N times in one invocation. It's among the fastest ways to smoke out a test that's green most runs and red occasionally, the exact flakiness that tends to surface only once it hits CI.

Repeats clean? Commit the test together with the markdown scenario behind it:

git add tests/logout.spec.ts tests/scenarios/logout.md
git commit -m "Add logout test with scenario spec"

Committing them together keeps intent attached to implementation. Whoever reads the diff later sees what the test was meant to prove, not just the assertions.

One thing to check before that first commit: your .gitignore. A freshly scaffolded Playwright project ignores playwright/.auth/, its default spot for storage state. But this setup writes to .auth/ at the repo root, a different path the default rule doesn't cover, which means a live authenticated session can slip into history unnoticed.

Add the right path first:

echo ".auth/" >> .gitignore

That one line keeps the session out of your repo history by intent rather than luck.

Verified, stable over repeats, and committed alongside its scenario, the test is a dependable addition. What gets harder from here: longer flows, data that shifts between runs, and auth beyond a plain login form.

Trust comes from being straight about where Claude Code struggles, not just where it shines. Four things break it fairly predictably.

Long multi-step flows. The longer the flow, the more the agent loses the thread, since each step starts from a fresh snapshot rather than the whole sequence before it, and long MCP sessions dropping browser context is a known issue. Splitting the flow into smaller scenarios and stitching them later holds up better.

Data that changes between runs. The agent tends to assert against whatever it saw at generation time, a timestamp, an order ID, a count, and those move on the next run. Assertions last longer when they target something stable: a confirmation state, a completed action, the presence of a result, not the exact value from one run.

Deep conditionals. While generating, the agent only travels one branch, one role, one account state, one flag. Branches it never walks stay invisible to it. Prompt each branch on its own, or write the conditional logic yourself, rather than expecting one pass to find every path.

OAuth and third-party auth. Usually the first wall you hit. Redirects out to Google, Microsoft, or any external identity provider leave your app, and the agent's control loop isn't built to follow through the handoff. One developer automating GitHub's OAuth earned a temporary IP ban for it, and other providers can react to automated logins the same way, which makes this genuinely fragile. The move is to authenticate in a separate setup step, save the session with storageState, and generate against an app that's already logged in.

None of these shrink the workflow's value. They just mark where a scenario needs prep before you hand it over.

We opened with a test that looked done and broke in CI regardless. Claude Code explored the real app through Playwright MCP, checked what was on the page, and produced a test you reviewed, verified, and committed.

That changes the economics of test creation. Once a trustworthy test costs minutes instead of hours, teams write more of them, and a bigger suite brings its own problem. Those tests still run in CI, and as the count climbs, the bottleneck moves from writing tests to running them fast without flakiness.

Endform runs every Playwright test on its own isolated machine in parallel, which keeps suite duration predictable no matter how many you add.

── more in #ai-agents 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-write-playwri…] indexed:0 read:17min 2026-10-05 · —