Auditable AI in QA: what Playwright MCP taught me about scaffolding tests that don't drift A developer built @vijaypjavvadi/bdd2pw v4.3.2, an MIT-licensed Gherkin .feature-to-Playwright TypeScript scaffolder built on Microsoft's Playwright MCP server, with a paper published in Elsevier SoftwareX in October 2026. The tool hashes the browser accessibility tree rather than the DOM to detect meaningful page drift, and mechanically rewrites common LLM emission errors (such as inverted fill() arguments or dropped awaits) before rejecting a generated binding. The developer reports that naive model-generated Playwright bindings run against the actual page only about 60% of the time. A Playwright test fails in staging. The button locator misses. You start the bisect, half an hour in you realise you've never seen the file before — an LLM generated it, someone shipped it, and nobody can tell you which model, which prompt, or which run produced any of the four broken lines. A test suite is a contract with the future. When you can't say where a line came from, you can't say what it means. I spent most of this year building @vijaypjavvadi/bdd2pw https://www.npmjs.com/package/@vijaypjavvadi/bdd2pw currently v4.3.2 as a way to think through that failure mode. It's an MIT-licensed Gherkin .feature → Playwright TypeScript scaffolder built on the Microsoft Playwright MCP server https://github.com/microsoft/playwright-mcp . A paper describing it was published in Elsevier SoftwareX https://www.sciencedirect.com/science/article/pii/S2352711026004243 in October 2026 DOI 10.1016/j.softx.2026.102933 https://doi.org/10.1016/j.softx.2026.102933 . This post isn't the paper. It's the three design intuitions behind it that I had to learn the hard way, and would have paid real money to be told at the start. The Playwright MCP server exposes a browser's accessibility tree to a caller — role, name, id, label, landmark structure. Ordinary Playwright codegen reads the DOM once, picks a locator heuristically, and hands you a spec. MCP lets you keep the tree around and interrogate it. I used to think the interesting part of MCP was locator picking. It isn't. The interesting part is that the accessibility tree gives you a stable, canonical representation of the page — an artifact you can hash. Two trees that describe the same UI hash to the same sha256, regardless of what CSS or whitespace churned underneath. That property is what I would build change-detection on top of in any future UI-testing tool. The DOM churns for reasons unrelated to what the page actually is: class names get regenerated, whitespace shifts, hidden comment nodes appear and disappear. Hash the DOM and every drift alarm has a 90% chance of being noise. Hash the accessibility tree and every drift alarm is a real drift alarm, because the accessibility tree is what your tests actually use. If you take one thing from this post: when you want to answer "has this page changed in a way that matters," ignore the markup and hash the accessibility tree. If you point a model at "generate a Playwright binding for this Gherkin step against this Page Object" , it will produce something. That something will look right. It will typecheck about half the time. On the apps I've tested informally, it runs against the actual page roughly 60% of the time. The failures aren't random. There are repeatable patterns — the model invents a fill locator, value helper on the Page Object that doesn't exist, drops a required await , imports from the wrong path, confuses expect page with expect page.locator ... . Each of these looks like code and compiles against ambient types; at runtime it crashes. Early on I treated these as flat failures and dropped them on the floor. The step landed as a // TODO and a human had to come back. That was correct — but it was too eager. Most of the common emission mistakes are mechanically recoverable before you ever ask a human. The model wants loginPage.usernameInput.fill "standard user" ; it emits loginPage.fill "usernameInput", "standard user" . The intent is unambiguous; the surface form is wrong. You can lift the intent out and rewrite the AST to the correct form, then run the rewritten version through the same gate that would have rejected it originally. If the rewrite survives, you ship it. If it doesn't, you drop it. Rewrite before you reject. Fail-closed is a floor, not a ceiling. If you're going to spend a model call producing a binding, spend a millisecond trying to repair the output before you throw it out. The hardest bug I had wasn't a wrong locator. It was a locator that was right six months ago, generated by a model I'd since swapped out, and I had no record of which version of the pipeline had produced it. Every "why" question — why did this fail now, which pack owns this rule, does the LLM actually earn its keep on this suite — was unanswerable because the emitted code carried no provenance. The minimum attribution surface I'd now insist on in any AI-assisted codegen tool, from day one: rule-14 , pack banking or a specific LLM call provider anthropic , model claude-sonnet-4-6 , cached or not, governance-sanitised or not, UTC timestamp . @llm-generated next to the existing tags is enough — then your trusted-only lane is --grep-invert @llm-generated with no new framework.