Stop Healing Your Tests: Why Throwaway Automation Fits the AI Era A developer argues that self-healing test automation solves the wrong problem by silently swapping broken locators for plausible ones, trading loud failures for quiet false positives. The proposed alternative is to skip maintained scripts entirely for long-tail checks: an AI agent reads a plain-language test case and executes it in a headless browser via a browser-control tool such as the Playwright MCP server, reporting PASS/FAIL with screenshots for human visual verification. Scenarios that prove valuable can then be promoted into small, reviewed, deterministic Playwright scripts, keeping the maintained suite deliberately small. Self-healing tests are solving the wrong problem. When a "healer" swaps a broken locator for a "close enough" element, the run stays green, but the test might now be asserting something nobody intended. You haven't fixed the test. You've traded a loud failure for a quiet false positive. Here's the alternative I keep coming back to: don't maintain a script at all. If a model can read a test case and execute it in a browser by itself, the script stops being the thing you need. The test case is. You can't break a script that doesn't exist. A self-healing locator answers "the button moved, how do I still click it?" It never asks "should this script still exist?" Worse, a healer can swap a broken selector for a nearby element that looks plausible. The test passes, but it now checks something the author never meant. You've traded a loud failure for a quiet false positive, and those are the most expensive kind of test debt. A locator that keeps breaking is also telling you something unstable hooks, a churning screen , and auto-repair mutes that message. Here's the alternative I keep coming back to. Your source of truth is the test case in plain language , the same one you'd hand to a manual tester. Instead of turning it into a Playwright script that someone has to own, you let an AI agent read it and run it. With a browser-control tool such as the Playwright MCP server, the agent can open a headless browser, follow the steps, look at the page, and report what happened. Because it works from the page as it is right now, there is nothing to heal. If the button moved, it finds the button, the way a human would. A test case in this world is just this: TC-CART-014: Promo banner shows for a new promo SKU Precondition: logged in as a standard test user, cart is empty 1. Add the item with SKU PROMO-001 to the cart 2. Open the cart page Expected: a promo banner is visible above the item list and mentions the discount Also check: no error toast, page has no layout overlap on mobile width And the instruction to the agent is equally short: Run TC-CART-014 against the staging URL in a headless browser. For each step, note what you did and what you saw. Take a screenshot at the end. Report PASS/FAIL per expected result, with evidence. The output is a short report with screenshots. You verify visually. A human skims the evidence in seconds, which is often faster than debugging a red CI job. No script exists, so no script can rot. Most suites have a long tail of checks that matter for a moment: a release, a migration, a risky refactor, a single bug report you want to reproduce. Writing and maintaining scripts for those is a bad trade. Here's the part I like most: this isn't a one-way door. Those background runs double as a discovery pipeline for your maintained suite . When a scenario keeps showing up, catches real bugs, or guards something critical, that's a signal it deserves a proper, deterministic script. At that point you promote it: take what the agent did the steps, the selectors it found, the state it needed and turn it into a reviewed test with real assertions. js import { test, expect } from '@playwright/test'; test 'promo banner shows for a new promo SKU', async { page, request } = { const res = await request.post '/api/cart', { data: { sku: 'PROMO-001', qty: 1 } } ; expect res.ok .toBeTruthy ; const { cartId } = await res.json ; await page.goto /cart/${cartId} ; await expect page.getByTestId 'promo-banner' .toBeVisible ; } ; My rule: if you wouldn't urgently fix a script when it breaks, don't promote it. Everything else stays as a test case the agent runs on demand. Your maintained suite stays small and trusted, because every script in it was chosen, not accumulated. I don't want to oversell it. Some honest limits: Self-healing optimizes for keeping every script alive. The throwaway approach optimizes for keeping the right scripts alive: let AI read your test cases and run them headless, verify the evidence visually, and only turn a scenario into maintained code when it has proven its worth. Your value as an SDET isn't how many scripts you maintain. It's the judgment about which checks deserve to become code at all. So here's my question: which part of your current suite could be replaced by a well-written test case and an agent, and what would you need to see before you trusted it? Originally published at luthfiferdian.com https://luthfiferdian.com/blog/throwaway-automation-instead-of-self-healing .