{"slug": "the-feature-shipped-the-ui-automation-didnt", "title": "The Feature Shipped. The UI Automation Didn’t.", "summary": "A quality engineer describes using four Model Context Protocol integrations — Jira, Figma, GitLab, and Playwright — to let an AI assistant assemble UI regression tests from manual test cases, design specs, and live application inspection within the same delivery workstream. The engineer emphasizes that the assistant gathers context and drafts tests while QE review and execution determine whether the results can be trusted, and warns that conflicting expectations between Jira, Figma, and the implementation must be clarified rather than resolved in favor of current application behavior.", "body_md": "*How I’m using Jira, Figma, GitLab, and Playwright MCPs to make automation part of delivery—not the task that comes after it.*\n\n“Can we finish testing this before the release?”\n\n“Yes.”\n\n“And automate it?”\n\n“We’ll pick that up next sprint.”\n\nIf you work in quality engineering, that conversation probably sounds familiar.\n\nIt’s not that nobody cares about automation. It’s that **delivery has a deadline, while automation often has a backlog.**\n\nThere are test cases to execute, defects to discuss, fixes to verify, and last-minute changes to validate. Once the feature ships, the next one is already waiting.\n\nSome teams have dedicated automation engineers while other QEs focus on delivery. In other teams, the same engineer does both.\n\nEither way, there’s a gap to manage: a handoff between people, or a competition for one person’s time.\n\nI’m using AI assistants connected to tools through MCP to make that gap smaller. The biggest benefit isn’t simply generating scripts faster.\n\nIt’s making it more practical to **test a feature and build its regression coverage in the same workstream.**\n\nConsider an invitation form.\n\nYou enter an invalid email address, check that a validation message appears, and confirm that the invitation cannot be submitted.\n\nA straightforward manual check.\n\nNow automate it.\n\nOpen developer tools. Inspect the input. Find a stable locator. Inspect the button. Locate the error message. Wire everything into the existing page object or helper. Add assertions. Run the script.\n\nThen investigate why it passes locally but fails in the pipeline.\n\nThen check the other browsers the team supports.\n\nGood frameworks make this easier, and not every locator requires a complicated XPath. But there is still plenty of mechanical work between:\n\n“I know what this feature should do.”\n\nAnd:\n\n“We have a reliable regression test for it.”\n\nThat’s the part I want help with.\n\nI’m using four MCP integrations:\n\n| Integration | What I use it for | \n|---|---|\n| **Jira MCP** | Load manual test cases and understand the scenario. | \n| **Figma MCP** | Review the designs and intended UI states. | \n| **GitLab MCP** | Reference the UI implementation. | \n| **Playwright MCP** | Inspect the running application, identify locators, and help verify automation behavior. | \n\nMCP—the Model Context Protocol—gives the AI assistant a way to access capabilities exposed by these tools.\n\nIt doesn’t make testing decisions by itself. The assistant brings the context together, and I review how that context is used.\n\nPreviously, I would move between these tools, collect the relevant details, and translate them into automation.\n\nNow I can ask the assistant to do more of that preparation and implementation, without explaining everything from scratch.\n\n*Four sources of context, one connected workflow. The assistant helps assemble the test; QE review and execution establish whether it can be trusted.*\n\nThere’s a big difference between:\n\nWrite a UI test for invalid email validation.\n\nLoad the Jira case for invalid email validation. Review the linked Figma frame and relevant UI component in GitLab. Inspect the form in the test environment using Playwright.\n\nAdd the test using our existing framework conventions. Verify the locators against the live page and run the generated test through our test runner.\n\nFlag conflicting expectations or missing information. Don’t change assertions just to make the test pass.\n\nThe first prompt leaves the assistant plenty of room to guess.\n\nThe second asks it to gather evidence before writing the test.\n\nFor an illustrative example, suppose Jira says the error should appear after submission, Figma shows validation when the user leaves the field, and the implementation validates while typing.\n\nWhich one should become the expected result?\n\nNot automatically the implementation.\n\nThat disagreement needs clarification. Otherwise, we risk generating a test that confirms what the application currently does rather than what it should do.\n\n**Sometimes the most useful output is a question—not a script.**\n\nLocators are a good example of where this workflow helps.\n\nPlaywright MCP can inspect the running page and expose information about elements, their accessible names, and their states. GitLab context can help explain how those elements are implemented.\n\nThe assistant can use that information to propose locators and check them in the browser.\n\nI still need to review the choices:\n\nAlso, a temporary element reference from an MCP browser snapshot is not a permanent locator for the test suite. The committed test still needs maintainable selectors.\n\nBut reviewing a locator that has already been investigated is different from starting every lookup myself.\n\nThe same applies to wiring methods, reusing fixtures, and following existing framework patterns.\n\n**I spend less time assembling the test and more time checking what it proves.**\n\nWhen I’m testing a feature, I know its details.\n\nI know which requirement needed clarification. I know which edge case exposed a defect. I know what changed after the latest fix.\n\nThat’s a useful time to create regression coverage.\n\nIf automation moves to a later sprint, someone needs to reconstruct that understanding. Sometimes that someone is me, trying to remember why a particular scenario mattered.\n\nWith this setup, the workflow becomes closer to:\n\n**Review the scenario → validate the feature → prompt → inspect the generated test → run → refine.**\n\nIt isn’t “one prompt and done.”\n\nBut it makes automation more manageable alongside my delivery responsibilities. Instead of postponing all the implementation work, I can get assistance with it while staying focused on the feature.\n\nDedicated automation engineers still have an important role. Framework architecture, test data, pipeline reliability, and suite health don’t take care of themselves.\n\nThe opportunity is to make contributing useful coverage easier for the QEs who are already testing the changes.\n\nThat helps keep the regression suite closer to the product—and gives the team a stronger basis for frequent, confident releases.\n\n*The opportunity isn’t to remove engineering effort—it’s to bring regression coverage closer to delivery. Token costs, browser execution time, and human review remain.*\n\nThis is one of the practical downsides I notice: the workflow can be token-heavy.\n\nOne scenario might involve reading a Jira case, retrieving design context, inspecting source files, capturing browser snapshots, generating code, and reviewing execution output.\n\nIf the test fails, another inspection-and-correction cycle follows.\n\nConnecting more tools doesn’t mean I should ask the assistant to read everything.\n\nA broad request such as:\n\nReview the application and automate the user-management module.\n\nLeaves a lot of room for expensive exploration.\n\nA narrower task gives it a clearer path:\n\nAutomate invalid-email validation using this Jira case, this Figma frame, and the invitation component. Reuse the existing login fixture.\n\nI find the important discipline is controlling scope: relevant files, specific scenarios, focused browser inspection, and a clear stopping point.\n\n**The goal isn’t maximum context. It’s enough context to make the next decision correctly.**\n\nCompared with API automation, the UI side still takes longer.\n\nAn API test can often send a request and inspect a response directly.\n\nA UI test may need to log in, navigate, open a dialog, wait for rendering, enter data, trigger validation, and observe the result.\n\nBetween those steps are loading states, overlays, animations, asynchronous updates, and browser differences.\n\nMCP doesn’t remove that complexity.\n\nIt can help with the investigation, but **less manual effort doesn’t necessarily mean less elapsed time.**\n\nThere’s another distinction I have to keep in mind: the assistant completing a flow in its browser session is not the same as the generated script passing independently.\n\nThat browser might already be authenticated or contain state from an earlier attempt.\n\nThe actual script still needs to run through the normal test runner, from a controlled starting state, in the required browsers and pipeline environment.\n\nA successful demonstration is useful. It isn’t the finish line.\n\nA generated test can look reasonable and still miss the point.\n\nFor example:\n\n```\nawait email.fill('invalid-email');\nawait expect(email).toHaveValue('invalid-email');\n```\n\nThat checks that the field contains the value the test entered.\n\nIt doesn’t verify that the application rejected the invalid address.\n\nIf the agreed behavior is validation on blur, an inline error, and disabled submission, the meaningful checks might look like this:\n\n``` js\nconst email = page.getByRole('textbox', {\n  name: 'Email address',\n  exact: true,\n});\n\nconst submit = page.getByRole('button', {\n  name: 'Send invitation',\n  exact: true,\n});\n\nawait email.fill('invalid-email');\nawait email.blur();\n\nawait expect(\n  page.getByText('Enter a valid email address', { exact: true })\n).toBeVisible();\n\nawait expect(submit).toBeDisabled();\n```\n\nThese labels and messages are illustrative; the real test should use the application’s verified behavior and existing conventions.\n\nEven this example verifies a specific UI contract, not backend enforcement.\n\nMy review question stays the same:\n\n**If the intended behavior broke, would this test fail for the right reason?**\n\nI don’t want the assistant to make a failure disappear by removing an assertion, selecting a different element, or adding delays without understanding the cause.\n\nA green test is useful only when it means something.\n\nI don’t see this workflow removing engineering effort from UI automation.\n\nIt changes where that effort goes.\n\nLess repetitive locator collection and method wiring. More attention to scenarios, assertions, maintainability, and execution evidence.\n\nThere are trade-offs: token consumption, browser execution time, and generated code that still needs review. I still need to verify the scripts in the environments where the team depends on them.\n\nBut I can stay aligned with my deliverables while maintaining the regression coverage those deliverables need.\n\nThat’s the improvement I care about.\n\n**I’m not trying to stop doing automation work. I’m trying to stop postponing it.**\n\n*Does automation move with feature delivery in your team—or usually follow a sprint behind?*", "url": "https://wpnews.pro/news/the-feature-shipped-the-ui-automation-didnt", "canonical_source": "https://dev.to/balasundar_veluchamy_ccbe/the-feature-shipped-the-ui-automation-didnt-lo1", "published_at": "2026-09-27 23:05:40+00:00", "updated_at": "2026-09-27 23:30:46.936725+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "developer-tools", "ai-tools", "mlops"], "entities": ["Jira", "Figma", "GitLab", "Playwright", "Model Context Protocol"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-feature-shipped-the-ui-automation-didnt", "markdown": "https://wpnews.pro/news/the-feature-shipped-the-ui-automation-didnt.md", "text": "https://wpnews.pro/news/the-feature-shipped-the-ui-automation-didnt.txt", "jsonld": "https://wpnews.pro/news/the-feature-shipped-the-ui-automation-didnt.jsonld"}}