# The Feature Shipped. The UI Automation Didn’t.

> Source: <https://dev.to/balasundar_veluchamy_ccbe/the-feature-shipped-the-ui-automation-didnt-lo1>
> Published: 2026-09-27 23:05:40+00:00

*How I’m using Jira, Figma, GitLab, and Playwright MCPs to make automation part of delivery—not the task that comes after it.*

“Can we finish testing this before the release?”

“Yes.”

“And automate it?”

“We’ll pick that up next sprint.”

If you work in quality engineering, that conversation probably sounds familiar.

It’s not that nobody cares about automation. It’s that **delivery has a deadline, while automation often has a backlog.**

There are test cases to execute, defects to discuss, fixes to verify, and last-minute changes to validate. Once the feature ships, the next one is already waiting.

Some teams have dedicated automation engineers while other QEs focus on delivery. In other teams, the same engineer does both.

Either way, there’s a gap to manage: a handoff between people, or a competition for one person’s time.

I’m using AI assistants connected to tools through MCP to make that gap smaller. The biggest benefit isn’t simply generating scripts faster.

It’s making it more practical to **test a feature and build its regression coverage in the same workstream.**

Consider an invitation form.

You enter an invalid email address, check that a validation message appears, and confirm that the invitation cannot be submitted.

A straightforward manual check.

Now automate it.

Open developer tools. Inspect the input. Find a stable locator. Inspect the button. Locate the error message. Wire everything into the existing page object or helper. Add assertions. Run the script.

Then investigate why it passes locally but fails in the pipeline.

Then check the other browsers the team supports.

Good frameworks make this easier, and not every locator requires a complicated XPath. But there is still plenty of mechanical work between:

“I know what this feature should do.”

And:

“We have a reliable regression test for it.”

That’s the part I want help with.

I’m using four MCP integrations:

| Integration | What I use it for | 
|---|---|
| **Jira MCP** | Load manual test cases and understand the scenario. | 
| **Figma MCP** | Review the designs and intended UI states. | 
| **GitLab MCP** | Reference the UI implementation. | 
| **Playwright MCP** | Inspect the running application, identify locators, and help verify automation behavior. | 

MCP—the Model Context Protocol—gives the AI assistant a way to access capabilities exposed by these tools.

It doesn’t make testing decisions by itself. The assistant brings the context together, and I review how that context is used.

Previously, I would move between these tools, collect the relevant details, and translate them into automation.

Now I can ask the assistant to do more of that preparation and implementation, without explaining everything from scratch.

*Four sources of context, one connected workflow. The assistant helps assemble the test; QE review and execution establish whether it can be trusted.*

There’s a big difference between:

Write a UI test for invalid email validation.

Load the Jira case for invalid email validation. Review the linked Figma frame and relevant UI component in GitLab. Inspect the form in the test environment using Playwright.

Add the test using our existing framework conventions. Verify the locators against the live page and run the generated test through our test runner.

Flag conflicting expectations or missing information. Don’t change assertions just to make the test pass.

The first prompt leaves the assistant plenty of room to guess.

The second asks it to gather evidence before writing the test.

For an illustrative example, suppose Jira says the error should appear after submission, Figma shows validation when the user leaves the field, and the implementation validates while typing.

Which one should become the expected result?

Not automatically the implementation.

That disagreement needs clarification. Otherwise, we risk generating a test that confirms what the application currently does rather than what it should do.

**Sometimes the most useful output is a question—not a script.**

Locators are a good example of where this workflow helps.

Playwright MCP can inspect the running page and expose information about elements, their accessible names, and their states. GitLab context can help explain how those elements are implemented.

The assistant can use that information to propose locators and check them in the browser.

I still need to review the choices:

Also, a temporary element reference from an MCP browser snapshot is not a permanent locator for the test suite. The committed test still needs maintainable selectors.

But reviewing a locator that has already been investigated is different from starting every lookup myself.

The same applies to wiring methods, reusing fixtures, and following existing framework patterns.

**I spend less time assembling the test and more time checking what it proves.**

When I’m testing a feature, I know its details.

I know which requirement needed clarification. I know which edge case exposed a defect. I know what changed after the latest fix.

That’s a useful time to create regression coverage.

If automation moves to a later sprint, someone needs to reconstruct that understanding. Sometimes that someone is me, trying to remember why a particular scenario mattered.

With this setup, the workflow becomes closer to:

**Review the scenario → validate the feature → prompt → inspect the generated test → run → refine.**

It isn’t “one prompt and done.”

But it makes automation more manageable alongside my delivery responsibilities. Instead of postponing all the implementation work, I can get assistance with it while staying focused on the feature.

Dedicated automation engineers still have an important role. Framework architecture, test data, pipeline reliability, and suite health don’t take care of themselves.

The opportunity is to make contributing useful coverage easier for the QEs who are already testing the changes.

That helps keep the regression suite closer to the product—and gives the team a stronger basis for frequent, confident releases.

*The opportunity isn’t to remove engineering effort—it’s to bring regression coverage closer to delivery. Token costs, browser execution time, and human review remain.*

This is one of the practical downsides I notice: the workflow can be token-heavy.

One scenario might involve reading a Jira case, retrieving design context, inspecting source files, capturing browser snapshots, generating code, and reviewing execution output.

If the test fails, another inspection-and-correction cycle follows.

Connecting more tools doesn’t mean I should ask the assistant to read everything.

A broad request such as:

Review the application and automate the user-management module.

Leaves a lot of room for expensive exploration.

A narrower task gives it a clearer path:

Automate invalid-email validation using this Jira case, this Figma frame, and the invitation component. Reuse the existing login fixture.

I find the important discipline is controlling scope: relevant files, specific scenarios, focused browser inspection, and a clear stopping point.

**The goal isn’t maximum context. It’s enough context to make the next decision correctly.**

Compared with API automation, the UI side still takes longer.

An API test can often send a request and inspect a response directly.

A UI test may need to log in, navigate, open a dialog, wait for rendering, enter data, trigger validation, and observe the result.

Between those steps are loading states, overlays, animations, asynchronous updates, and browser differences.

MCP doesn’t remove that complexity.

It can help with the investigation, but **less manual effort doesn’t necessarily mean less elapsed time.**

There’s another distinction I have to keep in mind: the assistant completing a flow in its browser session is not the same as the generated script passing independently.

That browser might already be authenticated or contain state from an earlier attempt.

The actual script still needs to run through the normal test runner, from a controlled starting state, in the required browsers and pipeline environment.

A successful demonstration is useful. It isn’t the finish line.

A generated test can look reasonable and still miss the point.

For example:

```
await email.fill('invalid-email');
await expect(email).toHaveValue('invalid-email');
```

That checks that the field contains the value the test entered.

It doesn’t verify that the application rejected the invalid address.

If the agreed behavior is validation on blur, an inline error, and disabled submission, the meaningful checks might look like this:

``` js
const email = page.getByRole('textbox', {
  name: 'Email address',
  exact: true,
});

const submit = page.getByRole('button', {
  name: 'Send invitation',
  exact: true,
});

await email.fill('invalid-email');
await email.blur();

await expect(
  page.getByText('Enter a valid email address', { exact: true })
).toBeVisible();

await expect(submit).toBeDisabled();
```

These labels and messages are illustrative; the real test should use the application’s verified behavior and existing conventions.

Even this example verifies a specific UI contract, not backend enforcement.

My review question stays the same:

**If the intended behavior broke, would this test fail for the right reason?**

I don’t want the assistant to make a failure disappear by removing an assertion, selecting a different element, or adding delays without understanding the cause.

A green test is useful only when it means something.

I don’t see this workflow removing engineering effort from UI automation.

It changes where that effort goes.

Less repetitive locator collection and method wiring. More attention to scenarios, assertions, maintainability, and execution evidence.

There are trade-offs: token consumption, browser execution time, and generated code that still needs review. I still need to verify the scripts in the environments where the team depends on them.

But I can stay aligned with my deliverables while maintaining the regression coverage those deliverables need.

That’s the improvement I care about.

**I’m not trying to stop doing automation work. I’m trying to stop postponing it.**

*Does automation move with feature delivery in your team—or usually follow a sprint behind?*
