Every time our Frontend (FE) changed a UI component, a little nightmare started: tests broke, the Test Automation (TA) team was called and what seemed to be endless back-and-forth happened.
I saw an opportunity to reduce this friction by introducing a headless agent into our workflow - one that picks up the broken tests and tries to fix them.
Developing automations like this is a hard process though - designing a meaningful flow, dealing with cross-dependencies and accesses that might turn into dead-ends, and other unknown issues along the way. In this article, I'll walk you through the problem, the design decisions and trade-offs, and how it was implemented end-to-end.
When working in frontend, we have some CI steps set up. One of them is to run test automation. More often than not, these tests fail because frontend changed the HTML structure or user flow altogether. When this happens, we have to go through these steps shown in the diagram.
Summarizing:
This workflow triggers for every feature branch opened by FE. The steps have to be done in sequence, and the flow involves at least four engineers (the ones writing the code and the ones reviewing it), context switching, and cross-team dependency. Also, there is a period in between that we lose test coverage: the tests pass, but only because they're skipped.
For our experiment, we decided to narrow our focus to something specific and easy to fix: when FE changes a UI component, needing a selector change; and only smoke tests would be used as the gate. I’ll give you more context below.
In frontend applications, we implement components that users interact with. These components are accessible - either by their text, input placeholder, role or by non-visible screen-reader aria-* attributes, falling back to data-testid. These identifiers, selectors (or locators, as some libraries call them), are used by testing libraries, such as Testing Library, Playwright or Cypress to conduct automated tests, which simulate user flows.
Whenever a text, label, or whatever is used to identify the testing component is changed, the tests break. You may think this is a rare problem, but it can happen in some situations:
Smoke tests are a small set of end-to-end tests that cover only the most critical flows of an application. The idea is not to test everything, but to quickly answer one question: is the app still working? Because they are fewer and faster than the full suite, they are cheaper to run on every feature branch - and for this experiment, that also meant fewer (and less flaky) failures for the agent to deal with. If you want to go deeper, this article is a good starting point.
This use case is not a one-size-fits-all solution, but it can give you some ideas. We had this specific setup:
The overall idea is that whenever there is a change in FE, the CI pipeline runs smoke tests, and if they fail, it would trigger a workflow agent in TA by sending a repository_dispatch event to the TA repo (through the GitHub REST API), which starts a GitHub Actions workflow there. This agent would then try to fix the non-passing tests. If it is successful, it would create a PR for a human to check. If this PR was merged, then the FE CI pipeline would pass.
Some notes to consider here:
After coming to this high-level design (after a lot of back and forth), I started to think about the details.
To implement this, some questions had to be answered:
Considering that every feature branch in FE triggers a CI pipeline and because of cost-efficiency and having more control, especially during the testing of this automation, I chose to trigger the agent under certain conditions. The feature branch changes would need to be deployed into a specific environment (we have more than 4 environments set up, so we can test multiple feature branches at the same time). Let’s call this environment the QA Env.
Having this environment is also needed because of the constraint that TA can only run smoke tests against deployed environments.
There are multiple steps of the automation design, and one of them is to decide where to trigger the agentic flow. As mentioned in the Constraints, we were using both Drone and GA.
Our Drone was becoming crowded already - it was running processes on all branches, including main. And I realised that with GA we would have more control on when and how to trigger - their interface was more friendly for testing and manual triggering. Also, I could create a separated flow only for the agent, not mixing with other processes.
For those reasons, I decided to go with GA for the agentic workflow, therefore the point of connection should be GA for triggering the agent.
The initiator of this automation is FE when two conditions are met: smoke tests fail AND it is deployed to QA Env. Since FE CI runs in Drone and FE CD runs in GA, these two conditions are checked in different places. I had two options:
The first option would be more costly - we would need to bring the smoke test step into the GA, and it would duplicate the step in both CI and CD pipelines.
The second option, on the other hand, would be a simple check. That’s because we use GitHub labels to track which environment is being deployed. We already had a flow set up where adding a label to a PR triggers a deployment to that environment.
Given that, we could check the GitHub label right after the Smoke tests step. If both smoke tests failed and the deployment label was set as QA Env, then Drone would send the dispatch event to TA.
Given that CI and CD happen in different environments, one trade-off here is that it is possible that Drone will receive and identify the label, but the GA deployment process can fail, or Drone check can happen before the branch is actually deployed. I accepted these trade-offs, because the worst that could happen is that the tests run against an environment without FE change, triggering a false negative.
For the race condition, as soon as the label is added, usually the CD happens in 5 minutes. The entire CI pipeline usually takes 15 minutes, making it a comfortable pass, since the check would be as the last step of the CI pipeline.
The FE created a new feature that broke smoke tests, it was deployed to QA Env, and the dispatch event was sent. Now what? What should the agent do?
The entire goal is for the broken smoke tests to be fixed, by updating the selectors. For that, I took some decisions:
So, with three main outcomes, plus two edge cases we would have:
Although it is possible to see all the logs in GA workflows page, I wanted to make provide as much up front transparency as possible, by ending the flow in FE PR, with a message on what was the outcome of the entire process.
Then, with most of the definitions set up, we can finally look at how the agent does this, and the harnesses around it, in the next sections.
As you will see, there are multiple steps to implement the flow end-to-end.
A harness is everything around the model that shapes how the agent works - instruction files, skills, commands, scripts and guardrails - so it behaves predictably without someone guiding it step by step.
The initial state of the codebase was not ready for AI at all. To avoid increasing the blast radius of the project, I decided to just create enough harnesses to implement this POC. Two assets were created:
CLAUDE.md pointing to it, which enables uses for both Claude and OpenAI models to read. Normal instructions file that is injected to every agent session (not going too deep into this as there's plenty of documentation on the web).getByRole, then by getByLabel and so on.
With those, I tested how it performed using my own agent (local Claude Code CLI).
Other commands were created too, but they belong to the automation per se, so they will be specified in the following sections.
Another problem was that the tests were not able to run against QA Env. Data was coupled into a specific environment, as well as some selectors. This isn’t related to AI nor to automation, but it would impact the ability to develop it. Therefore, I had an additional prep work: creating functions that would setup/create and teardown/delete data for every test run and making selectors environment agnostic.
A visual diagram of the entire flow is shown below:
As the flow starts in FE, naturally I had to start there, but the work was slim: I’ve added a new step in drone.yml, after the Smoke test step:
- name: Smoke test
commands:
- name: Dispatch selector repair
environment:
GITHUB_APP_TOKEN:
from_secret: GITHUB_APP_TOKEN
REPAIR_ENVIRONMENTS: QA_Env
commands:
- sh ci/drone/dispatch-selector-repair.sh
depends_on:
- Smoke test
when:
status:
- failure
(Note: this is an example piece of code and should not be used as it will not work).
As variables, a GitHub App token was needed to enable cross-repository communication and the REPAIR_ENVIRONMENTS is a mirror list of which environments are enabled. For now, only QA Env is enabled to trigger this automation.
I could add all the instructions in the YAML file, but I had extracted them into ci/drone/dispatch-selector-repair.sh to separate concerns and improve readability. In this script, a few things were done:
frontend-smoke-failed event
Then, TA would receive this event and payload n its end.
To receive the event, a GA workflow was created. selector-repair.yml is triggered when the event is received:
repository_dispatch:
types: [frontend-smoke-failed]
It can also be trigger manually, although this was mostly enabled for testing purposes, when evaluating the performance of the agent.
This workflow:
codex exec command, which runs the repair-selectors-ci)
As you can see, the LLM is only involved in in one specific step. Before and after, deterministic scripts were written, both to reduce AI costs and improve quality. The agent is only triggered if it can actually run the tests in an enabled environment; and after the agent finishes, a script puts together the message body for the PR comment back to FE.
On top of the shared harness from the setup section, the CI run adds two commands. Namely:
The Repair selector command is where the core agent logic lives. In summary, it does:
1. Reads inputs from environment variables: which environment to test, which FE PR triggered the run, and a list of failed specs.
2. Manages a persistent branch: one branch per frontend PR, reused and appended to across multiple runs rather than recreated each time, as each commit in the same branch triggers a new run. It never rebases or force-pushes, since the commit history works as an audit log. Also, if there is a new commit from FE side, GA cancels the current run via concurrency: cancel-in-progress, so it doesn’t keep working on outdated code. If the commit was already pushed, it stays there.
3. Determines what to test: either the specific failing specs or runs the whole smoke suite itself, in case of a manual run. If a previous repair run already committed a fix for it on the persistent branch, it's categorized as “Already repaired”.
4. Triages each failure into one of three buckets:
5. Repairs rot-classified selectors by probing the live page for accessible signals (role, label, text, placeholder, etc.), picking the most precise selector that resolves to exactly one element, and verifying the fix actually turns the test green before keeping it.
6. Runs quality gates — lint/prettier checks, cleanup of any scratch/probe files — before allowing anything to be committed.
7. Opens or updates a draft PR (never merges) - this is the Open repair PR command.
8. Writes a structured JSON summary (repair-summary.json) as its only communication channel back to the calling workflow — listing what was fixed, what was escalated as a regression, what couldn't be reproduced, what a prior run already fixed, and any environment error — with strict rules about which fields must stay empty in which scenarios, so the human-facing report is never misleading.
After that, the payload is sent back to selector-repair.yml to be picked up by the last step of "Comment back on FE PR". This is a deterministic code that executes the feedback to the FE repository as we already discussed in "Outcomes: what should the agent report?" section.
Before this experiment, a single breaking UI change meant skipping tests, merging with gaps in coverage, and coordinating at least four engineers across two teams. With this flow, FE gets feedback directly in its PR, TA reviews a draft fix instead of writing it from scratch, and, when the agent is successful, tests no longer need to sit skipped on main.
The biggest lesson for me was that the agent was the smallest part of the work. Most of the effort went into the pieces around it: deciding when it should run, connecting two repositories and two CI tools, making the tests run against a shared environment, and writing deterministic checks so the LLM is only called when it can actually help.
There's still room to improve and I believe this can serve as a case study for autonomous systems that go beyond using AI in local development. The future is agents working in the cloud without a human driving each step - and for that, we need solid harnesses and guardrails, and to keep observing their output to improve quality, reduce costs and speed things up.