What Makes Recorded UI Tests Survive Real Product Changes? An engineer working on CueCast, a web test recording and replay tool, detailed the technical challenges that cause recorded browser tests to break under real product changes, including ambiguous selectors, virtualized lists, non-standard input controls, and incidental hover steps. The engineer argues that replays must preserve the semantic meaning of targets—such as a node's parent path and role—rather than relying on positional selectors, and must capture run-specific values like generated project IDs for later verification. Recording a browser test can take minutes. Trusting it during a release six weeks later is harder. A menu gains another item. A tree has two nodes named “Settings.” A search option exists only while it is scrolled into view. The project you created on the first run already exists on the second. The recording may have captured every click correctly, yet the next run fails or, worse, clicks the wrong thing and passes. I work on CueCast https://www.icuecast.ai/ , a tool for recording and replaying web tests. These are some of the problems we have had to address. They apply to recorded tests broadly, whether you use a visual recorder, generate a script, or maintain browser tests by hand. Imagine a test that opens a tree and selects a node named “Production.” At recording time, there is one such node. Later, another team adds a second “Production” node under a different parent. A positional selector might still find a node. That is not enough. The test needs to preserve the meaning of the target: for example, Environments → Payments → Production , rather than “the third tree item.” Useful recording context can include the node's label, its parent path, its depth, its role, and nearby stable attributes. At replay time, those signals can help distinguish candidates. If the recorded parent path no longer exists, the test should report that mismatch instead of silently choosing another node with the same label. This is the same principle behind good hand-written browser tests. Prefer stable, user-facing meaning such as accessible roles and names or deliberate test IDs. Use positional selectors only when position is part of the behavior being tested. Even a context-aware locator cannot guarantee that every redesign will work automatically. When the product changes the meaning or structure of a workflow, the test needs review. The win is that a small, harmless layout change does not have to invalidate the whole case. Virtualized lists make a recorded selector especially fragile. A large dropdown may render only the visible options. The option recorded yesterday can be absent from the DOM today simply because the list opened at a different scroll position. A robust replay needs to recognize the scroll container, move through it, and search for the option as rows are rendered. It also needs a stopping rule: if the item never appears, fail with evidence about the container and target rather than waiting indefinitely. There is an important distinction here. Waiting for an element to appear helps when it is loading. A virtualized option may never appear until the list is scrolled. Adding another fixed delay does not solve that problem. Modern web apps contain controls that look like text fields but do not behave like a plain