Reliable Visual Regression Testing for Humans and Coding Agents JupyterLab cut its visual regression test suite runtime from an average of 55 minutes to 14-16 minutes and reduced flaky tests per run from 17 in January 2026 to 2.4 in August 2026, according to a post by a senior software engineer at OpenTeams and JupyterLab maintainer. The suite compares about 350 Playwright reference screenshots per pull request via Galata, and regenerating screenshots dropped from about 45 minutes started by a maintainer to about 1 minute requested by any contributor; runs with a hard failure fell from 54% to 14%. In the three months to 25 September 2026, 28 of the 30 merged pull requests that added or changed a UI test in JupyterLab were AI-assisted, and fixing the flaky tests uncovered 19 real bugs, all now fixed. - Senior Software Engineer at OpenTeams. JupyterLab maintainer and Jupyter Distinguished Contributor. JupyterLab uses visual regression testing to catch unintended changes to the interface before a release: it compares about 350 reference screenshots snapshots, in Playwright’s terms on every pull request. The tests use Galata https://github.com/jupyterlab/jupyterlab/tree/main/galata , JupyterLab’s test framework built on Playwright. Until February 2026 the suite took 55 minutes, and now it takes 14 to 16. Flaky tests, which fail once and pass on retry, went from 17 per run in January to 2.4 in August. Any contributor can now request new reference images with a comment. A local run on Linux with the CI fonts produces the same pixels as CI, and on Fedora 44 with its default fonts, 85% of the screenshots match. In the three months to 25 September 2026, 28 of the 30 merged pull requests that added or changed a UI test in JupyterLab were AI-assisted.