cd /news/developer-tools/reliable-visual-regression-testing-f… · home › topics › developer-tools › article
[ARTICLE · art-145773] src=openteams.com ↗ pub= topic=developer-tools verified=true sentiment=↑ positive

Reliable Visual Regression Testing for Humans and Coding Agents

JupyterLab cut its visual regression test suite runtime from an average of 55 minutes to 14-16 minutes and reduced flaky tests per run from 17 in January 2026 to 2.4 in August 2026, according to a post by a senior software engineer at OpenTeams and JupyterLab maintainer. The suite compares about 350 Playwright reference screenshots per pull request via Galata, and regenerating screenshots dropped from about 45 minutes started by a maintainer to about 1 minute requested by any contributor; runs with a hard failure fell from 54% to 14%. In the three months to 25 September 2026, 28 of the 30 merged pull requests that added or changed a UI test in JupyterLab were AI-assisted, and fixing the flaky tests uncovered 19 real bugs, all now fixed.

by read17 min views2 publishedOct 6, 2026
Reliable Visual Regression Testing for Humans and Coding Agents
Image: Openteams (auto-discovered)

Senior Software Engineer at OpenTeams. JupyterLab maintainer and Jupyter Distinguished Contributor.

JupyterLab uses visual regression testing to catch unintended changes to the interface before a release: it compares about 350 reference screenshots (snapshots, in Playwright’s terms) on every pull request. The tests use Galata, JupyterLab’s test framework built on Playwright. Until February 2026 the suite took 55 minutes, and now it takes 14 to 16. Flaky tests, which fail once and pass on retry, went from 17 per run in January to 2.4 in August. Any contributor can now request new reference images with a comment. A local run on Linux with the CI fonts produces the same pixels as CI, and on Fedora 44 with its default fonts, 85% of the screenshots match.

In the three months to 25 September 2026, 28 of the 30 merged pull requests that added or changed a UI test in JupyterLab were AI-assisted.<sup>1</sup> A developer who works with an agent on interface code needs a test that runs on their machine, a result they can trust, and an answer in minutes.

This post describes what we changed and which parts you can reuse for Playwright tests on GitHub Actions.

Before and after #

Before Now
Suite run time 42 min to 1 h 25 min, 55 min on average 14 to 16 min
Regenerating screenshots about 45 min, started by a maintainer about 1 min, requested by any contributor
Local run matches CI no yes, with the CI fonts
Looking at a failure download and unzip the report, start a web server click a badge in the pull request
Flaky tests per run 17 (January 2026) 2.4 (August 2026)
Runs with a hard failure 54% (January 2026) 14% (August 2026)

Fixing the flaky tests uncovered real bugs. Many tests were flaky because of a bug in JupyterLab that appeared only under CI timing. Examples are a notebook that took focus while it initialized, a race when several settings change quickly and a debugger that did not show its variables after F9. While pinning the fonts for the tests, we found three parts of the interface that ignored the configured font. All 19 bugs found this way are fixed.2

We also ran the Chromium tests on two setups other than the CI runner: Fedora 44 with its default fonts, and the Ubuntu 26.04 runner that CI will have to move to.3

Before Now
Fedora 44: screenshots that match CI 1% 85%
Fedora 44: tests that pass 56% 93%
Ubuntu 26.04 runner: screenshots that match CI 35% 100%
Ubuntu 26.04 runner: tests that pass 71% 99.8%

If you work with a coding agent #

If you would like to submit a PR to JupyterLab with a coding agent’s assistance:

  • Your agent can now check its work and generate reference images on your Linux machine.
  • A failure usually means a real problem. This is important because, given a false failure, an agent may attempt to edit working code, often making unnecessary edits or masking underlying issues (since debugging the unrelated test is not its prime objective).
  • After opening a PR you get an answer in 15 minutes not 55 minutes. This translates to 30 iteration rounds per working day, rather than 8 as previously.
  • The rules are in files the agent reads. JupyterLab’s AGENTS.md points to the contributing guide, and the guide lists the testing practices described below. Lint rules report fixed waits and hand-written selectors before a reviewer reads the test. Darshan Paudyal explainswhy an agent follows a lint rule more reliably than an instruction in AGENTS.md .
  • You and a reviewer can watch what the agent’s test does. The report has a video of every new or changed test.

Same pixels locally and on CI #

A screenshot of the interface is mostly text, and text rendering depends on the machine. We found these causes:

  • Fonts come from the operating system, so they differ between a laptop and the CI runner, and between two point releases of one Linux distribution.
  • Playwright’s install-deps installed about 80 system packages on CI, among them fonts that developer machines do not have (e.g.xfonts-cyrillic ). The theme asked forsystem-ui first, so the font used depended on the installed packages.
  • The browser defaults for subpixel antialiasing, font smoothing, optical sizing and kerning differ between platforms.
  • The terminal uses a WebGL renderer when WebGL is available and a DOM renderer when it is not, and the two draw text differently.
  • Text drawn on a canvas, as in the terminal and the data grid, does not use CSS.
  • The IPython console banner shows the IPython version, so console screenshots changed with every IPython upgrade.

JupyterLab 4.6 fixes all of them. The Galata helper extension ships its own fonts as npm dependencies, so the lockfile pins their versions:

"@fontsource/dejavu-sans": "^5.2.5",
"@fontsource/dejavu-mono": "^5.2.5",
"@fontsource-variable/noto-sans-sc": "^5.2.10"

It applies them over the system fonts, and sets the properties that the browser would otherwise choose:

:root {
  --jp-code-font-family-default: 'DejaVu Mono' !important;
  font-kerning: normal;
  -webkit-font-smoothing: none;
  -moz-osx-font-smoothing: none;
  font-optical-sizing: none;
}

The tests start Chromium with --disable-lcd-text, which turns off subpixel antialiasing, and --disable-webgl, which makes the terminal use its DOM renderer.

We stopped running install-deps. The only font it installed that the tests needed was one for Chinese characters, and Noto Sans SC in the list above took its place.

The terminal emulator, xterm.js, has no API for kerning or text rendering, so the tests set them on the canvas context:

ctx.fontKerning = 'normal';
ctx.textRendering = 'geometricPrecision';

The data grid did not apply the configured font-family to its canvases. Users saw this too, so we fixed it in JupyterLab.

The console banner is off in the default Galata settings. IPython now prints a static banner when SOURCE_DATE_EPOCH is set.

The new fonts changed every existing screenshot, and one pull request updated 310 reference images. No version of the fonts reproduced the old images, and neither did fonts copied from the CI runner. The update was due anyway: moving CI from Ubuntu 22.04 to 24.04 broke about 150 tests through font versions alone. With the fonts pinned, a runner upgrade no longer changes the fonts in the screenshots.

Fedora and openSUSE

We ran the Chromium tests in Fedora 44 and openSUSE Tumbleweed containers on GitHub Actions, against the reference images from the Ubuntu runner. With the runner’s font packages installed (DejaVu, Liberation, Lato and Noto Color Emoji), both matched all 255 screenshot comparisons. openSUSE also needed the three fontconfig rules that turn off hinting for small DejaVu text, which Ubuntu and Fedora ship with the font.

With the fonts that Fedora 44 installs by default, 39 screenshots still differ, all of Mermaid diagrams. JupyterLab shows a Mermaid diagram as an SVG image, and an image cannot use the fonts of the page, so its text uses a system font. Two Vega charts had the same problem because they asked for sans-serif, and the test chart now sets DejaVu Sans in its Vega config.

macOS and Windows

We tried reusing the Linux reference images on the macOS and Windows runners of GitHub Actions, with a sample of 20% of the screenshots, and at first 4% matched. We overrode navigator.platform, because macOS menus showed shortcuts as symbols and were up to 63 pixels narrower. We also turned off glyph hinting and subpixel positioning on Linux with --font-render-hinting=none and --disable-font-subpixel-positioning. After that no element differed in size, but only 14% of the screenshots matched on macOS and 4% on Windows. FreeType, CoreText and DirectWrite draw the same glyphs differently, and neither the operating systems’ font smoothing settings nor Chromium’s --text-contrast and --text-gamma switches changed that. The option left is a tolerance per platform: with maxDiffPixelRatio: 0.01, 59% of the failing screenshots would pass on macOS and 53% on Windows.

Snapshot updates from the failed run #

The old workflow for new reference images rebuilt JupyterLab, ran the whole suite with --update-snapshots and pushed the result. A maintainer had to start it, and it took about 45 minutes to produce a few images.

The failed run already has those images. Playwright writes the actual screenshot next to the expected one for every failed comparison, and its JSON reporter records which reference file each one belongs to. We enabled the JSON reporter on CI:

reporter: process.env.CI
  ? [['blob'], ['json', { outputFile: 'test-results/report.json' }]]
  : [['list'], ['html', { open: 'on-failure' }]],

A script, unpack_snapshots.py, copies each image to the path of its reference.

Now a contributor writes this comment on their pull request:

please open PR to update snapshots

After a maintainer approves the run, a bot waits for the test run of the head commit to finish, takes the screenshots from its artifacts, and opens a pull request against the contributor’s branch. A new request on the same pull request replaces the one in progress. The contributor accepts the new images by merging it. After the test run, the bot takes about a minute. The time grows with the number of changed screenshots and does not depend on the number of tests.

The bot covers Galata screenshots, JSON snapshots, documentation screenshots and example snapshots.

Why the bot opens a pull request

Pushing to the contributor’s branch needs a token with write access, in a job that anyone can start with a comment. For that reason the old workflow was limited to maintainers. Our bot commits to its own fork and opens a pull request against the contributor’s branch, so it has no write access to the repository. The write happens when the contributor merges.

A GitHub App cannot do this. An App can open pull requests inside its organisation, but not against a fork owned by someone else, so we use a plain bot account with a personal access token.

The workflow also limits what it does with files from the pull request:

  • The unpacking script comes from the default branch, so a pull request cannot change it.
  • A job without the bot token downloads the test artifact, unpacks it and compresses the images. It passes on only the snapshot files.
  • Only files in snapshot directories (galata/**/*-snapshots/ andexamples/**/*-snapshots/ ) are accepted. The job that holds the token checks each path again, and refuses a path that a symbolic link redirects.
  • Git hooks are off in both jobs. The commit flag --no-verify alone skips the pre-commit hook, but notpost-checkout orpost-commit .
  • The token is in a GitHub Actions environment. Its protection rules require the approval of a maintainer for each run.

Looking at a failure without a download #

A Playwright HTML report is a directory of HTML, scripts, screenshots and videos. As a GitHub Actions artifact it is a zip, so to look at a failure you downloaded it, unzipped it, started python -m http.server and opened localhost. Most reviewers did not.

In February 2026, GitHub Actions added artifact uploads without a zip (archive: false), and it serves a single uploaded HTML file directly. We asked Playwright to make its report self-contained, and the maintainers declined because the need is niche. Our action, inline-playwright-report, does it instead. The report keeps its assets in a base64 zip inside index.html, so the action rewrites the src attributes inside that zip.

The bot comment on every JupyterLab pull request now has a badge: green when all tests pass, orange with the number of flaky tests, red with the number of failures. The link opens the report filtered to the flaky or failed tests.

When a pull request adds or changes a test, the workflow runs that test again with video recording, and the video goes into the same report.

Posting results from a fork’s pull request #

A pull_request workflow from a fork cannot write comments. The usual solution is a second workflow, triggered by workflow_run, that has write permissions and reads an artifact from the first one. Code from the fork produced that artifact, so its content is untrusted. The comment built from it appears under a trusted bot account.

The comment takes two values from the artifact: the link to the report, and the failing and flaky counts. Without checks, a fork could make the trusted bot post a phishing link, or end the markdown link early and add its own text to the comment. The action makes these checks:

  • It parses the URL with new URL() instead of inserting the string. This removes newlines and encodes angle brackets, so the value cannot escape the markdown link.
  • It requires the URL to point at an artifact of the run being reported. A URL anywhere in your own repository is not safe enough, because GitHub serves the commits of a fork under the base repository: https://github.com/you/yourrepo/blob/<sha>/evil.html can be attacker content.
  • It compares the head commit of the run with the head commit of the pull request in the artifact. This check also skips runs from an older push.
  • It edits only comments written by a bot. A “Quote reply” of the badge by a person starts with the same text.
  • It accepts the failing and flaky counts only as non-negative integers, and shows Unknown otherwise.
  • It puts the URL in angle brackets in the markdown link, because a closing parenthesis in the URL ends the link early.

We also run zizmor on the workflow files, in CI and as a pre-commit hook. It correctly flags the workflow_run trigger, so that line has an inline exemption that explains the reason.

Measuring flakiness #

Flakiness measured on pull requests includes the effects of each change. Since December 2025, JupyterLab runs the suite every six hours on main. The inputs are the same every time, so any difference between runs is flakiness.

Every Monday, a script reads the last 28 scheduled runs and posts a table to a public issue. The table lists which tests failed, which passed only on retry, how often, and in which browser. It shows when the count starts to rise again.

To check a fix, you can start the workflow by hand with a test name pattern for Playwright’s --grep and a repeat count of up to 50.

Practices and lint rules

Most flaky tests, outdated screenshots and quirks of the snapshot updating workflows had one of a few causes, and the contributing guide now lists best practices for writing UI tests:

  • Compare only one screenshot per test, unless you use expect.soft . Otherwise the first failure hides the others, and the update takes several CI rounds.
  • Capture the smallest region that shows the change.
  • Do not use waitForTimeout() .
  • Compute crop dimensions instead of hard-coding them.
  • Open notebooks without a kernel when the test does not run code.

Lint rules enforce some of these:

  • eslint-plugin-playwright , with five rules set to error:no-wait-for-timeout ,no-element-handle ,no-networkidle ,prefer-to-have-count andprefer-web-first-assertions .
  • @jupyter/eslint-plugin , with rules that report hand-written selectors where a Galata helper exists, and require soft assertions before screenshot comparisons.
  • no-restricted-syntax rules againstscreenshot({ path }) , because the result should go throughtoMatchSnapshot() , and againsttest.describe.configure({ mode: 'serial' }) , because serial tests cannot be split across shards.
  • A stylelint rule that requires CSS variables for some properties, so a hard-coded font cannot return.

The Playwright rules highlighted 118 violations in 36 files. We enabled it gradually, suppressing the existing violations with inline eslint-disable comments and slowly working through the old ones in follow-up PRs.

Sharding #

With six shards per browser, the suite takes 15 minutes instead of 55.<sup>4</sup> Set fullyParallel: true, or Playwright cannot balance the shards.

The setup step runs once per shard, so each extra minute there costs six minutes of runner time. We replaced two lines:

- playwright install-deps
- playwright install chromium
+ playwright install chromium --only-shell

install-deps installs about 80 apt packages that headless tests do not use, and the apt step failed often enough to cancel jobs. If you test with Firefox or WebKit, remove install-deps but do not add --only-shell: it applies to Chromium only.

We also removed Python test dependencies that the browser tests did not use, and a second frontend build. The browser cache key pointed at a missing file, so the cache was not refreshed when Playwright was updated. We fixed the key.

When one shard fails, you can re-run only that shard. A re-run produces blobs for its own shard only, so the merge job starts from the merged blobs of the previous attempt.

Reusable parts #

Two composite actions in jupyterlab/maintainer-tools (BSD-3-Clause) work in any repository:

These parts need a few changes before you can reuse them:

  • unpack_snapshots.py : replacegalata , the test directory name in its paths, with your own.
  • galata-update-v2.yml : change the trigger phrases, the names of the test workflow and its artifact, and the list of paths the bot may commit.
  • fonts.ts is a JupyterLab plugin. In another application, load the same@fontsource packages and CSS from an entry point that only the tests use.
  • playwright.config.js : copy the two Chromium flags into your configuration.

Open work #

  • A few tests per run are still flaky (2.4 on average in August), and the weekly report lists them.
  • 36 waitForTimeout calls remain behindeslint-disable comments.
  • Mermaid diagrams still use system fonts, so their screenshots match only on a machine with the runner’s fonts.

Credits #

This work was a collaboration with Quansight PBC. At OpenTeams, @krassowski led the work, and @MUFFANUJ and @Darshan808 worked on the flaky test fixes, the CI tooling and the lint rules.

The Jupyter Foundation funded this work. @jtpio opened the 2023 issue that described the problem, @bollwyvl asked for a readable CI report, and @jasongrout set up the scheduled runs that give the flakiness numbers in this post.

Thank you to the reviewers: @jtpio, @jasongrout, @brichet, @Yann-P, @HaudinFlorence and @mfisher87.

Counted from the AI usage section of the JupyterLab pull request template . The section asks whether AI generated some or all of the content. The count covers pull requests merged intomain from 25 June to 25 September 2026 that changed a test file ingalata/test . It leaves out backports and pull requests from bots. The other two pull requests did not answer.↩ 2. Found through flaky tests: Found while pinning the fonts:

Measured on GitHub Actions on 26 September 2026 with the Chromium tests of the jupyterlab project, in a Fedora 44 container and on the Ubuntu 26.04 runner. “Before” is JupyterLab’s main branch on 4 February 2026, before the first change of this work, with its own lockfile and Playwright version, and Python packages as of that date. A test or screenshot that passes on the retry counts as passing, as on CI.

  • The Fedora 44 run had 41 failing tests: 39 Mermaid diagrams and 2 Vega charts, each failing on its screenshot only. The Vega charts got a pinned font after the run, so “Now” counts them as matching: 216 of 255 screenshots and 539 of 578 tests.
  • Screenshots: 235 comparisons before and 255 now. A test stops at its first failing screenshot, so fewer comparisons are reached when many fail.
  • Tests: 518 before and 578 now, without the tests that the CI runner skips too. On the CI runner itself, 517 and 518 of 518 tests passed in two runs before, and 578 of 578 now.
  • The Playwright version of February refuses to install Chromium on Ubuntu 26.04, so that run used the Chromium build for Ubuntu 24.04. Most of its screenshot mismatches are coloured fringes around text: the 26.04 image turns on subpixel rendering, and the Chromium flag that turns it off came with this work. ↩

The median of the 51 runs on main fromthe update to Playwright 1.58 on 26 January 2026 to the change to six shards on 4 February. After the update, the median Firefox job took 50 minutes instead of 42. The median Chromium job took 43 minutes instead of 42. From October 2025 to the update, the median run took 43 minutes. The Firefox shards are still about 20% slower than the Chromium shards, so we compare with the runs that used the same Playwright version.↩

── more in #developer-tools 4 stories · sorted by recency
── more on @jupyterlab 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reliable-visual-regr…] indexed:0 read:17min 2026-10-06 · —