← Workstreams

Workstream: ScreenshotTest

Status: Planned · Component: Maximize developer productivity

Goal

Make "did this change alter what the UI looks like?" a question the build answers. A WUI repository commits golden screenshots — full pages and per-component crops — alongside its code. Every build re-renders them and fails on an unintended pixel difference; an intentional change regenerates the goldens, and the PR's image diff shows reviewers exactly what changed. The committed gallery doubles as living design documentation: browsing screenshots/ shows what every page and feature looks like today, and its git history shows how the design evolved.

The core problem: pixels depend on the renderer

A screenshot is a function of far more than the page: the exact browser build, the font files available, the fontconfig configuration that selects and hints them, the viewport and device pixel ratio, and the rasterization path all leave fingerprints in the bytes. Two healthy machines rendering identical HTML routinely disagree — which is why naive golden tests flake, and why AiCliGui's screenshot tests (the prior art for committed UI screenshots) deliberately never compare: they re-record on every run, producing a valuable visual journal but detecting nothing.

The decision here is to make rendering a hosted, version-pinned service: url://screenshottest/ is the only renderer whose pixels count. Goldens are never rasterized on a developer machine or a CI droplet — both submit the WUI to the service and receive canonical pixels back. Consistency stops being a property every environment must maintain and becomes a property one service guarantees: exact Chromium build (pinned through Playwright), bundled font set with a pinned fontconfig, fixed viewport and device pixel ratio, software rasterization, frozen clock and disabled animations. Renderer versions are explicit; a version bump is a deliberate, service-side event paired with fleet-wide golden regeneration, never ambient drift.

Current state

The core is built, deployed, and in use: the ScreenshotTestApi / Runner / Embedded / ServerService family serves url://screenshottest/ (renderer do-img-226128685+playwright-core-1.61.1+chromium-1228), ScreenshotTestWui browses sessions and diff heatmaps at screenshottest.nursery.wasmserver.com (dogfooding its own goldens), and BuildTestWui is the first adopter — four committed goldens compared on every CI run, with fail-on-diff and record-mode regeneration proven end to end. The pieces this workstream built on:

  • The PerformanceTest family (Api / Embedded / Runner / ServiceServer / WUI, url://performancetest/) is the architectural template — including its submission model, where the caller uploads the classpath bundle and the service runs it on a worker it provisions.
  • The Testing Architecture already mandates real-browser (Playwright/Puppeteer) testing for WUIs, and every org WUI exposes the injected-API createServer(port, api) seam, so a WUI can serve representative fake data hermetically — the same launch pattern the webapp-screenshots Claude Code skill uses today for one-off PR screenshots hosted in PrScreenshots.
  • ComposeErrorBoundaries explicitly defers "screenshot tests of fallback/recovered states" to frontend consumers; this workstream is where that capability lands.

Why it accelerates developers

Today a WUI change ships with, at best, hand-curated PR screenshots; nothing catches the CSS refactor that silently moved a chart axis, and reviewing a visual change means checking out the branch. With goldens: unintended visual regressions fail the build like any other broken behavior; intentional changes are reviewed as image diffs GitHub already renders inline; and agents get a one-call, environment-independent way to both verify and regenerate the visuals of any WUI they touch.

Plan

Service family (per the standard layered architecture, mirroring PerformanceTest): ScreenshotTestApi (the url://screenshottest/ contract), ScreenshotTestRunner (on-worker: launch the submitted WUI, drive the pinned browser, capture, compare), ScreenshotTestEmbedded (session orchestration and worker provisioning), ScreenshotTestServerService (hosts url://screenshottest/), ScreenshotTestWui (session gallery, diff viewer, renderer-version history).

Submission model. A render session uploads the WUI's classpath bundle plus a launcher entry point (which starts the WUI hermetically with representative fake data, exactly like the repo's other tests), a scenario list — pages to capture, each with optional named CSS-selector component crops — and the current goldens. The runner starts the WUI on the worker, captures every scenario with the pinned renderer, compares against the submitted goldens, and returns per-screenshot verdicts, fresh PNGs, and diff heatmaps. Workers are provisioned per session and torn down, PerformanceTest-style; a warm worker pool is a latency optimization on the roadmap, not an architectural change.

Verdict semantics. A small per-pixel tolerance absorbs rasterizer noise; beyond it, a difference fails the test — unlike PerformanceTest's always-informational verdicts, because in a pinned renderer a pixel difference is deterministic evidence of change, not noise to be weighed. A missing golden fails too (new scenarios must commit their golden). Regeneration is explicit: the same session run in record mode returns PNGs the developer (or agent) commits — so an intentional change is a code diff plus an image diff in the same PR.

Where goldens live. In the WUI repository, full pages and component crops both, so GitHub's PR image diff is the review surface and the repository itself documents the design. Every golden set carries the renderer version that produced it; a session submitted with a mismatched renderer version fails with a message naming both versions rather than producing confusable diffs.

CI integration. Phase one needs no build-system changes: the golden check is an ordinary kompile test that calls url://screenshottest/ through a small client library and asserts the verdicts. Once Pluggable Execution Environments lands, the same check graduates to an @ExecutionEnvironment("url://screenshottest/...")-routed test — making this, alongside PerformanceTest, the second concrete runner that plan's routing generalizes.

The roadmap, in order:

  • [x] Contract and renderer. ScreenshotTestApi (session submission, scenarios, verdicts, record mode, renderer-version handshake) and ScreenshotTestRunner (pinned Playwright-Chromium + bundled fonts/fontconfig, hermetic WUI launch, capture, tolerance compare, diff heatmaps).
  • [x] Hosted service. ScreenshotTestEmbedded worker orchestration + ScreenshotTestServerService binding url://screenshottest/, deployed and holding recent sessions for inspection.
  • [x] First adopter. BuildTestWui commits goldens for its run-detail and runs-list pages (including the test-progress chart states) with the client-library test wired into its normal CI; an intentional-change demo proves fail-on-diff and record-mode regeneration end to end.
  • [x] Session WUI. ScreenshotTestWui: browse sessions, screenshots, and diff heatmaps; renderer-version history — dogfooding its own goldens.
  • [ ] Fleet adoption and regen tooling. Goldens across the WUI fleet; one-command regeneration per repo; a renderer-version bump runbook that pairs the bump with fleet-wide regeneration.
  • [ ] Warm worker pool to bring session latency from droplet-provisioning minutes toward seconds.
  • [ ] @ExecutionEnvironment routing once ExecutionEnvironments ships its annotation and dispatcher.

Related

Graduation

This workstream graduates when the service family is deployed, at least three WUI repositories run golden checks in their normal CI with goldens committed in-repo, and a renderer-version bump has been executed once with fleet regeneration. At that point the service documents itself as a project page in the Documentation Repository, the golden-test authoring guidance folds into the Testing Architecture documentation, and this page becomes a stub pointing at both.