AddThisFeature

Visual Regression Testing

Catch unintended visual changes before they reach users, not after.

involved Testing & QA

What it adds

An automated comparison of rendered screenshots against approved baselines, run on every change.

What your agent is told to do

5
  1. 1

    Identify the surfaces where a silent visual change would actually hurt — the shared components, the pricing table, the checkout, the dashboard — and capture those rather than screenshotting every route in the app.

  2. 2

    Freeze everything that varies between runs before capturing: load fonts and wait for them, fix the clock to a known instant, disable animations and transitions, and render from seeded fixtures rather than live data.

  3. 3

    Capture each surface in every theme and at every viewport the app claims to support, and label each capture with that combination so a diff says exactly which one moved.

  4. 4

    Fail the run on a diff and require a human to look at it, but make approving a genuine change a single deliberate action rather than a reason to disable the check.

  5. 5

    Storage, review, ownership, and expiry of the approved images belong to Screenshot Diff Baselines; this brief owns capture, stabilisation, and comparison. Do not build a second baseline store here.

Edge cases it handles

7
  • Unstable input produces diffs that mean nothing — a system font substituting for a web font, a relative timestamp ticking over, a spinner caught mid-rotation, or a random fixture name will all fail a run that should have passed.
  • A component that looks right in light mode at desktop width can be broken in dark mode on a phone, so a matrix that captures one theme at one viewport is a matrix that misses most regressions.
  • The run must make an intentional redesign distinguishable from flake, or reviewers learn to approve everything without looking and the suite stops protecting anything.
  • Granularity has to be chosen deliberately: whole-page shots turn one button change into forty diffs, and component-only shots miss layout collisions that only appear when components sit together.
  • Anti-aliasing and sub-pixel text rendering differ between machines, so a comparison with a zero-tolerance threshold will fail on a developer laptop even when nothing changed.
  • Content that legitimately varies in height — a truncated description, an error banner, a lazily loaded image — must be settled before capture, or the shot records a mid-load frame.
  • A newly added surface has no baseline. Treat the first capture as pending review rather than an automatic pass.

Definition of done

8
  • Fonts, dates, animations, and fixture data are deterministic, and re-running the suite on an unchanged codebase produces no diffs.
  • Every captured surface exists in each supported theme and viewport, and the diff report names the combination that changed.
  • A visual difference fails the run and blocks the merge until a person reviews it.
  • The capture set covers both isolated components and the assembled pages they appear on.
  • Approved images are read from the baseline store defined by Screenshot Diff Baselines rather than a second copy kept here.
  • A surface with no baseline is reported as new and awaiting review, never silently passed.
  • The feature matches the existing design system.
  • No existing functionality is broken.

Related features

How it works

  1. 1

    Copy the link

    Grab the Markdown instruction URL for this feature.

  2. 2

    Give it to your AI

    Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.

  3. 3

    It inspects, then implements

    Your agent reads your existing app first, then adds the feature to fit it.

Works with your stack

These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.

Need it tighter than that? Customize the feature and tell it exactly what you're running.