What Are Developers Actually Discussing When Visual Regression Tests Fail?
This program is tentative and subject to change.
Visual Regression Tests (VRTs) are widely adopted as a mechanism for detecting unintended visual changes in user interfaces. By design, VRTs operate on rendered pixel output, and the prevailing assumption is that they catch stylistic regressions such as layout shifts, color mismatches, and font alterations. We conduct an empirical analysis of 307 pull requests (PRs) from 103 GitHub repositories that incorporate VRT results via Chromatic, comparing them against 299 PRs that contain image attachments but no VRT (Visual PRs). Quantitatively, VRT-PRs exhibit a 5.3% higher acceptance rate, a 3.8 times longer median resolution time, 10 times more discussion comments, and 1.75 to 4.5 times larger code changes than Visual PRs.
VRT results are typically shared around the midpoint of the review process, sustaining ongoing discussion rather than serving only as a final check. Through a card-sorting analysis of 189 VRT-flagged issues, we identify seven defect categories: Layout (39.7%), Appearance (27.5%), Color (14.8%), Text (9.5%), State (6.9%), Test (6.3%), and Image (4.2%). A substantial fraction of these detections correspond to functional defects whose visual symptoms reveal underlying problems, including undefined component state (13 cases), contents disappearance (17 cases across multiple categories), and visually imperceptible regressions (5 cases). We further document cases in which VRT detected visual regressions originating from code changes in seemingly unrelated files, exposing non-local effects that no targeted test would have been written to catch. These observations reposition VRT from a stylistic verification tool to a detection mechanism for the unintended consequences of code changes, with implications for how VRT should be integrated into the maintenance toolchain.
This program is tentative and subject to change.
Thu 17 SepDisplayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change
11:00 - 12:30 | Session 16 - Scaling Smarter: Data, Deployment & DiscoveryTool Demonstration and Data Showcase Track / Industry Track / Visions and Emerging Results Track / Research Papers Track at A59S Theme: Software Architecture & Evolution | ||
11:00 20mPaper | Predictive Test Selection Without Historical Failure Data Industry Track Maximilian Jungwirth BMW Group, University of Passau, Raphael Nömmer Saarland University, CQSE GmbH, Andreas Stahlbauer CQSE GmbH, Sven Apel Saarland University, Gordon Fraser University of Passau | ||
11:20 20mPaper | Evaluation-Strategy Gap in Fault Diagnosis of Deep Learning Programs Research Papers Track Sigma Jahan Dalhousie University | ||
11:40 20mPaper | Taming Revocable Spot VM Instances with Diffused Research Papers Track Zhao Zhang Georgetown University, Zeying Zhu University of Maryland, Micah Sherr Georgetown University, Benjamin Ujcich Georgetown University, Wenchao Zhou Georgetown University | ||
12:00 10mShort-paper | GHARuns: A large dataset of GitHub Actions workflow execution metadata Tool Demonstration and Data Showcase Track Aref Talebzadeh Bardsiri University of Mons, Belgium, Tom Mens University of Mons, Alexandre Decan University of Mons; F.R.S.-FNRS Pre-print | ||
12:10 10mShort-paper | What Are Developers Actually Discussing When Visual Regression Tests Fail? Visions and Emerging Results Track Miku Watanabe Nara College, National Institute of Technology/Nara Institute of Science and Technology, Kosei Horikawa Nara Institute of Science and Technology, Brittany Reid Nara Institute of Science and Technology, Yutaro Kashiwa Nara Institute of Science and Technology, Hajimu Iida Nara Institute of Science and Technology Pre-print | ||