I Will Try to Fix You: Large Language Models for Mobile GUI Test Repair
Mobile applications evolve rapidly, forcing development teams to deliver updates at high frequency while maintaining reliable testing processes. In this context, GUI testing is essential for validating user-facing behavior, yet it is highly affected by fragility: tests break across application versions not because of functional defects, but due to changes in the GUI structure, appearance, or properties. This leads to substantial manual effort to diagnose failures and repair outdated tests.
This study aims to: (i) identify the most common causes of mobile GUI test breakages across application versions; (ii) assess how Large Language Models (LLMs) can reduce the effort required for repairing broken tests; and (iii) compare an LLM-based repair strategy with a state-of-the-art automated repair tool, Healenium-Appium.
A total of 61 broken GUI tests from 19 real-world Android applications were analyzed to identify the underlying causes of breakage. We then developed an LLM-based repair approach relying on iterative interactions with the model to generate repaired test scripts. The approach was experimentally evaluated against Healenium-Appium.
The LLM-based method successfully repaired 45 out of 61 tests (73.8%) after a single interaction and 56 out of 61 tests (91.8%) after multiple interactions, outperforming Healenium-Appium by 50.8% and 68.8%, respectively.
The results indicate that LLM-based repair is a highly effective solution for mitigating GUI test fragility. Integrating LLMs into mobile testing workflows can substantially reduce maintenance effort while delivering higher repair success rates than existing automated tools, thereby improving the reliability and scalability of test suites as applications evolve.
Tue 17 MarDisplayed time zone: Athens change
14:00 - 15:30 | |||
14:00 25mTalk | Metamorphic Testing for Sequential Prediction Models: A Survey of LSTMs and LLMs Workshops & Tutorials Alejandra Duque-Torres Software Competence Center Hagenberg (SCCH) GmbH, Stefan Fischer Software Competence Center Hagenberg, Claus Klammer Software Competence Center Hagenberg | ||
14:25 25mTalk | TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance Workshops & Tutorials Elena Bruches Siberian Neuronets LLC, Vadim Alperovich T-Technologies, Dari Baturova Siberian Neuronets LLC, Roman Derunets Siberian Neuronets LLC, Daniil Grebenkin Siberian Neuronets LLC, Georgiy Mkrtchyan T-Technologies, Oleg Sedukhin Siberian Neuronets LLC, Mikhail Klementev Siberian Neuronets LLC, Ivan Bondarenko Novosibirsk State University, Nikolay Bushkov T-Technologies, Stanislav Moiseev T-Technologies Pre-print | ||
14:50 25mTalk | I Will Try to Fix You: Large Language Models for Mobile GUI Test Repair Workshops & Tutorials Tommaso Fulcini Politecnico di Torino, Alessandro Poletti Politecnico di Torino, Anna Arnaudo Politecnico di Torino, Riccardo Coppola Politecnico di Torino | ||