Automated Reproduction of Android Application Bugs with LLMs: Are We There Yet?
Automated reproduction of software bugs from natural-language reports remains a challenging problem, particularly for mobile applications where reproduction often requires complex user interface interactions. Recent advances in large language models (LLMs) have sparked interest in leveraging generative models for Android bug reproduction. However, most existing approaches evaluate success primarily based on whether a failure is triggered during execution, providing limited insight into the semantic fidelity of generated reproduction sequences.
In this paper, we present a reflective empirical study assessing the capability of an LLM to reproduce Android application bugs directly from natural-language reports using the ANDROR2 dataset of 90 manually validated bug reports and ground-truth UIAUTOMATOR scripts. We evaluate two complementary tasks: failure-type classification and executable reproduction script generation. To assess reproduction fidelity, we introduce critical-step coverage, an intent-level metric that measures whether generated scripts capture the essential user actions required to trigger a failure.
Our results reveal an important asymmetry: while classification performance varies significantly across defect categories (70% overall accuracy), reproduction scripts achieve high intent coverage (mean 89.5%). These findings suggest that outcome-based success metrics alone are insufficient to characterize automated bug reproduction quality and motivate the need for intent-aware evaluation frameworks for LLM-based testing.
Wed 20 MayDisplayed time zone: Seoul change
10:30 - 12:00 | Web Application Automated TestingResearch Papers / Short Papers, Vision and Emerging Results at Room 101 Chair(s): Tommaso Fulcini Politecnico di Torino | ||
10:30 25mTalk | Leveraging Large Language Models for Trustworthiness Assessment of Web Applications Research Papers Oleksandr Yarotskyi University of Coimbra, José D'Abruzzo Pereira University of Coimbra, João R. Campos University of Coimbra | ||
10:55 25mTalk | Towards Automated Page Object Generation for Web Testing using Large Language Models Research Papers Betül Karagöz Technical University of Munich, Filippo Ricca DIBRIS, Università di Genova, Matteo Biagiola University of St. Gallen and Università della Svizzera italiana, Andrea Stocco Technical University of Munich, fortiss Pre-print | ||
11:20 25mTalk | Neural Embeddings for Web Testing Research Papers Kasun Kanaththage Technical University of Munich, Luigi Libero Lucio Starace Università degli Studi di Napoli Federico II, Matteo Biagiola University of St. Gallen and Università della Svizzera italiana, Paolo Tonella USI Lugano, Andrea Stocco Technical University of Munich, fortiss Pre-print | ||
11:45 15mTalk | Automated Reproduction of Android Application Bugs with LLMs: Are We There Yet? Short Papers, Vision and Emerging Results Dennis Carey Florida Polytechnic University, Karim Elish Florida Polytechnic University, Paniz Abedin Florida Polytechnic University | ||