Chatting about flaky tests with standard LLMs. An empirical exploration
Flaky tests yield inconsistent results without code changes, which undermines software reliability and may increase development costs, emphasizing the importance of effective detection methods. Despite various research efforts over the past 15 years, existing techniques often show limited accuracy and adoption. This study explores whether commonly available Large Language Models (LLMs) are suitable for detecting flaky tests in software. Using the International Dataset of Flaky Tests, we asked selected LLMs, including GPT and Gemini, to classify Java test cases as flaky or non-flaky. Results show that LLMs are unable to do so consistently. This research underscores the challenges of using general-purpose LLMs for flaky test detection and highlights the need for more effective solutions.
Mon 1 DecDisplayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change
16:00 - 17:30 | QuEMaLeS Session 2Quality Evaluation of ML-based Software Systems at Sala degli Affreschi (Fresco Room) | ||
16:00 20mOther | Discussion on Session 1 Quality Evaluation of ML-based Software Systems | ||
16:20 25mPaper | Software Product Quality: Some Thoughts about its Evolution and Perspectives in the AI years Quality Evaluation of ML-based Software Systems Luigi Buglione DXC Technology, Francesco Merola Istituto di Scienza e Tecnologie dell'Informazione "Alessandro Faedo" | ||
16:45 25mPaper | Chatting about flaky tests with standard LLMs. An empirical exploration Quality Evaluation of ML-based Software Systems Marcin Szwarc Poznań University of Technology, Poland, Bartosz Walter Poznań University of Technology, Poland | ||
17:10 20mOther | Discussion on Session 2 Quality Evaluation of ML-based Software Systems | ||