PROFES 2025
Mon 1 - Wed 3 December 2025 Salerno , Italy
Mon 1 Dec 2025 16:45 - 17:10 at Sala degli Affreschi (Fresco Room) - QuEMaLeS Session 2

Flaky tests yield inconsistent results without code changes, which undermines software reliability and may increase development costs, emphasizing the importance of effective detection methods. Despite various research efforts over the past 15 years, existing techniques often show limited accuracy and adoption. This study explores whether commonly available Large Language Models (LLMs) are suitable for detecting flaky tests in software. Using the International Dataset of Flaky Tests, we asked selected LLMs, including GPT and Gemini, to classify Java test cases as flaky or non-flaky. Results show that LLMs are unable to do so consistently. This research underscores the challenges of using general-purpose LLMs for flaky test detection and highlights the need for more effective solutions.

Mon 1 Dec

Displayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change

16:00 - 17:30
16:00
20m
Other
Discussion on Session 1
Quality Evaluation of ML-based Software Systems

16:20
25m
Paper
Software Product Quality: Some Thoughts about its Evolution and Perspectives in the AI years
Quality Evaluation of ML-based Software Systems
Luigi Buglione DXC Technology, Francesco Merola Istituto di Scienza e Tecnologie dell'Informazione "Alessandro Faedo"
16:45
25m
Paper
Chatting about flaky tests with standard LLMs. An empirical exploration
Quality Evaluation of ML-based Software Systems
Marcin Szwarc Poznań University of Technology, Poland, Bartosz Walter Poznań University of Technology, Poland
17:10
20m
Other
Discussion on Session 2
Quality Evaluation of ML-based Software Systems