ICST 2026
Mon 18 - Fri 22 May 2026 Daejeon, South Korea
Tue 19 May 2026 16:50 - 17:15 at Room 101 - LLM-Assisted Test Generation Chair(s): Shifat Sahariar Bhuiyan

Automated Program Repair (APR) suffers from the overfitting problem, in which generated patches pass the available test suite while remaining semantically incorrect. FixCheck, a recent Automated Patch Correctness Assessment (APCA) technique, aims to mitigate this issue by leveraging large language models (LLMs) to generate additional tests that expose incorrect patches. However, our empirical analysis shows that FixCheck’s LLM-based assertion generation frequently produces semantically weak assertions or non-compiling tests, substantially limiting its effectiveness. To address these limitations, we propose six prompt engineering strategies and systematically evaluate them on 109 incorrect patches from the Defects4J benchmark using LLMs with varying capacities, including GPT-4o, GPT-4o-mini, and Llama~3.2~3B. Our results demonstrate that the effectiveness of prompt engineering is strongly model-dependent. For GPT-based models, explicitly aligning the APCA objective through role definition improves incorrect patch detection by up to 7.5%. In contrast, for Llama~3.2~3B, enforcing strict output format constraints with illustrative examples reduces non-compiling assertions and improves detection performance by up to 19.9%. Overall, this study provides practical, model-aware prompt design guidelines for building reliable LLM-based APCA systems.

Tue 19 May

Displayed time zone: Seoul change

16:00 - 17:30
LLM-Assisted Test GenerationShort Papers, Vision and Emerging Results / Research Papers at Room 101
Chair(s): Shifat Sahariar Bhuiyan Università della Svizzera italiana
16:00
25m
Talk
Consistency Meets Verification: Enhancing Test Generation Quality in Large Language Models Without Ground-Truth Solutions
Research Papers
Hamed Taherkhani York University, Alireza Daghighfarsoodeh York University, Mohammad Chowdhury York University, Hung Viet Pham York University, Hadi Hemmati York University
16:25
25m
Talk
How well LLM-based test generation techniques perform with newer LLM versions?
Research Papers
Michael Konstantinou University of Luxembourg, Renzo Degiovanni Luxembourg Institute of Science and Technology, Mike Papadakis University of Luxembourg
16:50
25m
Talk
Improving Automated Patch Correctness Assessment by Designing LLM-Based OraclesArtifact ReviewedArtifact Available
Research Papers
Inyeong Jang Duksung Women's University, Jinyoung Kim Sungkyunkwan University
17:15
15m
Talk
Developer vs. DSpot vs. ChatGPT: A Comparative Study of JUnit Test Amplification
Short Papers, Vision and Emerging Results
David Onyango Owuor North Dakota State University, Ajay Jha North Dakota State University
Pre-print