ICSE 2026
Sun 12 - Sat 18 April 2026 Rio de Janeiro, Brazil
Wed 15 Apr 2026 17:00 - 17:15 at Asia IV - AI for Software Engineering 8 Chair(s): Yintong Huo

As large language models (LLMs) become increasingly capable and widely adopted, benchmarks play a central role in assessing their practical utility. For example, SWE-Bench Verified has emerged as a critical benchmark for evaluating LLMs’ software engineering abilities, particularly their aptitude for resolving real-world GitHub issues. Recent LLMs show impressive performance on SWE-Bench Verified, leading to optimism about their capacity for complex coding tasks. However, current evaluation protocols may overstate these models’ true capabilities. It is crucial to distinguish LLMs’ generalizable problem-solving ability from memorized patterns and other learned artifacts. In this work, we introduce two new diagnostic tasks: \textit{file path identification} from issue descriptions alone and \textit{ground truth function reproduction} with only the current file context and issue description to probe models’ underlying knowledge. We present empirical evidence that performance gains on SWE-Bench Verified may be partially driven by memorization rather than genuine problem-solving. We show that state-of-the-art (SoTA) models achieve up to 76% accuracy in identifying buggy file paths using only issue descriptions, without access to repository structure. This performance is merely up to 53% on tasks from repositories not included in SWE-Bench, pointing to possible data contamination or memorization. Similar patterns are also observed for the function reproduction task, where the verbatim similarity is much higher on SWE-Bench Verified than on other similar coding benchmarks (up to 35% consecutive 5-gram overlap ratio on SWE-Bench Verified and Full, but only up to 18% for tasks in other benchmarks). These findings raise concerns about the validity of existing results and underscore the need for more robust, contamination-resistant benchmarks to reliably evaluate LLMs’ coding abilities.

Wed 15 Apr

Displayed time zone: Brasilia, Distrito Federal, Brazil change

16:00 - 17:30
AI for Software Engineering 8Research Track / SE In Practice (SEIP) at Asia IV
Chair(s): Yintong Huo Singapore Management University, Singapore
16:00
15m
Talk
Quantifying Memorization Advantage in Code LLMs
Research Track
Alberick Euraste Djire University of Luxembourg, Abdoul Kader Kaboré University of Luxembourg, Jordan Samhi University of Luxembourg, Luxembourg, Earl T. Barr University College London, Jacques Klein University of Luxembourg, Tegawendé F. Bissyandé University of Luxembourg
16:15
15m
Talk
Assessing Coherency and Consistency of Code Execution Reasoning by Large Language ModelsVirtual Attendance
Research Track
Changshu Liu University of Illinois at Urbana-Champaign, Yang Chen University of Illinois at Urbana-Champaign, Reyhaneh Jabbarvand University of Illinois at Urbana-Champaign
Pre-print Media Attached
16:30
15m
Talk
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation BenchmarkVirtual Attendance
Research Track
Dewu Zheng Sun Yat-sen University, Yanlin Wang Sun Yat-sen University, Ensheng Shi Huawei, Xilin Liu Huawei Cloud, Yuchi Ma Huawei Cloud Computing Technologies, Hongyu Zhang Chongqing University, Zibin Zheng Sun Yat-sen University
Media Attached
16:45
15m
Talk
What’s in a Benchmark? The Case of SWE-Bench in Automated Program Repair
SE In Practice (SEIP)
Matias Martinez Universitat Politècnica de Catalunya (UPC), Xavier Franch Universitat Politècnica de Catalunya
17:00
15m
Talk
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of ReasonVirtual Attendance
SE In Practice (SEIP)
Shanchao Liang Purdue University, USA, Spandan Garg Microsoft Corporation, Roshanak Zilouchian Moghaddam Microsoft
Media Attached
17:15
15m
Talk
Rethinking the Evaluation of Secure Code Generation
Research Track
Shih-Chieh Dai University of Utah, USA, Jun Xu The University of Utah, Guanhong Tao University of Utah