The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
As large language models (LLMs) become increasingly capable and widely adopted, benchmarks play a central role in assessing their practical utility. For example, SWE-Bench Verified has emerged as a critical benchmark for evaluating LLMs’ software engineering abilities, particularly their aptitude for resolving real-world GitHub issues. Recent LLMs show impressive performance on SWE-Bench Verified, leading to optimism about their capacity for complex coding tasks. However, current evaluation protocols may overstate these models’ true capabilities. It is crucial to distinguish LLMs’ generalizable problem-solving ability from memorized patterns and other learned artifacts. In this work, we introduce two new diagnostic tasks: \textit{file path identification} from issue descriptions alone and \textit{ground truth function reproduction} with only the current file context and issue description to probe models’ underlying knowledge. We present empirical evidence that performance gains on SWE-Bench Verified may be partially driven by memorization rather than genuine problem-solving. We show that state-of-the-art (SoTA) models achieve up to 76% accuracy in identifying buggy file paths using only issue descriptions, without access to repository structure. This performance is merely up to 53% on tasks from repositories not included in SWE-Bench, pointing to possible data contamination or memorization. Similar patterns are also observed for the function reproduction task, where the verbatim similarity is much higher on SWE-Bench Verified than on other similar coding benchmarks (up to 35% consecutive 5-gram overlap ratio on SWE-Bench Verified and Full, but only up to 18% for tasks in other benchmarks). These findings raise concerns about the validity of existing results and underscore the need for more robust, contamination-resistant benchmarks to reliably evaluate LLMs’ coding abilities.
Wed 15 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
16:00 - 17:30 | AI for Software Engineering 8Research Track / SE In Practice (SEIP) at Asia IV Chair(s): Yintong Huo Singapore Management University, Singapore | ||
16:00 15mTalk | Quantifying Memorization Advantage in Code LLMs Research Track Alberick Euraste Djire University of Luxembourg, Abdoul Kader Kaboré University of Luxembourg, Jordan Samhi University of Luxembourg, Luxembourg, Earl T. Barr University College London, Jacques Klein University of Luxembourg, Tegawendé F. Bissyandé University of Luxembourg | ||
16:15 15mTalk | Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models Research Track Changshu Liu University of Illinois at Urbana-Champaign, Yang Chen University of Illinois at Urbana-Champaign, Reyhaneh Jabbarvand University of Illinois at Urbana-Champaign Pre-print Media Attached | ||
16:30 15mTalk | Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark Research Track Dewu Zheng Sun Yat-sen University, Yanlin Wang Sun Yat-sen University, Ensheng Shi Huawei, Xilin Liu Huawei Cloud, Yuchi Ma Huawei Cloud Computing Technologies, Hongyu Zhang Chongqing University, Zibin Zheng Sun Yat-sen University Media Attached | ||
16:45 15mTalk | What’s in a Benchmark? The Case of SWE-Bench in Automated Program Repair SE In Practice (SEIP) Matias Martinez Universitat Politècnica de Catalunya (UPC), Xavier Franch Universitat Politècnica de Catalunya | ||
17:00 15mTalk | The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason SE In Practice (SEIP) Shanchao Liang Purdue University, USA, Spandan Garg Microsoft Corporation, Roshanak Zilouchian Moghaddam Microsoft Media Attached | ||
17:15 15mTalk | Rethinking the Evaluation of Secure Code Generation Research Track Shih-Chieh Dai University of Utah, USA, Jun Xu The University of Utah, Guanhong Tao University of Utah | ||