Are We Building on Unreliable Ground? A Case Study on Unresolvable Pairs in Learning-based Code Repair
Learning-based code repair models rely heavily on benchmark datasets for both training and evaluation. CodeXGLUE, one of the most widely used benchmarks in the field, serves as the standard for evaluating many of these models. However, the quality of such benchmarks is often overlooked, which can have a substantial impact on the validity of models evaluation. This case study focuses on the issue of unresolvable bug-fix pairs within CodeXGLUE—pairs that cannot be resolved at the method-level granularity. These pairs are potentially skewing both model evaluation and training outcomes. Given the widespread use of CodeXGLUE, flaws in this benchmark mean that the evaluation of many models could be problematic, raising concerns about the reliability of results across the field. To assess the impact of these unresolvable pairs, we systematically classify and remove them from CodeXGLUE, then use it to reassess the models’ performance. Our findings demonstrate that the presence of unresolvable bug-fix pairs can mislead the evaluation and even result in incorrect ranking of different models performance. Moreover, after removing these pairs from CodeXGLUE, the latest four models show an average performance improvement of 5.3% when evaluated on Defects4J. Ultimately, our case study casts doubt on the reliability of benchmarks with unresolvable pairs for learning-based code repair.
Tue 14 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
11:00 - 12:30 | |||
11:00 5mTalk | Fixing Less by Preventing More: Semantic Checklists for Robust Code Translation Journal Ahead Workshop (JAWs) Penghao Jiang University of New South Wales, Ruijun Feng University of New South Wales, Xiao Cheng Macquarie University, Jiaojiao Jiang University of New South Wales, Yulei Sui University of New South Wales | ||
11:05 5mTalk | Executable but Not Reproducible? An Empirical Study of Code Clone Detection Tools Journal Ahead Workshop (JAWs) Palash Ranjan Roy University of Saskatchewan, Banani Roy University of Saskatchewan, Kevin Schneider University of Saskatchewan, Chanchal K. Roy University of Saskatchewan | ||
11:10 5mTalk | ARMS: A Vision for Actor Reputation Metric Systems in the Open-Source Software Supply Chain Journal Ahead Workshop (JAWs) Kelechi G. Kalu Purdue University, Sofia Okorafor Purdue University, Betül Durak Microsoft Research, Kim Laine Microsoft Research, Redmond, Radames Cruz Moreno Microsoft Research, Santiago Torres-Arias Purdue University, James C. Davis Purdue University Pre-print | ||
11:15 5mTalk | Practitioners’ Experiences and Expectations about Software Sustainability in Industry: A Semi-Structured Interview Study Journal Ahead Workshop (JAWs) Jennifer Gross Uppsala University, Aaliyah Chang Queen's University, Mariam Guizani Queen's University, Canada, Sofia Ouhbi Uppsala University, Tobias Wrigstad Uppsala University | ||
11:20 5mTalk | Compartmentalization-Aware Automated Program Repair Journal Ahead Workshop (JAWs) Jia Hu The University of Manchester, Youcheng Sun MBZUAI, Pierre Olivier The University of Manchester | ||
11:25 5mTalk | Code Comprehension Beyond Best Practices: Exploring the Developer Cognitive Spectrum Journal Ahead Workshop (JAWs) Faith Culas University of Auckland, Reid Holmes University of British Columbia, Thomas Fritz University of Zurich, Priyanka Dhopade University of Auckland, Kelly Blincoe University of Auckland | ||
11:30 5mTalk | Search-Based Evolutionary Data Pruning for Class-Level Code Summarization Journal Ahead Workshop (JAWs) Joseph Call William & Mary, Daniele Bifolco University of Sannio, Massimiliano Di Penta University of Sannio, Italy, Antonio Mastropaolo William and Mary, USA | ||
11:35 5mTalk | Are We Building on Unreliable Ground? A Case Study on Unresolvable Pairs in Learning-based Code Repair Journal Ahead Workshop (JAWs) Shihao Weng Nanjing University, Yang Feng Nanjing University, xinguohua Tianjin University, Zhenlun Zhang Nanjing University, Yining Yin Nanjing University, Jia Liu Nanjing University | ||
11:40 5mTalk | Operationalizing Research Software for Supply Chain Security Journal Ahead Workshop (JAWs) Kelechi G. Kalu Purdue University, Soham Rattan Purdue University, Taylor R. Schorlemmer Purdue University, George K. Thiruvathukal Loyola University Chicago, Jeff Carver University of Alabama, James C. Davis Purdue University Pre-print | ||
11:45 5mTalk | Static and Semantic Program Slicing for Quantum Programs Journal Ahead Workshop (JAWs) Hakam W. Alomari Miami University | ||
11:50 5mTalk | ScannerTrap: Benchmarking the Robustness of Web Vulnerability Scanners in Complex Application Environments Journal Ahead Workshop (JAWs) Weizhe Wang Tianjin University, Yao Zhang Tianjin University, Hao Liu QAX Technology Group Inc, Shuai Hu State Grid Xinjiang Electric Power Research Institute, Guangquan Xu School of Cybersecurity, Tianjin University, Bin Wu Tianjin University | ||
11:55 25mPanel | Panel Discussion: Repair, Evolution, Comprehension, and Security Journal Ahead Workshop (JAWs) | ||
12:20 10mAwards | Selection of award presentations Journal Ahead Workshop (JAWs) | ||