Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation
Research shows that errors in natural language can be corrected by translating texts to another language and back using language models. We explore to what extent this latent correction capability extends to Automated Program Repair (APR) by investigating Round-Trip Translation (RTT): translating code from one programming language into another programming or natural language and back, using Large Language Models (LLMs). We hypothesize that RTT restores patterns most commonly seen in the LLM’s training corpora through regression toward the mean, replacing infrequent bugs with more frequent, natural, bug-free code. To test this hypothesis, we employ nine LLMs and four common APR benchmarks in Java, and perform a detailed quantitative and qualitative analysis of RTT-generated patches. We find that RTT through English generates plausible patches for 100 of 164 bugs with GPT-4 on the HumanEval-Java benchmark, and 97 are found to be correct in our manual assessment. Moreover, RTT uniquely generates plausible patches for 46 bugs that were missed by LLMs specifically fine-tuned for APR. While this demonstrates the viability of RTT for APR, we also observe limitations, such as a lower overall bug fix rate than the state-of-the-art and diluting the original coding style. We analyze the impact of these limitations and discuss the potential of using RTT as a complementary component in APR frameworks.
Fri 17 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
14:00 - 15:30 | AI for Software Engineering 23Research Track / Demonstrations / Journal-first Papers at Asia I Chair(s): Wesley K.G. Assunção North Carolina State University | ||
14:00 15mTalk | CI-Bench: A Framework for Evaluating Large Language Model Tools on CI Failures Demonstrations Raian Latif Nabil University of California, Davis, Hao-Nan Zhu University of California, Davis, Cindy Rubio-González University of California at Davis | ||
14:15 15mTalk | Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation Journal-first Papers Fernando Vallecillos Ruiz Simula Research Laboratory, Anastasiia Grishina Simula Research Laboratory, Max Hort Simula Research Laboratory, Leon Moonen Simula Research Laboratory Link to publication Pre-print | ||
14:30 15mTalk | XRFix: Exploring Performance Bug Repair of Extended Reality Applications with Large Language Models Research Track Jingwen Wu Department of Computer Science, Hong Kong Baptist University, Hanyang Guo School of Software Engineering, Sun Yat-sen University, Hong-Ning Dai Department of Computer Science, Hong Kong Baptist University, Xiapu Luo Hong Kong Polytechnic University DOI Pre-print | ||
14:45 15mTalk | Synthetic Repo-level Bug Dataset for Training Automated Program Repair ModelsDistinguished Paper Award Research Track Minh V. T. Pham FPT Software AI Center, Huy N. Phan FPT Software AI Center, Hoang Nhat Phan Nanyang Technological University, Cuong Chi Le The University of Texas at Dallas, Tien N. Nguyen University of Texas at Dallas, Nghi D. Q. Bui Google Research | ||
15:00 15mTalk | PredicateFix: Repairing Static Analysis Alerts with Bridging Predicates Research Track Yuan-An Xiao Peking University, Weixuan Wang Peking University, Dong Liu Center Research Institute, ZTE Coporation, China, Junwei Zhou Center Research Institute, ZTE Coporation, China, Shengyu Cheng ZTE Corporation, Yingfei Xiong Peking University Pre-print | ||
15:15 15mTalk | Input Reduction Enhanced LLM-based Program Repair Research Track Boyang Yang Yanshan University, Luyao Ren Peking University, Xin Yin Zhejiang University, Jiadong Ren Yanshan University, Haoye Tian Aalto University, Shunfu Jin Yanshan University DOI Pre-print | ||