On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
Large Language Models (LLMs) have been recently exploited as judges for complex natural language processing tasks, such as Q&A (Question & Answer). The basic idea is to delegate to an LLM the assessment of the “quality” of the output provided by an automated technique (often another LLM) for tasks for which: (i) quantitative metrics would only tell part of the story, and; (ii) a large-scale human-based evaluation would be too expensive. LLMs-as-a-judge, if proven effective for a specific task, can also unlock new possibilities for automation, with several LLMs proposing a solution for a given instance of the task (e.g., an answer to a question) and others judging and deciding what is the best output to show the user. We study the effectiveness of LLMs-as-a-judge for two code- related tasks, namely code generation and code summarization. The rationale for choosing these tasks is two-fold. First, quantitative metrics are usually not enough for the assessment of code summa- rizers/generators. For example, it is well documented that metrics such as BLEU are quite weak proxies for the quality of the generated summaries. Second, even state-of-the-art techniques still struggle with handling complex instances of these tasks (e.g., summarizing a quite long / complex function), making them good candidates for benefiting from more advanced solutions envisioning collaboration among LLMs. For code generation, we check whether eight LLMs are able to judge the correctness of 1,405 Java methods and 1,281 Python functions generated by the same LLMs or implemented by humans. For code summarization, we compare the judgment of five LLMs to those provided by nine humans for ∼ 1.2k summaries, related to both Java and Python functions. Our findings show that GPT-4-turbo is the best LLM in terms of judging capabilities for both tasks, with “smaller” LLMs featuring tens of billions parameters not being able to cope with judging tasks. However, even the best-performing LLM frequently misjudges the correctness of the code and summary quality.
Fri 17 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
11:00 - 12:30 | AI for Software Engineering 20New Ideas and Emerging Results (NIER) / Research Track / Journal-first Papers at Asia I Chair(s): Ipek Ozkaya Carnegie Mellon University | ||
11:00 15mTalk | Is Hyper-Parameter Optimization Different for Software Analytics? Journal-first Papers Link to publication Pre-print | ||
11:15 15mTalk | On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization Journal-first Papers Giuseppe Crupi Università della Svizzera italiana, Rosalia Tufano Università della Svizzera Italiana, Alejandro Velasco William & Mary, Antonio Mastropaolo William and Mary, USA, Denys Poshyvanyk William & Mary, Gabriele Bavota Software Institute @ Università della Svizzera Italiana | ||
11:30 15mTalk | A Catalog of Data Smells for Coding Tasks Journal-first Papers Antonio Vitale Politecnico di Torino, University of Molise, Rocco Oliveto University of Molise, Simone Scalabrino University of Molise Link to publication | ||
11:45 15mTalk | Towards Automating Domain-Specific Data Generation for Text-to-SQL: A Comprehensive Approach Journal-first Papers Salmane Chafik UM6P College of Computing, Saad Ezzini King Fahd University of Petroleum and Minerals, Ismail Berrada UM6P College of Computing Link to publication DOI Pre-print File Attached | ||
12:00 15mTalk | Empirical and Sustainability Aspects of Software Engineering Research in the Era of Large Language Models: A Reflection New Ideas and Emerging Results (NIER) David Williams University College London, Maria Kechagia National and Kapodistrian University of Athens, Max Hort Simula Research Laboratory, Aldeida Aleti Monash University, Justyna Petke University College London, Federica Sarro University College London | ||
12:15 15mTalk | FORGE: An LLM-driven Framework for Large-Scale Smart Contract Vulnerability Dataset Construction Research Track Jiachi Chen Sun Yat-sen University, Yiming Shen Sun Yat-sen University, Jiashuo Zhang Peking University, China, Zihao Li Hong Kong Polytechnic University, John Grundy Monash University, Zhenzhe Shao Sun Yat-sen University, Yanlin Wang Sun Yat-sen University, Jiashui Wang Zhejiang University, Ting Chen University of Electronic Science and Technology of China, Zibin Zheng Sun Yat-sen University Pre-print Media Attached File Attached | ||