Can LLMs Capture Developers' Practical Needs for Code Summarization? An Empirical Study of Code Summary Optimization Behaviors on GitHub
Large Language Models (LLMs) have recently demonstrated strong performance in automatic code summarization. However, it remains unclear whether this capability truly aligns with developers’ practical needs during real-world software maintenance. Critically, the community lacks an effective benchmark and evaluation methodology to systematically assess this alignment. To address this gap, we first propose \textbf{CSO-Bench}, a novel benchmark comprising 4{,}845 Python instances (3{,}845 method-level and 1{,}000 class-level). It accurately captures real-world \textbf{C}ode \textbf{S}ummary \textbf{O}ptimization practices, where \emph{developers actively refine summaries without modifying the underlying code logic}. We characterize these refinement behaviors across four rigorous quality dimensions: factuality, completeness, clarity, and compliance. To enable systematic assessment, we innovatively introduce two complementary tasks—\emph{Summary Judgment} and \emph{Summary Editing}—to evaluate LLMs’ ability to recognize and improve summary quality. Leveraging a comprehensive framework that integrates automated metrics with a verified LLM-as-a-Judge protocol, we systematically evaluate seven mainstream LLMs (7B\textasciitilde 685B parameters) under four prompting strategies (i.e., Vanilla, CoT, Few-Shot, and RAG). Our extensive experiments reveals a significant gap between current LLMs’ capabilities and developers’ multi-dimensional quality requirements. Through an in-depth root cause analysis, we distill pervasive failure patterns, providing actionable insights into current model bottlenecks. We believe that CSO-Bench will effectively advance research on LLM-based code summarization, thereby supporting its application in real-world environments.
| (Internetware_Empirical.pdf) | 2.12MiB |
Sun 19 JulDisplayed time zone: Brisbane change
14:00 - 15:00 | Session 10: Empirical Software EngineeringResearch Track at Promenade Chair(s): Moustapha Awwalou DIOUF SnT, University of Luxembourg | ||
14:00 15mTalk | What Static Features Reveal About Ransomware: An Empirical Study Research Track | ||
14:15 15mTalk | When LLMs Invent Rust Crates: An Empirical Study of Hallucination Patterns and Mitigation Research Track Jieming Zheng Southern University of Science and Technology, Hao Guan Nankai University, Yepang Liu Southern University of Science and Technology Pre-print | ||
14:30 15mTalk | Studying the Effectiveness of Social Media on Open Source Donation Platforms: An Empirical Study Research Track Shuoxiao Zhang State Key Laboratory for Novel Software Technology Nanjing University, Nanjing, China, Enyi Tang Nanjing University, Xinyu Gao Nanjing University, Haoliang Cheng Nanjing University, An Guo The Hong Kong Polytechnic Universituy, Xu Zhou Nanjing University, Lingyun Situ Nanjing University, Xin Chen Nanjing University, Jianhua Zhao Nanjing University, China, Linzhang Wang Nanjing University, Xuandong Li Nanjing University | ||
14:45 15mTalk | Can LLMs Capture Developers' Practical Needs for Code Summarization? An Empirical Study of Code Summary Optimization Behaviors on GitHub Research Track xianwei wu Nanjing University, Haifeng Shen Southern Cross University, Guoping Rong Nanjing University File Attached | ||