Internetware 2026
Sat 18 - Mon 20 July 2026 Gold Coast, Australia

Large Language Models (LLMs) have recently demonstrated strong performance in automatic code summarization. However, it remains unclear whether this capability truly aligns with developers’ practical needs during real-world software maintenance. Critically, the community lacks an effective benchmark and evaluation methodology to systematically assess this alignment. To address this gap, we first propose \textbf{CSO-Bench}, a novel benchmark comprising 4{,}845 Python instances (3{,}845 method-level and 1{,}000 class-level). It accurately captures real-world \textbf{C}ode \textbf{S}ummary \textbf{O}ptimization practices, where \emph{developers actively refine summaries without modifying the underlying code logic}. We characterize these refinement behaviors across four rigorous quality dimensions: factuality, completeness, clarity, and compliance. To enable systematic assessment, we innovatively introduce two complementary tasks—\emph{Summary Judgment} and \emph{Summary Editing}—to evaluate LLMs’ ability to recognize and improve summary quality. Leveraging a comprehensive framework that integrates automated metrics with a verified LLM-as-a-Judge protocol, we systematically evaluate seven mainstream LLMs (7B\textasciitilde 685B parameters) under four prompting strategies (i.e., Vanilla, CoT, Few-Shot, and RAG). Our extensive experiments reveals a significant gap between current LLMs’ capabilities and developers’ multi-dimensional quality requirements. Through an in-depth root cause analysis, we distill pervasive failure patterns, providing actionable insights into current model bottlenecks. We believe that CSO-Bench will effectively advance research on LLM-based code summarization, thereby supporting its application in real-world environments.

Sun 19 Jul

Displayed time zone: Brisbane change

14:00 - 15:00
Session 10: Empirical Software EngineeringResearch Track at Promenade
Chair(s): Moustapha Awwalou DIOUF SnT, University of Luxembourg
14:00
15m
Talk
What Static Features Reveal About Ransomware: An Empirical Study
Research Track
Xinyu Liu , Kaifeng Huang Tongji University
14:15
15m
Talk
When LLMs Invent Rust Crates: An Empirical Study of Hallucination Patterns and Mitigation
Research Track
Jieming Zheng Southern University of Science and Technology, Hao Guan Nankai University, Yepang Liu Southern University of Science and Technology
Pre-print
14:30
15m
Talk
Studying the Effectiveness of Social Media on Open Source Donation Platforms: An Empirical Study
Research Track
Shuoxiao Zhang State Key Laboratory for Novel Software Technology Nanjing University, Nanjing, China, Enyi Tang Nanjing University, Xinyu Gao Nanjing University, Haoliang Cheng Nanjing University, An Guo The Hong Kong Polytechnic Universituy, Xu Zhou Nanjing University, Lingyun Situ Nanjing University, Xin Chen Nanjing University, Jianhua Zhao Nanjing University, China, Linzhang Wang Nanjing University, Xuandong Li Nanjing University
14:45
15m
Talk
Can LLMs Capture Developers' Practical Needs for Code Summarization? An Empirical Study of Code Summary Optimization Behaviors on GitHub
Research Track
xianwei wu Nanjing University, Haifeng Shen Southern Cross University, Guoping Rong Nanjing University
File Attached