FSE 2026
Sun 5 - Thu 9 July 2026 Montreal, Canada
Thu 9 Jul 2026 11:30 - 11:50 at MB 3.270 - Code and LLM 2 Chair(s): Ali Ouni

Code summarization plays a vital role in program comprehension and software maintenance by generating natural language descriptions to summarize the semantics of code. While Large Language Models (LLMs) have shown remarkable performance in this area, recent empirical studies reveal a critical limitation: LLMs are prone to hallucinations, producing summaries that are factually inaccurate or unfaithful to the source code, potentially misleading developers. In this paper, we propose to unveil, detect, and mitigate hallucinations in LLM-based code summarization. First, we construct Hallu-Eval, a novel dataset for unveiling hallucination phenomena and rigorously evaluating the effectiveness of hallucination detection and mitigation in LLM-based code summarization. It comprises both original code snippets to capture naturally occurring hallucinations and their semantically perturbed counterparts, which are designed to systematically induce challenging logical hallucinations, all complemented with manual hallucination annotations. Next, we propose Hallu-Det, a synergistic approach that combines direct entity-level detection to identify explicit hallucinations with a synonymous mutation-based refinement to reliably confirm or refute more ambiguous cases. Finally, we introduce Hallu-Shield, an inference-time mitigation approach that leverages an external value model to guide LLMs toward producing more faithful summaries—without costly retraining of the LLM itself. Extensive experiments show that Hallu-Eval effectively triggers hallucinations, increasing the hallucination rate of models such as Qwen2.5-Coder-7B from 17% to 97% on perturbed code. Our detection approach, Hallu-Det, achieves the best performance among baselines, reaching an F1-score of 0.95 for summaries generated by Qwen2.5-Coder-7B. Moreover, our mitigation method, Hallu-Shield, reduces hallucination rates—for example, from 66% to 59%, a 10.6% relative reduction, on DeepSeek-Coder-6.7B—while simultaneously improving summary quality, achieving a 71.5% win rate in a side-by-side evaluation using an LLM as a judge.

Thu 9 Jul

Displayed time zone: Eastern Time (US & Canada) change

10:30 - 12:30
Code and LLM 2Research Papers / Industry Papers at MB 3.270
Chair(s): Ali Ouni Ecole de Technologie Superieure (ETS)
10:30
20m
Talk
NES: An Instruction-Free, Low-Latency Next Edit Suggestion Framework Powered by Learned Historical Editing Trajectories
Industry Papers
Xinfang Chen Ant Group, Siyang Xiao Ant Group, Xianying Zhu Ant Group, Junhong Xie Ant Group, Ming Liang Ant Group, Dajun Chen Ant Group, Wei Jiang Ant Group, Yong Li Ant Group, Peng Di Kunlunxin & UNSW Sydney
10:50
20m
Talk
Balancing Latency and Accuracy of Code Completion via Local-Cloud Model Cascading
Research Papers
Lu Hanzhen Zhejiang University, Lishui Fan Zhejiang University, Jiachi Chen Zhejiang University, Qiuyuan Chen Tencent Technology, Zhao Wei Tencent, Zhongxin Liu Zhejiang University
Pre-print
11:10
20m
Talk
From Specifications to Implementation in the Gen-AI Era: Lessons from a Project-based Software Engineering Course
Research Papers
Yingying Wang University of British Columbia, Masih Beigi Rizi University of British Columbia, Fatemeh Khashei University of British Columbia, Julia Rubin The University of British Columbia
Pre-print
11:30
20m
Talk
Hallucinations in LLM-based Code Summarization: Unveiling, Detection, and Mitigation
Research Papers
Guanghua Wan Huazhong University of Science and Technology, Yuanning Feng Huazhong University of Science and Technology, Yao Wan Huazhong University of Science and Technology, Zhaoyang Chu University College London (UCL), Zhangqian Bi Huazhong University of Science and Technology, Junxiao Han Hangzhou City University, Zhou Zhao Zhejiang University, Hongyu Zhang Chongqing University, Pingpeng Yuan Huazhong University of Science and Technology, Xuanhua Shi Huazhong University of Science and Technology, Hai Jin Huazhong University of Science and Technology
DOI
11:50
20m
Talk
ReDef: Do Code Language Models Truly Understand Code Changes for Just-in-Time Software Defect Prediction?
Research Papers
Doha Nam Korea Advanced Institute of Science and Technology, Taehyoun Kim Korea Advanced Institute of Science and Technology; Agency for Defense Development, Duksan Ryu Jeonbuk National University, Jongmoon Baik Korea Advanced Institute of Science and Technology
DOI Pre-print Media Attached
12:10
20m
Talk
Mitigating Prompt-Induced Cognitive Biases in General-Purpose AI for Software Engineering
Research Papers
Francesco Sovrano USI Lugano, Switzerland, Gabriele Dominici Università della Svizzera italiana (USI), Alberto Bacchelli IfI, University of Zurich
Link to publication Pre-print