FSE 2026
Sun 5 - Thu 9 July 2026 Montreal, Canada
Thu 9 Jul 2026 14:10 - 14:30 at MB 3.430 - MSR 2 Chair(s): Shurui Zhou

Recent studies have explored the performance of Large Language Models (LLMs) on various Software Engineering (SE) tasks, such as code generation and bug fixing. However, these approaches typically rely on the context data from the current snapshot of the project, overlooking the potential of rich historical data residing in real-world software repositories. Additionally, the impact of prompt styles on LLM performance for SE tasks within a historical context remains underexplored. To address these gaps, we propose \textbf{HAFix}, which stands for \underline{H}istory-\underline{A}ugmented LLMs on Bug \underline{Fix}ing, a novel approach that leverages seven individual historical heuristics associated with bugs and aggregates the results of these heuristics (HAFix-Agg) to enhance LLMs’ bug-fixing capabilities. To empirically evaluate HAFix, we employ three Code LLMs (i.e., Code Llama, DeepSeek-Coder and DeepSeek-Coder-V2-Lite models) on 51 single-line Python bugs from BugsInPy and 116 single-line Java bugs from Defects4J. Our evaluation demonstrates that multiple HAFix heuristics (e.g., FN-modified and FN-all on Defects4J) achieve statistically significant improvements with large effect sizes compared to a non-historical baseline inspired by GitHub Copilot. Furthermore, the aggregated HAFix variant HAFix-Agg achieves substantial improvements with large effect sizes by combining the complementary strengths of individual heuristics, increasing bug-fixing rates relatively by an average of 45.05% on BugsInPy and 49.92% on Defects4J relative to the corresponding baseline. Moreover, within the context of historical heuristics, we identify the Instruction prompt style as the most effective template compared to the InstructionLabel and InstructionMask for LLMs in bug fixing. Finally, we evaluate the cost of HAFix in terms of inference time and token usage, and provide a pragmatic trade-off analysis of the cost and bug-fixing performance, offering valuable insights for the practical deployment of our approach in real-world scenarios.

Thu 9 Jul

Displayed time zone: Eastern Time (US & Canada) change

14:00 - 15:30
14:00
10m
Talk
Mapping GitHub Sponsorships: A Longitudinal Observatory for Open-Source Sustainability
Tool Demonstrations
Rylan Hiltz Trent University, Taher A. Ghaleb Trent University
14:10
20m
Talk
HAFix: History-Augmented Large Language Models for Bug Fixing
Journal-First Paper
Yu Shi Queen's University, Abdul Ali Bangash Lahore University of Management Sciences, Emad Fallahzadeh Queen's University, Bram Adams Queen's University, Ahmed E. Hassan Queen’s University
Pre-print
14:30
20m
Talk
Characterizing and Mitigating False-Positive Bug Reports in the Linux Kernel
Research Papers
jiashuo tian Tianjin University, Dong Wang Tianjin University, Chen Yang Tianjin University, Haichi Wang Tianjin University, Zan Wang Tianjin University, Junjie Chen Tianjin University
Pre-print
14:50
10m
Talk
Causal Software Engineering: A Vision and Roadmap
Ideas, Visions and Reflections
Roberto Pietrantuono Università di Napoli Federico II, Luca Giamattei Università di Napoli Federico II, Stefano Russo Università di Napoli Federico II, Julien Siebert Fraunhofer IESE, Neil Walkinshaw The University of Sheffield
15:00
10m
Talk
A Tool for Automatically Cataloguing and Selecting Pre-Trained Models and Datasets for Software Engineering
Tool Demonstrations
Alexandra González Universitat Politècnica de Barcelona - BarcelonaTech (UPC), Oscar Cerezo Universitat Politècnica de Catalunya - BarcelonaTech (UPC), Xavier Franch Universitat Politècnica de Catalunya, Silverio Martínez-Fernández UPC-BarcelonaTech
Pre-print Media Attached