ICSE 2026
Sun 12 - Sat 18 April 2026 Rio de Janeiro, Brazil
Fri 17 Apr 2026 14:00 - 14:15 at Oceania VII - Software Engineering for AI 7 Chair(s): Houari Sahraoui

Deep learning (DL) techniques have achieved significant success in various software engineering tasks (e.g., code completion by Copilot). However, DL systems are prone to bugs from many sources, including training data. Existing literature suggests that bugs in training data are highly prevalent, but little research has focused on understanding their impacts on the models used in software engineering tasks. In this paper, we address this research gap through a comprehensive empirical investigation focused on three types of data prevalent in software engineering tasks: code-based, text-based, and metric-based. Using state-of-the-art baselines, we compare the models trained on clean datasets with those trained on datasets with quality issues and without proper preprocessing. By analysing the gradients, weights, and biases from neural networks under training, we identify the symptoms of data quality and preprocessing issues. Our analysis reveals that quality issues in code data cause biased learning and gradient instability, whereas problems in text data lead to overfitting and poor generalisation of models. On the other hand, quality issues in metric data result in exploding gradients and model overfitting, and inadequate preprocessing exacerbates these effects across all three data types. Finally, we demonstrate the validity and generalizability of our findings using six new datasets. Our research provides a better understanding of the impact and symptoms of data bugs in software engineering datasets. Practitioners and researchers can leverage these findings to develop better monitoring systems and data-cleaning methods to help detect and resolve data bugs in deep learning systems.

Fri 17 Apr

Displayed time zone: Brasilia, Distrito Federal, Brazil change

14:00 - 15:30
Software Engineering for AI 7Research Track / Journal-first Papers at Oceania VII
Chair(s): Houari Sahraoui DIRO, Université de Montréal
14:00
15m
Talk
Towards Understanding the Impact of Data Bugs on Deep Learning Models in Software Engineering
Journal-first Papers
Mehil Shah Dalhousie University, Masud Rahman Dalhousie University, Foutse Khomh Polytechnique Montréal
Link to publication Pre-print
14:15
15m
Talk
T4PC: Training Deep Neural Networks for Property Conformance
Journal-first Papers
Felipe Toledo , Trey Woodlief University of Virginia, Sebastian Elbaum University of Virginia, Matthew B Dwyer University of Virginia
14:30
15m
Talk
A Comprehensive Study of Deep Learning Model Fixing ApproachesDistinguished Paper Award
Research Track
Hanmo You Tianjin University, Zan Wang Tianjin University, Zishuo Dong College of Intelligence and Computing, Tianjin University, Luanqi Mo College of Intelligence and Computing, Tianjin University, Jianjun Zhao Kyushu University, Junjie Chen Tianjin University
14:45
15m
Talk
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
Research Track
Mehil Shah Dalhousie University, Masud Rahman Dalhousie University, Foutse Khomh Polytechnique Montréal
DOI Pre-print
15:00
15m
Talk
TypeCare: Boosting Python Type Inference Models via Context-Aware Re-Ranking and AugmentationArtifact Award Winner
Research Track
Wonseok Oh Korea University, Hakjoo Oh Korea University
15:15
15m
Talk
Aligning Requirement for Large Language Model's Code Generation
Research Track
Zhao Tian Tianjin University, Junjie Chen Tianjin University
Pre-print