ICSE 2026
Sun 12 - Sat 18 April 2026 Rio de Janeiro, Brazil
Fri 17 Apr 2026 15:00 - 15:15 at Asia IV - AI for Software Engineering 24 Chair(s): Matteo Esposito

As Large Language Models for Code (LM4Code) become integral to software engineering, establishing trust in their output becomes critical. However, standard accuracy metrics obscure the underlying reasoning of generative models, offering little insight into how decisions are made. Although post-hoc interpretability methods attempt to fill this gap, they often restrict explanations to local, token-level insights, which fail to provide a developer-understandable global analysis. Our work highlights the urgent need for \textbf{global, code-based} explanations that reveal how models reason across code. To support this vision, we introduce \textit{code rationales} (CodeQ), a framework that enables global interpretability by mapping token-level rationales to high-level programming categories. Aggregating thousands of these token-level explanations allows us to perform statistical analyses that expose systemic reasoning behaviors. We validate this aggregation by showing it distills a clear signal from noisy token data, reducing explanation uncertainty (Shannon entropy) by over 50%. Additionally, we find that a code generation model (codeparrot-small) consistently favors shallow syntactic cues (e.g., \textbf{indentation}) over deeper semantic logic. Furthermore, in a user study with 37 participants, we find its reasoning is significantly misaligned with that of human developers. These findings, hidden from traditional metrics, demonstrate the importance of global interpretability techniques to foster trust in LM4Code.

Fri 17 Apr

Displayed time zone: Brasilia, Distrito Federal, Brazil change

14:00 - 15:30
AI for Software Engineering 24Journal-first Papers / New Ideas and Emerging Results (NIER) / Research Track at Asia IV
Chair(s): Matteo Esposito University of Oulu
14:00
15m
Talk
Exploring Fine-Grained Bug Report Categorization with Large Language Models and Prompt Engineering: An Empirical Study
Journal-first Papers
Anil Koyuncu Bilkent University
14:15
15m
Talk
Mapping the Trust Terrain: LLMs in Software Engineering — Insights and Perspectives
Journal-first Papers
Dipin Khati William & Mary, Yijin Liu College of William and Wary, David Nader Palacio Microsoft, Yixuan Zhang William & Mary, Denys Poshyvanyk William & Mary
14:30
15m
Talk
On-the-Fly Input Adaptation for Reliable Code IntelligenceVirtual Attendance
New Ideas and Emerging Results (NIER)
Ravishka Rathnasuriya The University of Texas - Dallas, Wei Yang UT Dallas
14:45
15m
Talk
From RSE to AI4RSE: A Quadrant Model for AI-Augmented Research Software
New Ideas and Emerging Results (NIER)
Siamak Farshidi Wageningen University & Research, Kwabena Ebo Bennin Wageningen University & Research, Bedir Tekinerdogan Wageningen University & Research
15:00
15m
Talk
Enabling Global, Human-Centered Explanations for LLMs: From Tokens to Interpretable Code and Test Generation
Research Track
Dipin Khati William & Mary, Daniel Rodriguez-Cardenas William & Mary, David Nader Palacio Microsoft, Alejandro Velasco William & Mary, Michele Tufano Google, Denys Poshyvanyk William & Mary
Pre-print
15:15
15m
Talk
When to Answer and When to Defer: A Decision Framework for Reliable Code PredictionsVirtual Attendance
New Ideas and Emerging Results (NIER)
Ravishka Rathnasuriya The University of Texas - Dallas, Wei Yang UT Dallas