ICSE 2026
Sun 12 - Sat 18 April 2026 Rio de Janeiro, Brazil
Fri 17 Apr 2026 14:00 - 14:15 at Asia IV - AI for Software Engineering 24 Chair(s): Matteo Esposito

Accurate classification of issues is essential for effective project management and timely responses, as the volume of issue reports continues to grow. Manual classification is labor-intensive and error-prone, necessitating automated solutions. While large language models (LLMs) show promise in automated issue labeling, most research focuses on broad categorization (e.g., bugs, feature requests), with limited attention to fine-grained categorization. Understanding specific bug types is crucial, as different bugs require tailored resolution strategies. This study addresses this gap by evaluating LLMs and prompt engineering strategies for fine-grained bug report categorization. We analyze 221,184 fine-grained bug report category labels generated by selected LLMs using various prompt engineering strategies for 1,024 bug reports. We examine how LLMs and prompt engineering influence output characteristics, control over outputs, and categorization performance. Our findings highlight that LLMs and prompt engineering significantly impact output consistency and classification capability, with some yielding consistent results and others introducing variability. Based on these findings, we analyze the agreements and disagreements between LLM-generated labels and human annotations to assess category correctness. Our results suggest that examining label consistency and discrepancies can serve as a complementary method for validating bug report categories, identifying unclear reports, and detecting misclassifications in human annotations.

Fri 17 Apr

Displayed time zone: Brasilia, Distrito Federal, Brazil change

14:00 - 15:30
AI for Software Engineering 24Journal-first Papers / New Ideas and Emerging Results (NIER) / Research Track at Asia IV
Chair(s): Matteo Esposito University of Oulu
14:00
15m
Talk
Exploring Fine-Grained Bug Report Categorization with Large Language Models and Prompt Engineering: An Empirical Study
Journal-first Papers
Anil Koyuncu Bilkent University
14:15
15m
Talk
Mapping the Trust Terrain: LLMs in Software Engineering — Insights and Perspectives
Journal-first Papers
Dipin Khati William & Mary, Yijin Liu College of William and Wary, David Nader Palacio Microsoft, Yixuan Zhang William & Mary, Denys Poshyvanyk William & Mary
14:30
15m
Talk
On-the-Fly Input Adaptation for Reliable Code IntelligenceVirtual Attendance
New Ideas and Emerging Results (NIER)
Ravishka Rathnasuriya The University of Texas - Dallas, Wei Yang UT Dallas
14:45
15m
Talk
From RSE to AI4RSE: A Quadrant Model for AI-Augmented Research Software
New Ideas and Emerging Results (NIER)
Siamak Farshidi Wageningen University & Research, Kwabena Ebo Bennin Wageningen University & Research, Bedir Tekinerdogan Wageningen University & Research
15:00
15m
Talk
Enabling Global, Human-Centered Explanations for LLMs: From Tokens to Interpretable Code and Test Generation
Research Track
Dipin Khati William & Mary, Daniel Rodriguez-Cardenas William & Mary, David Nader Palacio Microsoft, Alejandro Velasco William & Mary, Michele Tufano Google, Denys Poshyvanyk William & Mary
Pre-print
15:15
15m
Talk
When to Answer and When to Defer: A Decision Framework for Reliable Code PredictionsVirtual Attendance
New Ideas and Emerging Results (NIER)
Ravishka Rathnasuriya The University of Texas - Dallas, Wei Yang UT Dallas