FSE 2026
Sun 5 - Thu 9 July 2026 Montreal, Canada
Thu 9 Jul 2026 15:20 - 15:30 at MB 1.210 - LLM for SE 7 Chair(s): Zhijie Wang

Evaluating LLM-generated content at scale is a critical challenge in production systems. Structured checklists enable both humans and LLMs to provide detailed, interpretable, and reliable evaluations—but crafting effective checklists remains difficult and often requires iterative refinement based on real-world data. We present AutoChecklist, a system deployed at Microsoft that augments checklist-based LLM judges with a free-form comment option and dynamically updates the checklist based on this feedback. This ensures checklists adapt to outlier instances and capture dataset-specific criteria that initial designs may miss. We evaluate AutoChecklist across three production-relevant domains and demonstrate that (1) it consistently surfaces valuable new evaluation criteria, (2) the comment option improves answer quality on the original checklist, (3) it is flexible with respect to seed checklist properties like tone and answer type, and (4) it is robust across different datasets. The results show that automated checklist refinement significantly reduces the manual effort required to build reliable evaluation pipelines.

Thu 9 Jul

Displayed time zone: Eastern Time (US & Canada) change

14:00 - 15:30
LLM for SE 7Industry Papers / Research Papers / Journal-First Paper at MB 1.210
Chair(s): Zhijie Wang Concordia University
14:00
20m
Talk
Not All RAGs Are Created Equal: A Component-Wise Empirical Study for Software Engineering Tasks
Research Papers
Qiang Ke Huazhong University of Science and Technology, Yanjie Zhao Huazhong University of Science and Technology, Hongjin Leng Xiamen University Malaysia, Shengming Zhao Fudan University, Haoyu Wang Huazhong University of Science and Technology
Pre-print
14:20
20m
Talk
ScanCoder: Leveraging Human Attention Patterns to Enhance LLMs for Code
Research Papers
Yueke Zhang Vanderbilt University, Yifan Zhang Vanderbilt University, Zihan Fang Vanderbilt University, Greg Trafton Naval Research Laboratory, Daniel Levin Vanderbilt University, Kevin Leach Vanderbilt University, Yu Huang Vanderbilt University
14:40
20m
Talk
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
Journal-First Paper
Martin Weyssow DIRO, Université de Montréal, Aton Kamanda DIRO, Université de Montréal, Xin Zhou Singapore Management University, Singapore, Houari Sahraoui DIRO, Université de Montréal
15:00
20m
Talk
Empirical Studies of Parameter Efficient Methods for Large Language Models of Code and Knowledge Transfer to R
Journal-First Paper
Amirreza Esmaeili University of British Columbia, Iman Saberi University of British Columbia Okanagan, Fatemeh Hendijani Fard University of British Columbia, Okanagan
15:20
10m
Talk
AutoChecklist: Automated Checklist Refinement for LLM Judges
Industry Papers
Mansi Uniyal Microsoft, Mukul Singh Microsoft, Gust Verbruggen Microsoft, Vu Le Microsoft, Sumit Gulwani Microsoft