Evaluating LLM-generated content at scale is a critical challenge in production systems. Structured checklists enable both humans and LLMs to provide detailed, interpretable, and reliable evaluations—but crafting effective checklists remains difficult and often requires iterative refinement based on real-world data. We present AutoChecklist, a system deployed at Microsoft that augments checklist-based LLM judges with a free-form comment option and dynamically updates the checklist based on this feedback. This ensures checklists adapt to outlier instances and capture dataset-specific criteria that initial designs may miss. We evaluate AutoChecklist across three production-relevant domains and demonstrate that (1) it consistently surfaces valuable new evaluation criteria, (2) the comment option improves answer quality on the original checklist, (3) it is flexible with respect to seed checklist properties like tone and answer type, and (4) it is robust across different datasets. The results show that automated checklist refinement significantly reduces the manual effort required to build reliable evaluation pipelines.
Thu 9 JulDisplayed time zone: Eastern Time (US & Canada) change
14:00 - 15:30 | LLM for SE 7Industry Papers / Research Papers / Journal-First Paper at MB 1.210 Chair(s): Zhijie Wang Concordia University | ||
14:00 20mTalk | Not All RAGs Are Created Equal: A Component-Wise Empirical Study for Software Engineering Tasks Research Papers Qiang Ke Huazhong University of Science and Technology, Yanjie Zhao Huazhong University of Science and Technology, Hongjin Leng Xiamen University Malaysia, Shengming Zhao Fudan University, Haoyu Wang Huazhong University of Science and Technology Pre-print | ||
14:20 20mTalk | ScanCoder: Leveraging Human Attention Patterns to Enhance LLMs for Code Research Papers Yueke Zhang Vanderbilt University, Yifan Zhang Vanderbilt University, Zihan Fang Vanderbilt University, Greg Trafton Naval Research Laboratory, Daniel Levin Vanderbilt University, Kevin Leach Vanderbilt University, Yu Huang Vanderbilt University | ||
14:40 20mTalk | CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences Journal-First Paper Martin Weyssow DIRO, Université de Montréal, Aton Kamanda DIRO, Université de Montréal, Xin Zhou Singapore Management University, Singapore, Houari Sahraoui DIRO, Université de Montréal | ||
15:00 20mTalk | Empirical Studies of Parameter Efficient Methods for Large Language Models of Code and Knowledge Transfer to R Journal-First Paper Amirreza Esmaeili University of British Columbia, Iman Saberi University of British Columbia Okanagan, Fatemeh Hendijani Fard University of British Columbia, Okanagan | ||
15:20 10mTalk | AutoChecklist: Automated Checklist Refinement for LLM Judges Industry Papers Mansi Uniyal Microsoft, Mukul Singh Microsoft, Gust Verbruggen Microsoft, Vu Le Microsoft, Sumit Gulwani Microsoft | ||