FSE 2026
Sun 5 - Thu 9 July 2026 Montreal, Canada
Wed 8 Jul 2026 14:00 - 14:20 at MB 3.270 - Code and LLM 1 Chair(s): Zhijie Wang

Code assistants powered by large language models are now embedded in integrated development environments, yet developers lack reliable signals for when to trust generated code. Model confidence could serve as such a signal, but only if it accurately reflects the likelihood of acceptance. Post-hoc calibration aims to achieve this alignment, though its efficacy in production settings remains understudied. We investigate in-IDE confidence calibration from two perspectives: (1) scalable methods for calibrating confidence signals and (2) interface design for communicating reliability to developers. We introduce a \textbf{flexible calibration framework} for open-source models and evaluate calibration against developer acceptance behavior using over \textbf{24 million real-world IDE interactions} across multiple languages. We find that a general Platt-scaling calibrator \textit{does not}, consistently improve the usefulness of confidence as a reliability signal, while personalized calibration can help when sufficient user interaction data is available. Complementing this, a multi-phase design study with expert designers and \textbf{153 professional developers} indicates a preference for non-numerical, color-coded reliability indicators embedded in the in-editor generation workflow.

Wed 8 Jul

Displayed time zone: Eastern Time (US & Canada) change

14:00 - 15:30
14:00
20m
Talk
Does In-IDE Calibration of Large Language Models work at Scale?
Industry Papers
Roham Koohestani JetBrains Research & Delft University of Technology, Agnia Sergeyuk JetBrains Research, David Gros University of California, Davis, Claudio Spiess University of California, Davis, Sergey Titov JetBrains Research, Prem Devanbu University of California at Davis, Mali Izadi Google & TU Delft
14:20
10m
Talk
Projectional Decoding: Towards Semantic-Aware LLM Generation
Ideas, Visions and Reflections
Boqi Chen University of Ottawa, José Antonio Hernández López Department of Computer Science and Systems, University of Murcia, Aren Babikian University of Toronto
14:30
10m
Talk
The Stylistic Blind Spot: Uncovering the Hidden Implicit Bias of Coding Style on LLM Code Evaluation
Ideas, Visions and Reflections
Zhiyuan Liu Nanjing University, Yingying Jiang Nanjing University, Huiyan Wang Nanjing University
DOI
14:40
10m
Talk
TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models
Tool Demonstrations
Amirreza Esmaeili University of British Columbia, Fatemeh Hendijani Fard University of British Columbia, Okanagan
14:50
20m
Talk
Bash-Commenter: Leveraging Syntax-Aware Preference Optimization to Reinforce Large Language Model for Bash Code Comment Generation
Research Papers
Lei Yu Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, China, Jingyuan Zhang Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, China, Xin Wang Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, Li Yang Institute of Software, Chinese Academy of Sciences, Fengjun Zhang Institute of Software, Chinese Academy of Sciences, China, Peng Wang Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, Jia Xu Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, Jiajia Ma Institute of Software, Chinese Academy of Sciences, China