The Stylistic Blind Spot: Uncovering the Hidden Implicit Bias of Coding Style on LLM Code Evaluation
Rapid iteration of large language models (LLMs) improved their performance on software engineering tasks. While prevailing benchmarks effectively measure LLMs’ functional correctness, they often overlook non-functional aspects, such as coding styles, which are crucial in real-world development and can introduce evaluation bias. To investigate such implicit bias, we propose BenchPrism that automatically disperses diverse coding styles and generates style-fulfilled benchmark variants for LLM evaluation, and thus explore how the bias behaves and is influenced across different benchmarks, LLMs, and code tasks. Experiments confirm the existence of bias, which indeed leads to unstable performance deviations in LLM evaluation (from a 14.92% drop to a 7.51% increase), even distorting model rankings. We also raise valuable takeaways and hope to inspire future researchers to be mindful of the robustness of LLM evaluation when considering the possibility of such hidden bias.
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
14:00 - 15:30 | Code and LLM 1Research Papers / Ideas, Visions and Reflections / Tool Demonstrations / Industry Papers at MB 3.270 Chair(s): Zhijie Wang Concordia University | ||
14:00 20mTalk | Does In-IDE Calibration of Large Language Models work at Scale? Industry Papers Roham Koohestani JetBrains Research & Delft University of Technology, Agnia Sergeyuk JetBrains Research, David Gros University of California, Davis, Claudio Spiess University of California, Davis, Sergey Titov JetBrains Research, Prem Devanbu University of California at Davis, Mali Izadi Google & TU Delft | ||
14:20 10mTalk | Projectional Decoding: Towards Semantic-Aware LLM Generation Ideas, Visions and Reflections Boqi Chen University of Ottawa, José Antonio Hernández López Department of Computer Science and Systems, University of Murcia, Aren Babikian University of Toronto | ||
14:30 10mTalk | The Stylistic Blind Spot: Uncovering the Hidden Implicit Bias of Coding Style on LLM Code Evaluation Ideas, Visions and Reflections DOI | ||
14:40 10mTalk | TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models Tool Demonstrations Amirreza Esmaeili University of British Columbia, Fatemeh Hendijani Fard University of British Columbia, Okanagan | ||
14:50 20mTalk | Bash-Commenter: Leveraging Syntax-Aware Preference Optimization to Reinforce Large Language Model for Bash Code Comment Generation Research Papers Lei Yu Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, China, Jingyuan Zhang Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, China, Xin Wang Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, Li Yang Institute of Software, Chinese Academy of Sciences, Fengjun Zhang Institute of Software, Chinese Academy of Sciences, China, Peng Wang Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, Jia Xu Institute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences, Jiajia Ma Institute of Software, Chinese Academy of Sciences, China | ||