Code Roulette: How Prompt Variability Affects LLM Code Generation
Code generation is one of the most active areas of application of Large Language Models (LLMs). While LLMs lower barriers to writing code and accelerate development process, the overall quality of generated programs depends on the quality of given prompts. Specifically, functionality and quality of generated code can be sensitive to user’s background and familiarity with software development. It is therefore important to quantify LLM’s sensitivity to variations in the input. To this end we propose an evaluation pipeline for LLM code generation with a focus on measuring sensitivity to prompt augmentations, completely agnostic to a specific programming tasks and LLMs, and thus widely applicable. We provide extensive experimental evidence illustrating utility of our method and share our code for the benefit of the community.
Tue 14 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
11:35 - 12:25 | Developer Experience and Human-AI ProgrammingLLM4Code at Oceania I Chair(s): Yiling Lou University of Illinois at Urbana-Champaign | ||
11:35 10mTalk | Achieving Productivity Gains with AI-based IDE features: A Journey at Google LLM4Code Maxim Tabachnyk Google, Inc., Xu Shu Google, Inc., Alexander Frömmgen Google, Inc., Pavel Sychev Google, Inc., Vahid Meimand Google, Inc., Ilia Krets Google, Inc., Stanislav Pyatykh Google, Inc., Abner Araujo Google, Inc., Kristof Molnar Google, Inc., Satish Chandra Meta Platforms, Inc. | ||
11:45 10mTalk | Usage, Effects and Requirements for AI Coding Assistants in the Enterprise: An Empirical Study LLM4Code Michele Merler IBM Research, Rangeet Pan IBM Research, Rahul Krishna IBM Research, Tin Kam Ho IBM Research, Raju Pavuluri IBM T.J. Watson Research Center, Maja Vukovic IBM Research | ||
11:55 10mTalk | Code Roulette: How Prompt Variability Affects LLM Code Generation LLM4Code Andrei Paleyes Department of Computer Science and Technology, Univesity of Cambridge, Diana Robinson University of Cambridge, UK, Radzim Sendyka University of Cambridge, Christian Cabrera University of Cambridge, Neil D. Lawrence Department of Computer Science and Technology, Univesity of Cambridge | ||
12:05 5mTalk | English or Chinese? Investigating the Impact of Prompt Language on Large Language Models for Code Summarization LLM4Code Yijia Tang College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics, Nanjing, China, Zhiqiu Huang Nanjing University of Aeronautics and Astronautics, Jian Xie Informationization Department (Information Technology Center), Nanjing University of Aeronautics and Astronautics, Nanjing, China, Yaoshen Yu Informationization Department (Information Technology Center), Nanjing University of Aeronautics and Astronautics, Nanjing, China, Bowei Xia College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics, Nanjing, China, Enya Shen School of Software, Tsinghua University, Beijing, China, Yukun Cao School of Computer Science and Artificial Intelligence, Wuhan Textile University | ||
12:10 10mTalk | The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution LLM4Code Norbert Tihanyi Technology Innovation Institute, Bilel Cherif Technology Innovation Institute, ABU Dhabi, UAE, Mohamed Amine Ferrag United Arab Emirates University, ABU Dhabi, UAE, Richard A. Dubniczky Eötvös Loránd University, Budapest, Hungary, Tamas Bisztray University of Oslo | ||
12:20 10mTalk | An Initial Exploration of Contrastive Prompt Tuning to Generate Energy-Efficient Code LLM4Code | ||