One Size Does Fit All: Exploring Model Fusion for Software Engineering Tasks
Large language models (LLMs) have achieved remarkable performance in software engineering (SE), and fine-tuning LLM for specific SE tasks has gradually become a new paradigm. However, storing fine-tuned checkpoints for multiple tasks incurs heavy storage and deployment complexity. Model fusion, which operates on fine-tuned parameters, offers excellent parameter compression and scalability, yet its effectiveness in the SE domain remains underexplored, making such an investigation essential for guiding the development of customized fusion techniques for the SE domain. To bridge this gap, we conduct a systematic study of model fusion in the SE contexts and reveal the following major findings: (1) when fusing programming languages (PLs) within the same task, model fusion usually works well and can enhance the performance of PLs with fewer data when PLs share similar features. (2) when fusing SE tasks of the same category within a same PL, all methods except TALL-Masks generally suffer substantial performance degradation on specific tasks; (3) when fusing SE tasks of different categories across different PLs, all existing model fusion methods exhibit significant performance degradation on certain tasks. In our evaluation results, TALL-Masks, which introduces a mask for each task to extract the most relevant dimensions from the fusion parameters, achieves promising performance. However, during parameters fusion, weak features (i.e., small variation in fine-tuned parameters) are easily overshadowed by strong ones (i.e., large variation in fine-tuned parameters) during parameter fusion, causing the constructed masks to fail to extract the most relevant parameters. To overcome this situation, we propose an improved version of TALL-Masks, called Scaling-Masks. The key idea is to amplify weak features to prevent them from being overshadowed by strong ones, which is achieved by scaling the value range of weak features to match that of strong features. Experimental results demonstrate that Scaling-Masks can significantly improve fusion performance for tasks with extremely weak features without affecting other tasks, with normalized accuracy improved by 63.49% for vulnerability detection when fusing SE tasks of different categories and 24.02% for PHP when fusing PLs in the code repair task.
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
14:00 - 15:30 | SE and AI 2Industry Papers / Research Papers / Tool Demonstrations at MB 3.210 Chair(s): Rafal Wlodarski Carnegie Mellon Silicon Valley | ||
14:00 10mTalk | IssueGuard: Real-Time Secret Leak Prevention Tool for GitHub Issue Reports Tool Demonstrations Md Nafiu Rahman Brac University , Sadif Ahmed Bangladesh University of Engineering and Techonology, Zahin Wahab The University of British Columbia, Gias Uddin York University, Canada, Rifat Shahriyar Bangladesh University of Engineering and Technology Dhaka, Bangladesh Pre-print | ||
14:10 20mTalk | ProofFusion: Improving Neural Theorem Proving via Adaptive Retrieval-Augmented Reasoning Research Papers Manqing Zhang Northwestern Polytechnical University, Yunwei Dong Northwestern Polytechnical University, School of Computer Science and Engineering, Lingru Zhou Northwestern Polytechnical University, Bingxu Xiao Northwestern Polytechnical University, Yepang Liu Southern University of Science and Technology Pre-print | ||
14:30 20mTalk | Eagle: Leveraging Operations Documents for Comprehensive Benchmark Question Generation Industry Papers Yuhe Liu Tsinghua University, Changhua Pei Computer Network Information Center at Chinese Academy of Sciences, Hang Wang Computer Network Information Center, Chinese Academy of Sciences, Longlong Xu Tsinghua University, Xiaogang Dong Huawei, Zhen Feng Huawei, Li Zheng China Academy of Information and Communications Technology, Kehang Ji China Academy of Information and Communications Technology, Dan Pei Tsinghua University | ||
14:50 20mTalk | Unveiling AI-Driven Web Applications: Insights into Characteristics, Functionality, and Compliance Research Papers Liuhuo Wan , Zicong Liu University of Queensland, Chuan Yan University of Queensland, Liujia Wan Northeastern University, Naipeng Dong The University of Queensland, Australia, Zi Huang University of Queensland, Guangdong Bai City University of Hong Kong Pre-print | ||
15:10 20mTalk | One Size Does Fit All: Exploring Model Fusion for Software Engineering Tasks Research Papers Yinggang Qiu National University of Defense Technology, Yihao Qin , Mingyang Geng National University of Defense Technology, Shangwen Wang National University of Defense Technology, Dezun Dong NUDT Link to publication DOI | ||