Enhancing LLM Code Generation with Ensembles: A Similarity-Based Selection Approach
Ensemble learning has been widely used in machine learning to improve model robustness, accuracy, and generalization, but has not yet been applied to code generation tasks with large language models (LLMs). We propose an ensemble approach for LLMs in code generation. Instead of relying on the output of a single model, we generate multiple candidate programs from different LLMs and apply a structured voting mechanism to select the most reliable solution. For voting, we compute syntactic and semantic similarity using CodeBLEU and behavioral equivalence using CrossHair’s differential behavior analysis. By aggregating these similarity scores, we select the program that best aligns with the consensus among the candidates. We show through experiments that our ensemble approach consistently outperforms standalone LLMs on the well-known HumanEval and the more challenging LiveCodeBench datasets, achieving an accuracy of 90.2% and 50.2%, respectively, on the two datasets. In comparison, the best-performing LLM (GPT-4o) has an accuracy of 83.5% and 43.4%, respectively. Furthermore, even when restricted to free open source models, our method achieves an accuracy of 80.5% and 41.6%, respectively, demonstrating the viability of our approach in resource-constrained settings.
Fri 17 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
16:00 - 17:30 | AI for Software Engineering 26Research Track / Demonstrations / New Ideas and Emerging Results (NIER) at Asia I Chair(s): Jiakun Liu Harbin Institute of Technology | ||
16:00 15mTalk | AdapTrack: Constrained Decoding without Distorting LLM's Output Intent Research Track Yongmin Li Peking University, Jia Li Tsinghua University, Ge Li Peking University, Zhi Jin Peking University, Wuhan University | ||
16:15 15mTalk | Evaluating Generated Commit Messages with Large Language ModelsDistinguished Paper Award Research Track Qunhong Zeng Beijing Institute of Technology, Yuxia Zhang Beijing Institute of Technology, Zexiong Ma Peking University, Bo Jiang Bytedance Network Technology, Ningyuan Sun ByteDance, Klaas-Jan Stol Lero; University College Cork; SINTEF Digital , Xingyu Mou Beijing Institute of Technology, Hui Liu Beijing Institute of Technology Pre-print | ||
16:30 15mTalk | Automating Just-In-Time Python Type Annotation UpdatingDistinguished Paper Award Research Track Zhipeng Xue Zhejiang University, Zhipeng Gao Shanghai Institute for Advanced Study - Zhejiang University, Xing Hu Zhejiang University, Jingyuan Chen Zhejiang University, Xin Xia Zhejiang University, Shanping Li Zhejiang University | ||
16:45 15mTalk | Unveiling the Potential of Diffusion Large Language Models in Software Engineering Tasks: An Empirical Study New Ideas and Emerging Results (NIER) Jingyao Zhang Xi'an Jiaotong-Liverpool University, Li Tianlin NTU, Xiaoyu Zhang Nanyang Technological University, Singapore, Qiang Hu Tianjin University, Bin Shi Xi'an Jiaotong University Media Attached File Attached | ||
17:00 15mTalk | Enhancing LLM Code Generation with Ensembles: A Similarity-Based Selection Approach Research Track Tarek Mahmud Texas State University, Bin Duan University of Queensland, Corina S. Păsăreanu Carnegie Mellon University; NASA Ames, Guowei Yang University of Queensland | ||
17:15 15mTalk | Code4MeV2: a Research-oriented Code-completion Platform Demonstrations Roham Koohestani Delft University of Technology, Parham Bateni Delft University of Technology, Aydin Ebrahimi Delft University of Technology, Behdad Etezadi Delft University of Technology, Kiarash Karimi Delft University of Technology, Mali Izadi TU Delft | ||