RlDecompiler: Enhancing LLM-based Decompilation via Reinforcement Learning with a Multi-Faceted Reward Function
Decompiling binary code into human-readable, high-level source code is a core challenge in reverse engineering. While traditional methods often rely on brittle, pattern-based heuristics, the advent of Large Language Models (LLMs) offers a more flexible and robust approach. However, current LLM-based decompilation efforts are often limited by their training methodologies, which typically treat the task as a simple sequence-to-sequence translation and struggle to enforce the functional correctness of the output. To address these issues, this paper proposes an innovative framework for training LLMs to perform high-fidelity decompilation. A core contribution of our work is a novel data processing pipeline that enriches the model’s input. This pipeline integrates Ghidra-based static analysis to directly embed crucial context, such as static resources (strings, floating-point numbers) and relabeled basic blocks—from the binary into an LLM-friendly prompt. Building on this enriched input, we employ reinforcement learning fine-tuning guided by a multi-faceted reward function that comprehensively evaluates syntactic correctness, AST similarity, compilability, and functional correctness via test cases. Using this framework, we trained the RlDecompiler family of models (1.3B and 3B). Experimental results demonstrate that RlDecompiler achieves state-of-the-art performance, and its generated code quality is also higher than that of the baseline models. The RlDecompiler 1.3B and 3B models achieve rerunnable rates of 27.96% and 40.70%, respectively, outperforming existing baselines. The code is available at https://anonymous.4open.science/r/rldecompile-19D0/.
Sun 12 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
11:00 - 12:30 | Session 1 - Code AnalysisResearch Track / ICPC Program / Early Research Achievements (ERA) at Europa II Chair(s): Igor Wiese Federal University of Technology | ||
11:00 10mTalk | Pretraining on Call Graphs: When Binary Analysis Tasks Profit From Context Research Track Pre-print Media Attached | ||
11:10 10mTalk | LuaReSym: Recovering Variables Liveness Range in Stripped Lua Bytecode via Multi-Stage Static Analysis Research Track Weilong Li School of Computer Science and Engineering,Sun Yat-sen University, Ruizhi Xiao School of Computer Science and Engineering,Sun Yat-sen University, Yabo Wang School of Computer Science and Engineering,Sun Yat-sen University, Jiakun Sun School of Computer Science and Engineering,Sun Yat-sen University, Yuqing Shao School of Information Science and Engineering, East China University of Science and Technology, Shuyuan Jin School of Computer Science and Engineering,Sun Yat-sen University | ||
11:20 10mTalk | Modubin: A Binary Modularization Approach Based on the Locality of Homologous Functions Research Track Wenyan Yu Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Lei Cui Zhongguancun Laboratory, Jiayuan Li Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, liyubo Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Hong Li Institute of Information Engineering at Chinese Academy of Sciences, Kai Cheng Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Hongsong Zhu Institute of Information Engineering at Chinese Academy of Sciences; University of Chinese Academy of Sciences DOI Media Attached | ||
11:30 10mTalk | RlDecompiler: Enhancing LLM-based Decompilation via Reinforcement Learning with a Multi-Faceted Reward Function Research Track Yuchi Su University of Electronic Science and Technology of China, Weina Niu University of Electronic Science and Technology of China, Jiacheng Gong University of Electronic Science and Technology of China, Ran Yan University of Electronic Science and Technology of China, Song Li The State Key Laboratory of Blockchain and Data Security, Zhejiang University, Xin Liu Lanzhou University, Xiaosong Zhang University of Electronic Science and Technology of China | ||
11:40 10mTalk | A Multi-Agent Framework for Automated Exploit Generation with Constraint-Guided Comprehension and Reflection Research Track Siyi Chen Alibaba Group, Tianhan Luo Alibaba Group, Shijian Wu Alibaba Group, Xiangyu Liu Alibaba Group, Yilin Zhou Wuhan University, Qi Li Alibaba Group, Wenyuan Xu Aarhus University Pre-print | ||
11:50 10mTalk | Typify: A Lightweight Usage-driven Static Analyzer for Precise Python Type Inference Research Track Ali Aman University of Windsor, Muhammad Asaduzzaman University of Windsor, Shaowei Wang University of Manitoba Pre-print | ||
12:00 10mTalk | To GOTO or Not to GOTO: Measuring Structural Complexity of (Decompiled) Code Research Track | ||
12:10 5mTalk | Understanding Type Hints in Python Libraries and Frameworks: Early Insights Early Research Achievements (ERA) | ||
12:15 10mLive Q&A | Joint QA and Discussion ICPC Program | ||