Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents
The Rust programming language presents a steep learning curve and significant coding challenges, making the automation of issue resolution essential for its broader adoption. Recently, LLM-powered code agents have shown remarkable success in resolving complex software engineering tasks, yet their application to Rust has been limited by the absence of a large-scale, repository-level benchmark. To bridge this gap, we introduce Rust-SWE-bench, a benchmark comprising 500 real-world, repository-level software engineering tasks from 34 diverse and popular Rust repositories. We then perform a comprehensive study on Rust-SWE-bench with four representative agents and four state-of-the-art LLMs to establish a foundational understanding of their capabilities and limitations in the Rust ecosystem.
Our extensive study reveals that while ReAct-style agents are promising, i.e., resolving up to 21.2% of issues, they are limited by two primary challenges: comprehending repository-wide code structure and complying with Rust’s strict type and trait semantics. We also find that issue reproduction is rather critical for task resolution. Inspired by these findings, we propose RUSTFORGER, a novel agentic approach that integrates an automated test environment setup with a Rust metaprogramming-driven dynamic tracing strategy to facilitate reliable issue reproduction and dynamic analysis. The evaluation shows that RUSTFORGER using Claude-Sonnet-3.7 significantly outperforms all baselines, resolving 28.6% of tasks on Rust-SWE-bench, i.e., a 34.9% improvement over the strongest baseline, and, in aggregate, uniquely solves 46 tasks that no other agent could solve across all adopted advanced LLMs.
Wed 15 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
11:00 - 12:30 | AI for Software Engineering 2Research Track at Asia IV Chair(s): Mike Papadakis University of Luxembourg | ||
11:00 15mTalk | Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents Research Track Jiahong Xiang Southern University of Science and Technology, Wenxiao He Southern University of Science and Technology, Xihua Wang Southern University of Science and Technology, Hongliang Tian Ant Group, Yuqun Zhang Southern University of Science and Technology | ||
11:15 15mTalk | SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution Research Track Han Li Shanghai Jiao Tong University, China, Yuling Shi Shanghai Jiao Tong University, Shaoxin Lin , Xiaodong Gu Shanghai Jiao Tong University, Heng Lian Xidian University, Wang Xin , Yantao Jia Huawei, huangtao , Qianxiang Wang Huawei Technologies Co., Ltd | ||
11:30 15mTalk | More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents Research Track | ||
11:45 15mTalk | ADARULE: LLM-Driven Natural Language to LTL Conversion via Pattern-Adaptive Rule Induction Research Track Jiayi Hu East China Normal University, Jingling Sun University of Electronic Science and Technology of China, Chong Wang Nanyang Technological University, Yihao Huang East China Normal University, jincaofeng , Yilongfei Xu East China Normal University, Yong Li Institute of Software, Chinese Academy of Sciences, Kailong Wang Huazhong University of Science and Technology, Weikai Miao Shanghai Key Lab for Trustworthy Computing, School of Computer Science and Software Engineering, East China Normal University, Jin Song Dong National University of Singapore, Geguang Pu East China Normal University, China | ||
12:00 15mTalk | Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents Research Track Ratnadira Widyasari Singapore Management University, Singapore, Martin Weyssow Singapore Management University, Ivana Clairine Irsan Singapore Management University, Han Wei Ang GovTech, Frank Liauw Government Technology Agency Singapore, Eng Lieh Ouh Singapore Management University, Singapore, Lwin Khin Shar Singapore Management University, Hong Jin Kang University of Sydney, David Lo Singapore Management University | ||
12:15 15mTalk | Agent-Based Ensemble Reasoning for Repository-Level Issue Resolution Research Track Zhao Tian Tianjin University, Pengfei Gao ByteDance, Junjie Chen Tianjin University, Chao Peng ByteDance Pre-print | ||