Hydra-Reviewer: A holistic multi-agent system for automatic code review comment generation
Review comment generation is a crucial task in code review, and significant progress has been made in automating. Previous research has generated review comments by fine-tuning pre-trained models or Large Language Models (LLMs). However, these studies have overlooked the necessity of conducting code reviews from multiple perspectives, resulting in the omission of potential issues in code changes. Additionally, the complexity of review comments often hinders the accurate quantitative evaluation of automated tools’ effectiveness.
In this paper, we first conduct an empirical study to propose a comprehensive taxonomy of code review dimensions. We also identify three major limitations of existing automated code review (ACR) methods: lack of comprehensiveness, incorrectness, and vagueness. Building on the insights from our empirical study, we introduce HYDRA-REVIEWER, a collaborative multi-agent framework powered by large language models, designed to automatically generate high-quality code reviews. We utilize the CodeReview and CodeReviewNew benchmark datasets, along with a newly constructed review comment generation dataset. We compare HYDRA-REVIEWER with several baselines, including CodeReviewer, LLaMA-Reviewer, ChatGPT, Comprehensive-ChatGPT, and DeepSeek-V3.
The experimental results show that HYDRA-REVIEWER achieves a BLEU score of 8.20, outperforming the state-of-the-art baseline, DeepSeek-V3, which scores 7.85. In qualitative evaluation, HYDRA-REVIEWER‘s generated comments span an average of 7.8 review dimensions, addressing the limitations of existing ACR methods effectively. Additionally, HYDRA-REVIEWER demonstrates strong generalization capabilities on an unseen dataset. We further validate the contributions of each component of HYDRA-REVIEWER through an ablation study and confirm the helpfulness and readability of the generated comments via a User Study. Finally, a cost analysis reveals that HYDRA-REVIEWER generates review comments at an average cost of 0.018 dollars and 62.63 seconds per code change.
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
10:30 - 12:30 | Code Review 1Industry Papers / Journal-First Paper / Research Papers at MB 2.210 Chair(s): Tao Xiao Kyushu University | ||
10:30 20mTalk | SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation Research Papers Zhengran Zeng Peking University, Ruikai Shi Peking University, Keke Han Peking University, Yixin Li Peking University, Kaicheng Sun Northwestern Polytechnical University, Yidong Wang Peking University, Zhuohao Yu Peking University, Rui Xie Peking University, Wei Ye Peking University, Shikun Zhang Peking University | ||
10:50 20mTalk | Assessing Harmful Comments and Specificity in Code Review Feedback at Scale using Large Language Models Industry Papers Audrey You University of Auckland, Jingyi (Jenny) Wang University of Auckland, Youxiang Lei Multitudes, Lauren Peate Multitudes, Kelly Blincoe University of Auckland | ||
11:10 20mTalk | HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation Industry Papers Kla Tantithamthavorn Monash University, Hong Yi Lin The University of Melbourne, Patanamon Thongtanunam University of Melbourne, Wachiraphan (Ping) Charoenwet University of Melbourne, Minwoo Jeong Atlassian, Ming Wu Atlassian | ||
11:30 20mTalk | Hydra-Reviewer: A holistic multi-agent system for automatic code review comment generation Journal-First Paper Xiaoxue Ren Zhejiang University, Chaoqun Dai Zhejiang Gongshang University, Qiao Huang Zhejiang Gongshang University, Ye Wang Zhejiang Gongshang University, Chao Liu Chongqing University, Bo Jiang Zhejiang Gongshang University Link to publication | ||
11:50 20mTalk | AI-Assisted Fixes to Code Review Comments at Scale Industry Papers Chandra Sekhar Maddila Meta Platforms, Inc., Negar Ghorbani Meta Platforms Inc., James Saindon Meta, Parth Thakkar Meta Platforms, Inc., Vijayaraghavan Murali Meta Platforms Inc., Rui Abreu Meta, Jingyue Shen Meta Platforms Inc., Brian Zhou Meta Platforms Inc., Nachiappan Nagappan Meta Platforms, Inc., Peter C Rigby Meta / Concordia University | ||
12:10 20mTalk | Code Reviewer Recommendation for High Risk Diffs at Scale: Workflow, Recommender, and Live Experiments Industry Papers Aishwarya Girish Paraspatki Meta Platforms, Inc., Brandon Reznicek Meta Platforms, Inc., Rui Abreu Meta, Ford Garberson Meta Platforms, Inc., Audris Mockus University of Tennessee, Nachiappan Nagappan Meta Platforms, Inc., Peter C Rigby Meta / Concordia University | ||