Code Reviewer Recommendation for High Risk Diffs at Scale: Workflow, Recommender, and Live Experiments
Code review is an important gate for code changes, but review effort and expertise are limited. Defect prediction and just-in-time risk modeling aim to identify changes that are most likely to introduce defects so that teams can focus additional attention where it is most useful. We study whether risk-aware interventions, targeted specifically at these high-risk changes, can improve review outcomes while preserving developer velocity. In our setting, risk refers to the likelihood that a code change will introduce a defect after landing. We quantify this defect risk with the Diff Risk Score (DRS) used to identify high-risk diffs and investigate three research questions related to code review: (RQ1) whether a dedicated workflow improves review behavior on high-risk diffs, (RQ2) whether we can design a reviewer recommender tailored to risky diffs, and (RQ3) how well that new recommender performs in an A/B test across thousands of high risk diffs.
For RQ1, we introduce the Risky Diff Reviewer (RDR) workflow, which assigns an additional recommended reviewer for high-risk diffs and shows a bypassable pre-land dialog when authors attempt to land without resolving the risk signal. In a randomized trial on the top 5% of diffs by diff defect risk (over 15k diffs), RDR increases review depth and collaboration and increases the fraction of diffs whose final version falls below the high-risk threshold by 9.56%, with a minor 0.45% increase in diff reviewing time.
For RQ2, we design a new reviewer recommender (RecRDR) for risky diffs. RecRDR uses a continuous, weighted ground truth that emphasizes file experience, author collaboration, interaction quality, and recency. In offline backtests, RecRDR improves Top-1 and Top-3 action rates over the baseline recommender.
For RQ3, we evaluate RecRDR using an A/B test within the RDR workflow on over 5k diffs. RecRDR increases engagement from the recommended reviewer, with 9.94% more comments and 7.77% more acceptances, and we do not observe statistically significant regressions in guardrail metrics such as diff reviewing and processing time.
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
10:30 - 12:30 | Code Review 1Industry Papers / Journal-First Paper / Research Papers at MB 2.210 Chair(s): Tao Xiao Kyushu University | ||
10:30 20mTalk | SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation Research Papers Zhengran Zeng Peking University, Ruikai Shi Peking University, Keke Han Peking University, Yixin Li Peking University, Kaicheng Sun Northwestern Polytechnical University, Yidong Wang Peking University, Zhuohao Yu Peking University, Rui Xie Peking University, Wei Ye Peking University, Shikun Zhang Peking University | ||
10:50 20mTalk | Assessing Harmful Comments and Specificity in Code Review Feedback at Scale using Large Language Models Industry Papers Audrey You University of Auckland, Jingyi (Jenny) Wang University of Auckland, Youxiang Lei Multitudes, Lauren Peate Multitudes, Kelly Blincoe University of Auckland | ||
11:10 20mTalk | HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation Industry Papers Kla Tantithamthavorn Monash University, Hong Yi Lin The University of Melbourne, Patanamon Thongtanunam University of Melbourne, Wachiraphan (Ping) Charoenwet University of Melbourne, Minwoo Jeong Atlassian, Ming Wu Atlassian | ||
11:30 20mTalk | Hydra-Reviewer: A holistic multi-agent system for automatic code review comment generation Journal-First Paper Xiaoxue Ren Zhejiang University, Chaoqun Dai Zhejiang Gongshang University, Qiao Huang Zhejiang Gongshang University, Ye Wang Zhejiang Gongshang University, Chao Liu Chongqing University, Bo Jiang Zhejiang Gongshang University Link to publication | ||
11:50 20mTalk | AI-Assisted Fixes to Code Review Comments at Scale Industry Papers Chandra Sekhar Maddila Meta Platforms, Inc., Negar Ghorbani Meta Platforms Inc., James Saindon Meta, Parth Thakkar Meta Platforms, Inc., Vijayaraghavan Murali Meta Platforms Inc., Rui Abreu Meta, Jingyue Shen Meta Platforms Inc., Brian Zhou Meta Platforms Inc., Nachiappan Nagappan Meta Platforms, Inc., Peter C Rigby Meta / Concordia University | ||
12:10 20mTalk | Code Reviewer Recommendation for High Risk Diffs at Scale: Workflow, Recommender, and Live Experiments Industry Papers Aishwarya Girish Paraspatki Meta Platforms, Inc., Brandon Reznicek Meta Platforms, Inc., Rui Abreu Meta, Ford Garberson Meta Platforms, Inc., Audris Mockus University of Tennessee, Nachiappan Nagappan Meta Platforms, Inc., Peter C Rigby Meta / Concordia University | ||