From Static to Dynamic: Benchmarking Real-world Code Review with MCR-bench
In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve software quality, making the process costly and time-consuming. Although recent work explores large language models (LLMs) for automated code review, most approaches oversimplify code review into a single-round, static decision task, which fails to capture the multi-round interactive nature and the complex problem-solving processes inherent in realistic review scenarios. To bridge this gap, we introduce MCR-bench, the first defect state–aware benchmark designed for realistic multi-round code review. MCR-bench covers five commonly-used programming languages and consists of 2,269 real-world multi-round code review tasks. each of which annotated with fine-grained defect information and cross-round state labels. Each task in MCR-bench is equipped with fine-grained defect metadata (e.g., description, type, severity) alongside dynamic state annotations, capturing the complete evolutionary trajectory of a defect throughout the multi-round process. We obtain several findings through extensive experiments on MCR-bench with mainstream LLMs. \textbf{(1) Limited overall capability:} experiments reveal that mainstreaming LLMs exhibit limited overall performance in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases; \textbf{(2) Defect-sensitive performance:} LLMs’ performance varies substantially across different defect types and severity levels, with semantically complex or low-salience defects being significantly more likely to be missed; \textbf{(3) Underlying Failure Mechanisms:} our in-depth error analysis dissects the distinct drivers of false positives and false negatives, revealing critical weaknesses such as cross-round temporal misalignment and the inadequate long-range memory.