FSE 2026
Sun 5 - Thu 9 July 2026 Montreal, Canada
Thu 9 Jul 2026 14:20 - 14:40 at MB 2.210 - Code review 2 Chair(s): Binhang Qi

While large language models (LLMs) have shown efficacy in automated code review, existing benchmarks suffer from three limitations: (1) a lack of rich semantic context like issue descriptions; (2) data quality issues due to insufficient validation; and (3) coarse granularity (file or commit level) that misses fine-grained, line-level nuances. To address these, we present ContextCRBench, a high-quality, context-rich benchmark designed for fine-grained evaluation of LLMs in code review tasks.

Our construction pipeline consists of three main modules. First, the Raw Data Crawling module collects over 153.7k issues and PRs from the selected most popular repositories. Next, the Comprehensive Context Extraction module establishes rich context by rigorously linking issue-PR pairs for textual context and extracting the full surrounding function or class for code context. Finally, our Multi-stage Data Filtering module applies a series of checks to remove entries that are outdated, improperly formatted, or identified as low-value by an LLM-based classifier. This rigorous process yields the final benchmark of 67,910 entries. Each entry in our benchmark is enriched with both textual context and code context.

Our benchmark supports three evaluation scenarios: hunk-level quality assessment, line-level defect localization, and review comment generation. Evaluations on eight popular LLMs reveal great limitations in current models, and textual context improves performance more than code context alone. Furthermore, when deployed at ByteDance as a reward signal for a self-evolving tool, ContextCRBench guided a 61.98% performance improvement, demonstrating its practical industrial utility. All scripts and datasets are available at https://github.com/kinesiatricssxilm14/ContextCRBench.

Thu 9 Jul

Displayed time zone: Eastern Time (US & Canada) change

14:00 - 15:30
Code review 2Industry Papers / Journal-First Paper / Research Papers / Tool Demonstrations at MB 2.210
Chair(s): Binhang Qi National University of Singapore
14:00
20m
Talk
The Price of Precision: The Cost of Preprocessing for Automated Code Revision in Code Review
Journal-First Paper
Shirin Pirouzkhah University of Zurich, Pooja Rani University of Zurich, Francesco Sovrano USI Lugano, Switzerland, Vincent Hellendoorn Google DeepMind, USA, Alberto Bacchelli IfI, University of Zurich
14:20
20m
Talk
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
Industry Papers
Ruida Hu Harbin Institute of Technology, Shenzhen, Xinchen Wang Harbin Institute of Technology, Xin-Cheng Wen Harbin Institute of Technology, Zhao Zhang Bytedance Network Technology, Bo Jiang Bytedance Network Technology, Pengfei Gao ByteDance, Chao Peng Tencent, Cuiyun Gao Harbin Institute of Technology, Shenzhen
14:40
20m
Talk
Mitigating the Risk of Defects and Improving Knowledge Distribution with Code Reviewer Recommenders
Research Papers
Mohammadali Sefidi Esfahani Concordia University, Peter Rigby Concordia University; Meta
Pre-print
15:00
20m
Talk
The Interaction of Complexity and Provenance in Code Review Decisions: Evidence from a Controlled Experiment
Research Papers
Neha Singh University of Zurich, Francesco Sovrano USI Lugano, Switzerland, Vincent Hellendoorn Google DeepMind, USA, Alberto Bacchelli IfI, University of Zurich
DOI Pre-print
15:20
10m
Talk
SmartPatchLinker: An Open-Source Tool to Linked Changes Detection for Code Review
Tool Demonstrations
Islem Khemissi Concordia University, Moataz Chouchen Concordia University, Dong Wang Tianjin University, Raula Gaikovina Kula The University of Osaka