AI-to-AI Code Review of GitHub Pull Requests
This program is tentative and subject to change.
AI coding agents are increasingly integrated into software development workflows, operating on both sides of the pull request (PR) process: AI authoring agents, which create or modify PRs, and AI reviewers, which evaluate them. This creates a closed loop where one AI coding agent reviews contributions of another AI coding agent. In this paper, we construct an AI-to-AI code review dataset by linking AI-authored pull requests with AI-attributed review events from CodAGE, a public dataset of coding agent–generated GitHub events. Our dataset contains 248,641 PRs across 12 AI coding agents with at least one AI review, including 45,269 reviewed by a different AI product and 208,145 by the same product. We observe that cross-product AI-to-AI code review occurs in only about 1.6% of identified agent-authored PRs but is substantial in absolute terms: 45k PRs written by one identifiable AI product and reviewed by another. This activity grows by more than two orders of magnitude from 2025-Q1 to 2025-Q3. We measure reviewer behavior using CodeRabbit comment categories, per-PR comment volume, and time to first review, and find that it varies across author–reviewer pairs. For example, Claude-Code PRs receive more refactor comments from CodeRabbit than Copilot PRs (35.0% vs. 10.5%), a difference that may stem from PRs themselves rather than the reviewer. For three of four dual-role reviewers, mean comments per PR were 58–65% higher in the same-product group, though effect sizes were small or negligible and the difference was concentrated in the upper tail. Median time from PR creation to first AI review was 1.2 minutes for cross-product pairs and 4.7 minutes for same-product pairs, reflecting which reviewer bots are in each group rather than the product pairing. Overall, our large-scale characterization shows that closed-loop AI-to-AI code review is on the rise but remains a minority phenomenon, with review output varying across authoring-agent groups and author–reviewer configurations.
This program is tentative and subject to change.
Fri 9 OctDisplayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change
14:00 - 15:30 | Trust, Review and Evaluation of AI-Generated CodeESEM - Journal First Track / ESEM - Registered Reports Track / ESEM - Technical Track / ESEM - Software Engineering in Practice Track / ESEM - Emerging Results, Vision, and Reflection Papers Track at Terra | ||
14:00 12mTalk | AI-to-AI Code Review of GitHub Pull Requests ESEM - Emerging Results, Vision, and Reflection Papers Track | ||
14:12 12mTalk | How Developers Use Relation Chains in Code Review: An Empirical Study Across Three Open-Source Ecosystems ESEM - Technical Track Ahmed Belhouchette ENSI, Mannouba University, Moataz Chouchen Concordia University, Marouene Chaieb National School of Computer Science, Mohammad Hamdaqa Polytechnique Montreal, Abdelwahab Hamou-Lhadj Concordia University, Montreal, Canada | ||
14:25 12mTalk | How Do Software Professionals Evaluate AI-Generated Code? (Registered Report) ESEM - Registered Reports Track Samuli Määttä University of Oulu, Hera Arif Dalhousie University, Burak Turhan University of Oulu, Paul Ralph Dalhousie University, Markus Kelanti University of Oulu Pre-print | ||
14:38 12mTalk | CWEFT: CWE-aware Evaluation of Free-text vs. Typed Prompts ESEM - Emerging Results, Vision, and Reflection Papers Track | ||
14:51 12mTalk | Trust-Calibrated Code Review: A Participatory Design Study of Review Workflows for LLM-Generated Multi-File Changes ESEM - Software Engineering in Practice Track Lo Heander Lund University, Agnia Sergeyuk JetBrains Research, Ilya Zakharov JetBrains Research, Emma Söderberg Lund University, Nikita Mukhortov JetBrains | ||
15:04 12mTalk | Code Review as Decision-Making - Building a Cognitive Model from the Questions Asked During Code Review ESEM - Journal First Track | ||
15:17 12mTalk | How Reliable Is LLM-as-Judge for Patch Correctness Assessment? An Empirical Study ESEM - Technical Track Shanggui Zhan School of Computer Science and Technology, Hangzhou Dianzi University; Zhejiang Key Laboratory of New Industrial Internet Control Technology, Xingqi Wang School of Computer Science and Technology, Hangzhou Dianzi University; Zhejiang Key Laboratory of New Industrial Internet Control Technology, Dan Wei School of Computer Science and Technology, Hangzhou Dianzi University; Zhejiang Key Laboratory of New Industrial Internet Control Technology, Xin Xiang chool of Computer Science and Technology, Hangzhou Dianzi University | ||