A Preliminary Analysis of the Impact of AI Assisted/Generated Code on Quality, Centrality and Review Time at Scale
This program is tentative and subject to change.
AI code-generation tools are widely deployed, yet AI is not applied uniformly: its use varies systematically with the context of the task, the code being modified, and the engineer. Studies that do not account for these contextual differences risk attributing to AI what is actually driven by context. Using data from over 10K developers and 400K diffs at Meta, with character-level provenance tracing of AI-suggested code through editing, review, and landing, we model (a) which context factors predict AI code landing, (b) how landed AI code relates to review time, and (c) its association with production outages (Sevs)—while controlling for contextual confounds. We find that AI code lands more often in less central code, larger changes, test files, and for authors with higher tenure. Review time decreases and Sev rates are lower for diffs with AI code, but both associations are confounded by AI code appearing in less central contexts. These findings demonstrate that naive comparisons of AI vs. non-AI work may reach misleading conclusions, and we provide actionable guidance for controlling these confounds.
This program is tentative and subject to change.
Fri 9 OctDisplayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change
11:00 - 12:30 | Mining Software Repositories: Large-Scale Evidence on Code, CI and EvolutionESEM - Technical Track / ESEM - Emerging Results, Vision, and Reflection Papers Track / ESEM - Journal First Track / ESEM - Software Engineering in Practice Track at Mars | ||
11:00 15mTalk | Characterizing Feedback Statements in Machine Learning Jupyter Notebooks ESEM - Technical Track Arumoy Shome Delft University of Technology, Luís Cruz TU Delft, Diomidis Spinellis AUEB & TU Delft, Arie van Deursen TU Delft | ||
11:15 15mTalk | Does GCC Compiler Development Pay for Itself? GCC Versions’ Impact on Energy Consumption ESEM - Technical Track Brage Isak Bakkane Keiserås Department of Informatics, University of Oslo, Bo Markussen Department of Mathematics, University of Copenhagen, Michael Kirkedal Thomsen University of Oslo & University of Copenhagen, Maja H. Kirkeby Roskilde University, Ingrid Chieh Yu University of Oslo, Ken Friis Larsen Department of Computer Science, University of Copenhagen | ||
11:30 15mTalk | A Preliminary Analysis of the Impact of AI Assisted/Generated Code on Quality, Centrality and Review Time at Scale ESEM - Software Engineering in Practice Track Audris Mockus University of Tennessee, Peter Rigby Concordia University; Meta, Don Stewart Meta, Chandra Sekhar Maddila Meta Platforms, Inc., Nachiappan Nagappan Meta Platforms, Inc. | ||
11:45 15mTalk | Test Alert Snooze: An Empirical Study of Consecutive Test Failures on CI ESEM - Technical Track Ayane Shirakawa Nara Institute of Science and Technology, Tatsuya Shirai Nara Institute of Science and Technology, Yutaro Kashiwa Nara Institute of Science and Technology, Masanari Kondo Kyushu University, Yasutaka Kamei Kyushu University, Hajimu Iida Nara Institute of Science and Technology | ||
12:00 15mTalk | Development and evolution of Xtext-based DSLs on GitHub: an empirical investigation ESEM - Journal First Track Weixing Zhang Karlsruhe Institute of Technology (KIT), Daniel Strüber Chalmers | University of Gothenburg / Radboud University, Regina Hebig Universität Rostock, Rostock, Germany | ||
12:15 10mTalk | QModel: A Time-Aware GitHub Mining Framework for Empirical Software Quality Studies ESEM - Emerging Results, Vision, and Reflection Papers Track Dmytro Polishchuk Jagiellonian University | ||