ESEIW 2026
Sun 4 - Fri 9 October 2026 München, Germany

This program is tentative and subject to change.

Background: Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.

Aims: We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.

Method: In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a conceptual design and a high-fidelity semi-interactive prototype evaluated through a follow-up survey with N=43 practitioners.

Results: Participants identified trust-calibration as the central challenge. The study yielded a three-level review workflow (overview, file-analysis, code snippet review) supported by seven design constructs (chunk, risk-per-line, risk-per-file, judge, walk-through, zooming in/out, and security cage). In the validation survey, all three workflow levels scored above the neutral midpoint (means 3.50–3.91 on a five-point scale). Of the respondents, 63% expected reduced overall review effort, and 52% reduced trust-assessment effort, relative to their current tools. These findings suggest that the design constructs indicate a positive direction for future tool development.

Conclusions: Reviewing LLM-generated multi-file changes is a trust-calibration problem rather than a diffing problem. The three-level workflow and the seven constructs we report give tool designers a conceptual framework for building AI-ready code review tools that surface risk and confidence signals at the granularity at which developers allocate attention.

Data Availability: https://doi.org/10.5281/zenodo.20124352

This program is tentative and subject to change.

Fri 9 Oct

Displayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change

14:00 - 15:30
14:00
15m
Talk
How Developers Use Relation Chains in Code Review: An Empirical Study Across Three Open-Source Ecosystems
ESEM - Technical Track
Ahmed Belhouchette ENSI, Mannouba University, Moataz Chouchen Concordia University, Marouene Chaieb National School of Computer Science, Mohammad Hamdaqa Polytechnique Montreal, Abdelwahab Hamou-Lhadj Concordia University, Montreal, Canada
14:15
15m
Talk
How Reliable Is LLM-as-Judge for Patch Correctness Assessment? An Empirical Study
ESEM - Technical Track
Shanggui Zhan School of Computer Science and Technology, Hangzhou Dianzi University; Zhejiang Key Laboratory of New Industrial Internet Control Technology, Xingqi Wang School of Computer Science and Technology, Hangzhou Dianzi University; Zhejiang Key Laboratory of New Industrial Internet Control Technology, Dan Wei School of Computer Science and Technology, Hangzhou Dianzi University; Zhejiang Key Laboratory of New Industrial Internet Control Technology, Xin Xiang chool of Computer Science and Technology, Hangzhou Dianzi University
14:30
15m
Talk
Trust-Calibrated Code Review: A Participatory Design Study of Review Workflows for LLM-Generated Multi-File Changes
ESEM - Software Engineering in Practice Track
Lo Heander Lund University, Agnia Sergeyuk JetBrains Research, Ilya Zakharov JetBrains Research, Emma Söderberg Lund University, Nikita Mukhortov JetBrains
14:45
15m
Talk
Code Review as Decision-Making - Building a Cognitive Model from the Questions Asked During Code Review
ESEM - Journal First Track
Lo Heander Lund University, Emma Söderberg Lund University, Christofer Rydenfält Lund University
15:00
10m
Talk
AI-to-AI Code Review of GitHub Pull Requests
ESEM - Emerging Results, Vision, and Reflection Papers Track
Niruthiha Selvanayagam Ecole de Technologie Supérieure, Taher A. Ghaleb Trent University
15:10
10m
Talk
How Do Software Professionals Evaluate AI-Generated Code? (Registered Report)
ESEM - Registered Reports Track
Samuli Määttä University of Oulu, Hera Arif Dalhousie University, Burak Turhan University of Oulu, Paul Ralph Dalhousie University, Markus Kelanti University of Oulu
Pre-print
15:20
10m
Talk
CWEFT: CWE-aware Evaluation of Free-text vs. Typed Prompts
ESEM - Emerging Results, Vision, and Reflection Papers Track
Lanaya Aziza Carbonell Bilkent University, Anil Koyuncu Bilkent University