ICSE 2026
Sun 12 - Sat 18 April 2026 Rio de Janeiro, Brazil

AI coding agents are rapidly evolving and capable of performing complex software engineering tasks. As these systems move toward real-world deployment, rigorous and meaningful evaluation becomes increasingly critical for understanding their capabilities, limitations, and risks. In this talk, I present a series of evaluation methods that aim to bridge this gap, covering code generation, question answering and agentic systems. I will discuss lessons learned from building realistic evaluation frameworks and outline future directions toward trustworthy and effective AI software engineering agents.

I am a Principal Research Scientist at ByteDance (字节跳动). I received my PhD degree from Laboratory for Foundations of Computer Science (LFCS), The University of Edinburgh under supervision of Dr. Ajitha Rajan.

At ByteDance, I lead the Trae Research team (ByteDance Software Engineering Lab), where we conduct research on AI agents for software engineering including the application and evaluation of AI agents, and training LLMs for agents. I am also responsible for academic development and university collaboration.

I am passionate about building practical software testing, analysis, and debugging systems to predict, detect, diagnose, and fix bugs for all kinds of software systems.

Outside of work, I enjoy going to the gym.

Tue 14 Apr

Displayed time zone: Brasilia, Distrito Federal, Brazil change

14:00 - 15:30
Evaluation, Reliability, and Engineering PracticeAGENT at Oceania VIII
Chair(s): Oshani Weerakoon Department of Computing, University of Turku
14:00
30m
Keynote
Keynote: On the Evaluation of AI Coding Agents
AGENT
K: Chao Peng ByteDance
14:30
6m
Talk
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
AGENT
Sreemaee Akshathala IIIT Hyderabad, Bassam Adnan IIIT Hyderabad, Mahisha Ramesh IIIT Hyderabad, Karthik Vaidhyanathan IIIT Hyderabad, Basil Muhammed MontyCloud, Kannan Parthasarathy MontyCloud
14:36
6m
Talk
PerfBench: Can Agents Resolve Real-World Performance Bugs?Virtual Attendance
AGENT
Spandan Garg Microsoft Corporation, Roshanak Zilouchian Moghaddam Microsoft, Neel Sundaresan Microsoft
Pre-print Media Attached
14:42
6m
Talk
SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs
AGENT
Arihant Tripathy IIIT Hyderabad, India, Ch Pavan Harshit IIIT Hyderabad, India, Karthik Vaidhyanathan IIIT Hyderabad
Link to publication DOI Pre-print
14:48
6m
Talk
Context Matters: Evaluating MCP-Based Context-Aware AI — A Case Study of Email Communications in Nonprofit Organizations
AGENT
Nitin Gupta University of Victoria, Jayani Samaraweera University of Victoria, Raaj Chatterjee Meaningful Technology Inc., Riya Shrestha University of Victoria, Trinity West University of Victoria, Dana Damian University of Victoria
14:54
6m
Talk
Toward Agentic Software Project Management: A Vision and Roadmap
AGENT
Lakshana Assalaarachchi Monash University, Australia, Zainab Masood Prince Sultan University, Rashina Hoda Monash University, John Grundy Monash University
15:00
6m
Talk
Not All Problems Are Nails, Not All Tools Should Be Hammers: A Position Paper on Agent Usage in Software Engineering Tasks
AGENT
Juuso Rytilahti Department of Computing, University of Turku, Panu Puhtila University of Turku, Oshani Weerakoon Department of Computing, University of Turku, Erkki Kaila Department of Computing, University of Turku, Tuomas Mäkilä University of Turku
15:06
24m
Live Q&A
Session 3 Joint Q&A
AGENT