Keynote: On the Evaluation of AI Coding Agents
AI coding agents are rapidly evolving and capable of performing complex software engineering tasks. As these systems move toward real-world deployment, rigorous and meaningful evaluation becomes increasingly critical for understanding their capabilities, limitations, and risks. In this talk, I present a series of evaluation methods that aim to bridge this gap, covering code generation, question answering and agentic systems. I will discuss lessons learned from building realistic evaluation frameworks and outline future directions toward trustworthy and effective AI software engineering agents.
I am a Principal Research Scientist at ByteDance (字节跳动). I received my PhD degree from Laboratory for Foundations of Computer Science (LFCS), The University of Edinburgh under supervision of Dr. Ajitha Rajan.
At ByteDance, I lead the Trae Research team (ByteDance Software Engineering Lab), where we conduct research on AI agents for software engineering including the application and evaluation of AI agents, and training LLMs for agents. I am also responsible for academic development and university collaboration.
I am passionate about building practical software testing, analysis, and debugging systems to predict, detect, diagnose, and fix bugs for all kinds of software systems.
Outside of work, I enjoy going to the gym.
