CAIN 2026
Sun 12 - Sat 18 April 2026 Rio de Janeiro, Brazil
co-located with ICSE 2026
Sun 12 Apr 2026 11:08 - 11:20 at Oceania X - Engineering Agentic Systems Chair(s): Henry Muccini

Current benchmarks for evaluating software engineering agents, such as SWE-Bench Verified, are predominantly derived from GitHub issues and fail to accurately reflect how developers interact with chat-based coding assistants in integrated development environments (IDEs). We posit that this mismatch leads to a systematic overestimation of agent’s capabilities in real-world scenarios, especially bug fixing. We introduce a novel benchmarking framework that transforms existing formal benchmarks into realistic user queries through systematic analysis of developer interaction patterns with chat-based agents. Our methodology is flexible and can be easily extended to existing benchmarks. In this paper, we apply our testing framework to SWE-Bench Verified, the TypeScript subset of Multi-SWE-Bench and a private benchmark, SWE-Bench C# and transform formal GitHub issue descriptions into realistic user-style queries based on telemetry analysis of a popular chat-based agent interactions. Our findings reveal that existing benchmarks significantly overestimate agent capabilities for some models by $>$50% over baseline performance for public benchmarks and $\sim$10-16% for our internal benchmark. This work establishes a new paradigm for evaluating interactive chat-based software engineering agents through benchmark mutation techniques.

Sun 12 Apr

Displayed time zone: Brasilia, Distrito Federal, Brazil change

11:00 - 12:30
Engineering Agentic SystemsIndustry Track / Research Track / CAIN Program at Oceania X
Chair(s): Henry Muccini University of L'Aquila, Italy
11:00
8m
Short-paper
Towards an Approach for Specifying Intelligent Systems Involving Foundation Model Based AgentsShort Paper
Research Track
Júlia Condé Araújo Department of Informatics - Pontifical Catholic University of Rio de Janeiro (PUC-Rio), Marina Condé Araújo Department of Informatics - Pontifical Catholic University of Rio de Janeiro (PUC-Rio), José M. C. Boaro Pontifical Catholic University of Rio de Janeiro, Marcos Kalinowski Pontifical Catholic University of Rio de Janeiro (PUC-Rio)
11:08
12m
Full-paper
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent EvaluationFull Paper
Research Track
Spandan Garg Microsoft Corporation, Benjamin Steenhoek Microsoft, Yufan Huang
Pre-print
11:20
12m
Industry talk
Context Sharing Strategies for Production Multi-Agent AI Systems: An Industrial EvaluationFull PaperVirtual Attendance
Industry Track
Minav Suresh Patel Independent Researcher, Rohit Dhawan Independent Researcher, Priyank Desai Amazon.com, Ankush Dhar Amazon
Media Attached
11:32
8m
Industry talk
Managing Variability in Industrial AI Agents for Manufacturing: Experiences at HitachiShort Paper
Industry Track
DOI
11:40
8m
Short-paper
Architecting AgentOps Needs CHANGEShort Paper
Research Track
Shaunak Biswas IIIT Hyderabad, Hiya Bhatt IIIT Hyderabad, Karthik Vaidhyanathan IIIT Hyderabad
Pre-print
11:48
12m
Industry talk
How to Build AI Agents by Augmenting LLMs with Codified Human Expert Domain Knowledge? A Software Engineering FrameworkFull Paper
Industry Track
Choro Ulan Uulu Eindhoven University of Technology, Mikhail Kulyabin , Iris Fuhrmann , Jan Joosten , Nuno Miguel Martins Pacheco , Filippos Petridis , Rebecca Johnson , Jan Bosch Chalmers University of Technology, Helena Holmström Olsson Malmö University
12:00
12m
Full-paper
Agentic AI Architecture for Evaluating and Improving Reinforcement Learning PipelinesFull Paper
Research Track
Evangelos Ntentos University of Vienna, Uwe Zdun University of Vienna
12:12
18m
Live Q&A
Joint Q&A (Engineering Agentic Systems)
CAIN Program