The Illusion of Agentic Complexity in README.md Generation: Evaluating Single-Agent vs. Multi-Agent RAG
This program is tentative and subject to change.
Large Language Models (LLMs) are increasingly utilized to automate software engineering tasks, such as generating repository-level documentation. While Multi-Agent Systems (MAS) are often adopted under the premise that task decomposition improves performance, the impact of architectural complexity on practical efficiency remains under-examined. This study empirically evaluates Retrieval-Augmented Generation (RAG) dependent architectures for the generation of README files for GitHub repositories. A systematic comparison is conducted between a Single-Agent pipeline, a specialized MAS, and a developer-guided planning (Dev-Plan) variant, benchmarked against LARCH, a state-of-the-art baseline, and the original ground truth. Results indicate that our proposed RAG architectures significantly outperform LARCH. A critical architectural trade-off is identified: the Single-Agent pipeline achieves lexical quality comparable to MAS while reducing token consumption by 86% and operating at twice the speed. In contrast, manual taxonomy analysis demonstrates that MAS achieves high structural consistency (98%), resolving the formatting issues observed in single-agent approaches. Autonomous planning is identified as the primary pipeline bottleneck; incorporating lightweight developer-guided plans (Dev-Plan) eliminates structural failures and produces the highest overall documentation quality, surpassing both autonomous AI, baseline, and ground-truth.
This program is tentative and subject to change.
Fri 18 SepDisplayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change
14:00 - 15:30 | Session 26 - Facts, Faults & Future ChallengesResearch Papers Track / Registered Reports / Visions and Emerging Results Track / Tool Demonstration and Data Showcase Track at A59S Chair(s): Michael J. Decker Bowling Green State University Theme: Empirical Software Engineering | ||
14:00 10mPaper | A Security-by-Design Evaluation Framework for Ontology-Grounded Code Generation: A Pre-Registered Confirmatory Study (Stage 1) Registered Reports Pedro Farinha Independent Researcher Link to publication Pre-print | ||
14:10 10mPaper | Can we disentangle Size from Quality to better understand Software Maintenance & Evolution? A Study with 2M Scratch projects Registered Reports Pre-print | ||
14:20 20mPaper | Assessing the User Interaction Cost of Keyboard Navigation in Web Applications Research Papers Track Christina Z. Chaniotaki University of Southern California, Paul T. Chiou University of Southern California, William G.J. Halfond University of Southern California Pre-print | ||
14:40 20mPaper | Demystifying Checker Bugs in Deep Learning Libraries Research Papers Track Nima Shiri Harzevili York University, Jiho Shin York University, Gias Uddin York University, Canada, Jinqiu Yang Concordia University, Junjie Wang Institute of Software at Chinese Academy of Sciences, Song Wang York University, Zhen Ming (Jack) Jiang York University, Nachiappan Nagappan Meta Platforms, Inc. | ||
15:00 10mShort-paper | IBN-Drift: A Multi-Dimensional Benchmark Dataset for Intent Drift Tool Demonstration and Data Showcase Track Hanlin Liu Inner Mongolia University, Da Li Inner Mongolia University, Shuai Ma Inner Mongolia University, Jiaxing Pi Inner Mongolia University, Hua Li College of Computer Science, Inner Mongolia University, Chunyan An College of Computer Science, Inner Mongolia University | ||
15:10 10mShort-paper | The Illusion of Agentic Complexity in README.md Generation: Evaluating Single-Agent vs. Multi-Agent RAG Visions and Emerging Results Track Abu Saleh University of L'Aquila, Tesfay Welegebreal Tesfay University of L'Aquila, Phuong T. Nguyen University of L’Aquila, Juri Di Rocco University of L'Aquila, Muhammad Umar Zeshan University of L’Aquila, Davide Di Ruscio University of L'Aquila Pre-print | ||