Eagle: Leveraging Operations Documents for Comprehensive Benchmark Question Generation
The rapid growth in scale and complexity of modern software systems has intensified the need for intelligent and reliable IT operations. While Artificial Intelligence for IT Operations (AIOps) addresses some challenges, existing solutions predominantly rely on isolated, task-specific models that struggle with interpreting multimodal data, incur high maintenance costs, and lack sufficient transparency. Operations Large Language Models (OpsLLMs) offer unified, knowledge-rich reasoning capabilities, yet their evaluation faces significant barriers, including the absence of Ops-centric evaluation taxonomies, limited availability of public datasets, simplistic question-generation methods, and inadequate quality standards for comprehensive operations tasks.
We present Eagle, a comprehensive benchmarking framework tailored for evaluating OpsLLMs. Deployed inside Huawei, Eagle ingests enterprise product documentation and synthesizes 4{,}845 domain-grounded QA pairs across logs, metrics, traces, and configurations, which are paired with a standardized model evaluation system. This deployment supported multiple evaluations of OpsLLM and culminated in an internal horizontal benchmarking report that informed model selection and rollout decisions. Methodologically, Eagle (i) defines an operations-centric taxonomy aligning core LLM abilities with end-to-end operations tasks; (ii) implements an automated question-generation pipeline with multi-granular quality controls validated by human annotation; and (iii) provides reproducible evaluation suites and metrics for scenario-driven reasoning. In offline studies, Eagle-generated test suites improve expert-rated rubric scores by 22%–49% over state-of-the-art baselines, enabling more precise assessments of anomaly detection, fault diagnosis, and root-cause analysis abilities in OpsLLMs. To foster community adoption and reproducibility, we open-source the framework (https://github.com/NickLennonLiu/eagle_code/) and a sanitized dataset (https://github.com/NickLennonLiu/eagle_data/). By bridging general LLM evaluation and operations practice, Eagle delivers a deployable foundation for advancing large-model applications in AIOps.
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
14:00 - 15:30 | SE and AI 2Industry Papers / Research Papers / Tool Demonstrations at MB 3.210 Chair(s): Rafal Wlodarski Carnegie Mellon Silicon Valley | ||
14:00 10mTalk | IssueGuard: Real-Time Secret Leak Prevention Tool for GitHub Issue Reports Tool Demonstrations Md Nafiu Rahman Brac University , Sadif Ahmed Bangladesh University of Engineering and Techonology, Zahin Wahab The University of British Columbia, Gias Uddin York University, Canada, Rifat Shahriyar Bangladesh University of Engineering and Technology Dhaka, Bangladesh Link to publication DOI Pre-print | ||
14:10 20mTalk | ProofFusion: Improving Neural Theorem Proving via Adaptive Retrieval-Augmented Reasoning Research Papers Manqing Zhang Northwestern Polytechnical University, Yunwei Dong Northwestern Polytechnical University, School of Computer Science and Engineering, Lingru Zhou Northwestern Polytechnical University, Bingxu Xiao Northwestern Polytechnical University, Yepang Liu Southern University of Science and Technology Pre-print | ||
14:30 20mTalk | Eagle: Leveraging Operations Documents for Comprehensive Benchmark Question Generation Industry Papers Yuhe Liu Tsinghua University, Changhua Pei Computer Network Information Center at Chinese Academy of Sciences, Hang Wang Computer Network Information Center, Chinese Academy of Sciences, Longlong Xu Tsinghua University, Xiaogang Dong Huawei, Zhen Feng Huawei, Li Zheng China Academy of Information and Communications Technology, Kehang Ji China Academy of Information and Communications Technology, Dan Pei Tsinghua University | ||
14:50 20mTalk | Unveiling AI-Driven Web Applications: Insights into Characteristics, Functionality, and Compliance Research Papers Liuhuo Wan , Zicong Liu University of Queensland, Chuan Yan University of Queensland, Liujia Wan Northeastern University, Naipeng Dong The University of Queensland, Australia, Zi Huang University of Queensland, Guangdong Bai City University of Hong Kong Pre-print | ||
15:10 20mTalk | One Size Does Fit All: Exploring Model Fusion for Software Engineering Tasks Research Papers Yinggang Qiu National University of Defense Technology, Yihao Qin , Mingyang Geng National University of Defense Technology, Shangwen Wang National University of Defense Technology, Dezun Dong NUDT Link to publication DOI | ||