When LLMs Listen to Experts: Accurate Failure Diagnosis in Operating Systems
Efficient failure diagnosis is critical to maintaining the stability and reliability of operating systems (OSes) in modern industrial environments. In practice, manual failure analysis by on-call engineers (OCEs) faces increasing challenges due to the growing complexity of OS failures, while recent automated diagnosis methods often suffer from low explainability, limiting their practical adoption. Large Language Models (LLMs) hold promise for advancing automated failure diagnosis through their sophisticated reasoning and language generation capabilities. However, traditional LLM-based solutions struggle to integrate domain knowledge and lack the effective interaction mechanisms required for industrial troubleshooting. To address these practical challenges, we present \textit{OScope}, an automated, explainable failure diagnosis framework powered by LLMs. \textit{OScope} leverages historical failure cases to enhance the semantic understanding of anomalies, enabling precise retrieval of relevant troubleshooting guides. The framework further structures the diagnosis process using standard operating procedure (SOP) templates, supporting step-by-step verification and correction with corresponding SOP documents. Importantly, \textit{OScope} facilitates human-in-the-loop collaboration, allowing OCEs to interact with the system for report refinement and practical feedback. We have evaluated \textit{OScope} on a real-world OS failure dataset collected from \textit{Alibaba}. Results show that \textit{OScope} achieves an $AC@5$ of 90%, significantly outperforming baseline methods and demonstrating high diagnostic value in production settings. The diagnostic reports generated by \textit{OScope} have also received positive feedback from OCEs for readability and practical usefulness. Since deployment at \textit{Alibaba}, \textit{OScope} has substantially improved the efficiency of engineers in resolving failures, highlighting its tangible impact as a successful case of applying automated software engineering methods in industry.
Thu 16 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
16:00 - 17:30 | AI for Software Engineering 17Research Track / SE In Practice (SEIP) at Asia IV Chair(s): Martin Monperrus KTH Royal Institute of Technology | ||
16:00 15mTalk | TAAF: A Trace Abstraction and Analysis Framework Synergizing Knowledge Graphs and LLMs Research Track Alireza Ezaz Brock University, Ghazal Khodabandeh Brock University, Majid Babaei University of the Fraser Valley, Naser Ezzati-Jivan Brock University Pre-print | ||
16:15 15mTalk | InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix Caching Research Track Yilun Wang School of Systems Science and Engineering, Sun Yat-sen University, Guangzhou, China, Pengfei Chen Sun Yat-sen University, Haiyu Huang Sun Yat-sen University, Zilong He Sun Yat-sen University, Gou Tan School of Systems Science and Engineering, Sun Yat-sen University, Guangzhou, China, Chuanfu Zhang Sun Yat-Sen University, Jingkai He School of Systems Science and Engineering,Sun Yat-sen University, Guangzhou, China, Zibin Zheng Sun Yat-sen University Pre-print Media Attached | ||
16:30 15mTalk | Order Matters! An Empirical Study on Large Language Models' Input Order Bias in Software Fault Localization Research Track Md Nakhla Rafi Concordia University, Dong Jae Kim DePaul University, Tse-Hsun (Peter) Chen Concordia University, Shaowei Wang University of Manitoba | ||
16:45 15mTalk | When LLMs Listen to Experts: Accurate Failure Diagnosis in Operating Systems SE In Practice (SEIP) Yongxin Zhao , Shenglin Zhang Nankai University, Yuxin Sun Nankai University, Wenwei Gu Nankai University, Yongqian Sun Nankai University, Luping Wang Alibaba Group, Li Shi Alibaba Group, Cheng Huang Alibaba Group, Guodong Yang Alibaba Group, Liping Zhang Alibaba Group, Dan Pei Tsinghua University Media Attached | ||
17:00 15mTalk | MagmaScope: Identifying Root-Cause Changes for Emergency Incident in Large-Scale Cloud Infrastructure SE In Practice (SEIP) Zongyang Li Peking University, Ning Wang Bytedance, Jiliang Liu Bytedance, Yaping Zhang Bytedance, Feifan Tong Bytedance, Zhaoxing Chen Bytedance, Chan Li Bytedance, Ming Liu Bytedance, Xiang Zhang Bytedance, Yifan Wu Peking University, Tong Jia Institute for Artificial Intelligence, Peking University, Beijing, China, Ying Li School of Software and Microelectronics, Peking University, Beijing, China | ||
17:15 15mTalk | Correctness isn’t Efficiency: Runtime Memory Divergence in LLM-Generated Code SE In Practice (SEIP) Prateek Kumar Rajput Zortify and University of Luxembourg, Yewei Song University of Luxembourg, Abdoul Aziz Bonkoungou B Medical Systems and University of Luxembourg, Iyiola E. Olatunji University of Luxembourg, Abdoul Kader Kaboré University of Luxembourg, Jacques Klein University of Luxembourg, Tegawendé F. Bissyandé University of Luxembourg | ||