An Empirical Study of Fine-Grained Entity Relationships for Tracing Natural Language and Code Vulnerability Artifacts
Understanding software vulnerabilities requires analyzing fine-grained entities and their complex relationships across natural language (NL) artifacts and source code. Vulnerability descriptions often interweave vulnerability triggers (VT), crash phenomena (CP), and after-fix (AF) actions within the same sentence, making it challenging to distinguish root causes, failure symptoms, and remediation strategies. Additionally, missing key NL entities further hinders traceability, limiting an analyst’s ability to determine why and how a vulnerability was introduced and resolved. To address these challenges, we conduct an empirical study on fine-grained entity relationships using a manually curated dataset of 1,000 vulnerabilities. We extract phrase-level VT, CP, and AF entities, categorize them into structured taxonomies, and analyze cross-entity relationships within NL artifacts and source code, uncovering recurring patterns in vulnerability evolution and remediation strategies. Furthermore, we investigate the automation of vulnerability entity extraction using different approaches, showing that ELECTRA [8], a state-of-the-art pre-trained language model, along with other LLM-based approaches, outperforms other methods.
Thu 16 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
16:00 - 17:30 | Evolution 3Research Track / New Ideas and Emerging Results (NIER) at Oceania VIII Chair(s): Antu Saha William & Mary | ||
16:00 15mTalk | MINES: Explainable Anomaly Detection through Web API Invariant Inference Research Track Wenjie Zhang National University of Singapore, Yun Lin Shanghai Jiao Tong University, Kwok Chun Fung Amos National University of Singapore, Xiwen Teoh National University of Singapore, Xiaofei Xie Singapore Management University, Frank Liauw Government Technology Agency Singapore, Hongyu Zhang Chongqing University, Jin Song Dong National University of Singapore | ||
16:15 15mTalk | Actionable Warning Is Not Enough: Recommending Valid Actionable Warnings with Weak Supervision Research Track Zhipeng Xue Zhejiang University, Zhipeng Gao Shanghai Institute for Advanced Study - Zhejiang University, Tongtong Xu Huawei, Xing Hu Zhejiang University, Xin Xia Zhejiang University, Shanping Li Zhejiang University | ||
16:30 15mTalk | SeRe: A Security-Related Code Review Dataset Aligned with Real-World Review Activities Research Track Zixiao Zhao , Yanjie Jiang Tianjin University, Hui Liu Beijing Institute of Technology, Kui Liu Huawei, Lu Zhang Peking University | ||
16:45 15mTalk | Translating PL/I Macro Procedures into Java Using Automatic Templatization and Large Language Models New Ideas and Emerging Results (NIER) | ||
17:00 15mTalk | An Empirical Study of Fine-Grained Entity Relationships for Tracing Natural Language and Code Vulnerability Artifacts Research Track Simin Wang Department of Computer Science, Southern Methodist University, Dallas, Texas, USA 75275-0122, Liguo Huang Southern Methodist University, Shiyi Wei University of Texas at Dallas, Amiao Gao Department of Computer Science, Southern Methodist University, Dallas, Texas, USA 75275-0122, Ruiqi Hu Department of Statistics and Data Science, Vincent Ng Human Language Technology Research Institute, University of Texas at Dallas, Richardson, TX 75083-0688 | ||
17:15 15mTalk | Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted Retrieval Research Track Huihui Huang Singapore Management University, Singapore, Ratnadira Widyasari Singapore Management University, Singapore, Ting Zhang Monash University, Ivana Clairine Irsan Singapore Management University, Jieke Shi Singapore Management University, Han Wei Ang GovTech, Frank Liauw Government Technology Agency Singapore, Eng Lieh Ouh Singapore Management University, Singapore, Lwin Khin Shar Singapore Management University, Hong Jin Kang University of Sydney, David Lo Singapore Management University | ||