Exploring and Improving Real-World Vulnerability Data Generation via Prompting Large Language Models
Data-driven approaches have proven to be promising for vulnerability analysis, contingent on quality and sizable training data being available. Several dedicated vulnerability data generation techniques have demonstrated impressive merits, yet they are limited to simple (single-line injection induced) vulnerabilities only and suffer from overfitting to seed samples. Large language models (LLMs) may overcome these fundamental limitations as they are known to be effective at generative tasks. yet it remains unclear how they would fare for the task of vulnerable sample generation. In this paper, we explore the potentials and gaps of seven state-of-the-art (SOTA) LLMs for that task via prompting.
We reveal that the LLMs are capable of injecting vulnerabilities, with advanced prompting strategies such as few-shot in-context learning and our new vulnerability-introducing code-change semantics (VICS) guided prompting boosting the effectiveness, achieving up to 93% success rate on a synthetic dataset and 88% on real-world code. The LLMs can effectively perform both single- and multi-line injections, addressing a key limitation of prior work. Notably, they exhibit a strong preference to replacement edits, different from ground-truth patterns, and their effectiveness varies across CWE types. Furthermore, LLMs, particularly with VICS, outperform existing SOTA vulnerability generators %(VulGen and VGX) with success rate improvements of up to 210%-343%. Crucially, the LLM-generated data substantially improves the performance of downstream DL-based vulnerability analysis models, especially with multi-line injections, boosting their accuracy by up to 36.2%. The generated samples also enhances the effectiveness of other LLMs for vulnerability analysis via RAG, with multi-line-injected samples yielding up to 16.5% gains. Most importantly, our findings reveal that, augmenting existing DL models with high-quality, LLM-generated data can lead to vulnerability analysis performance (up to 67.50% accuracy) superior to those of even the most advanced LLMs performing the same analysis (up to 27.40%), indicating the usefulness of the LLM-generated vulnerability data at present and in longer term.
| Guangbei_paper2613_presentation_Slides (wsu-beamer.pdf) | 381KiB |
Thu 16 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
16:00 - 17:30 | Dependability and Security 7Research Track at Oceania X Chair(s): Kaixuan Li Nanyang Technological University | ||
16:00 15mTalk | WhisperCatcher: Demystifying Unauthorized and Encrypted Private Data Transmission in Android ApplicationsDistinguished Paper Award Research Track Zhaoyu Qiu Xi'an Jiaotong University, Ming Fan Xi'an Jiaotong University, Bocan Ma Xi'an Jiaotong University, Yutian Tang University of Glasgow, United Kingdom, Lei Xue Sun Yat-Sen University, Haijun Wang Xi'an Jiaotong University, Ting Liu Xi'an Jiaotong University | ||
16:15 15mTalk | Exploring and Improving Real-World Vulnerability Data Generation via Prompting Large Language Models Research Track Guangbei Yi Washington State University, Yu Nong University at Buffalo, SUNY, Minzhang Li Washington State University, Haipeng Cai University at Buffalo, SUNY DOI Pre-print Media Attached File Attached | ||
16:30 15mTalk | TaintP2X: Detecting Taint-Style Prompt-to-Anything Injection Vulnerabilities in LLM-Integrated Applications Research Track HeJunjie , Shenao Wang Huazhong University of Science and Technology, Yanjie Zhao Huazhong University of Science and Technology, Xinyi Hou Huazhong University of Science and Technology, Zhao Liu 360 AI Security Lab, Quanchen Zou 360 AI Security Lab, Haoyu Wang Huazhong University of Science and Technology | ||
16:45 15mTalk | CoBrA: Context-, Branch-sensitive Static Analysis for Detecting Taint-style Vulnerabilities in PHP Web Applications Research Track Yichao Xu , Mingqing Kang Johns Hopkins University, Neil Thimmaiah University of Illinois Chicago, Rigel Gjomemo University of Illinois Chicago, V. N. Venkatakrishnan University of Illinois Chicago, Yinzhi Cao Johns Hopkins University | ||
17:00 15mTalk | Project-Level Resource Leak Detection through Agent-based Ownership Analysis and Repair Pattern Verification Research Track Chengxin Xu Institute of Information Engineering, Chinese Academy of Sciences, xiu zhang Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China, Xiaorui Gong Institute of Information Engineering, Chinese Academy of Science Media Attached | ||
17:15 15mTalk | Understanding DevOps Security of Google Workspace Apps Research Track Liuhuo Wan , Chuan Yan University of Queensland, Zicong Liu University of Queensland, Haoyu Wang Huazhong University of Science and Technology, Guangdong Bai City University of Hong Kong Media Attached | ||