Evaluating Large Language Models for Security Bug Report Prediction
Early detection of security bug reports (SBRs) is critical for timely vulnerability mitigation. We present an evaluation of prompt engineering and fine-tuning approaches for predicting SBRs using Large Language Models (LLMs). Our results revealed a distinct trade-off between the two approaches. Using the prompted proprietary models, we observed the highest sensitivity to SBRs with an average G-measure of 0.77 among all the tested datasets, albeit at the cost of a higher false-positive rate, resulting in a precision of only 0.22. In contrast, fine-tuned models offered more viable solutions, achieving an average G-measure of 0.51 with a 10 to 50 times reduction in inference latency compared to proprietary LLMs. While proprietary models offer accessibility, specialized fine-tuned models remain the optimal choice for privacy-preserving and cost-effective SBR prediction.
Tue 17 MarDisplayed time zone: Athens change
16:00 - 17:30 | |||
16:00 20mTalk | Evaluating Large Language Models for Security Bug Report Prediction Workshops & Tutorials Farnaz Soltaniani Technische Universität Clausthal, Shoaib Razzaq Technical University of Clausthal, Mohammad Ghafari TU Clausthal | ||
16:20 20mTalk | Towards Project-Aware Actionability Detection for Coding Rule Violations Workshops & Tutorials Széles Csoma Lázár University of Szeged, Department of Software Engineering, Gergő Balogh Department of Software Engineering, University of Szeged | ||
16:40 10mTalk | Don’t Mind the Mesh: An Empirical Study of Istio Service Mesh Security in GitHub Workshops & Tutorials Kohsuke Sonoda Aalto university, Jose Luis Martin-Navarro Aalto University, Tuomas Aura Aalto University | ||
16:50 20mTalk | Can Developers rely on LLMs for Secure IaC Development? Workshops & Tutorials | ||
17:10 20mTalk | Persistent Human Feedback, LLMs, and Static Analyzers for Secure Code Generation and Vulnerability Detection Workshops & Tutorials | ||