InDe-LLM: Defending Against Jailbreak Attacks in LLM-Powered Systems via Intention Disentangling
Jailbreak attacks have been regarded as a crucial threat to LLM-powered software systems. Recent studies indicate the existence of a steering vector within models’ internal activations, which can adjust a model’s propensity to reject user requests, and thus is regarded as an effective approach for training-free defense. However, attackers may wrap their malicious intentions within a seemingly benign context, which shifts the distribution of harmful prompts toward benign inputs along the steering vector, effectively bypassing existing defense approaches. In this work, we propose a defense framework InDe-LLM based on intention disentangling. By projecting the embedding of inputs into a benign-invariant subspace, we could disentangle the harmful intentions of jailbreak prompts without affecting benign inputs. Next, such disentangled harmful intentions can be easily identified based on LLMs’ well-aligned concept of harmfulness, and rejected through activation steering. Our experiments show that InDe-LLM achieves high defense effectiveness, outperforming baselines by 27.2%–43.5% across three models and ten attacks while preserving high utility on benign inputs. Moreover, our evaluation demonstrates that it exhibits high transferability to unseen attacks.
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
10:30 - 12:30 | Security 2Research Papers / Industry Papers at MB 3.435 Chair(s): Amin Milani Fard New York Institute of Technology | ||
10:30 20mTalk | YASA: Scalable Multi-Language Taint Analysis on the Unified AST at Ant Group Industry Papers Yayi Wang Ant Group, Shenao Wang Huazhong University of Science and Technology, Jian Zhao Huazhong University of Science and Technology, Shaosen Shi Ant Group, Ting Li Ant Group, Yan Cheng Ant Group, Lizhong Bian Ant Group, Kan Yu Ant Group, Yanjie Zhao Huazhong University of Science and Technology, Haoyu Wang Huazhong University of Science and Technology | ||
10:50 20mTalk | InDe-LLM: Defending Against Jailbreak Attacks in LLM-Powered Systems via Intention Disentangling Research Papers YujueWang Tsinghua University, Quan Zhang East China Normal University, Chijin Zhou East China Normal University, Gwihwan Go Tsinghua University, Dalong Shi AVIC International Digital Network Technology Co., Ltd., Yu Jiang Tsinghua University | ||
11:10 20mTalk | Characterizing Trust Boundary Vulnerabilities in TEE Container Systems: An Empirical Study Research Papers Weijie Liu Nankai University, Hongbo Chen Indiana University Bloomington, Shuo Huai Nankai University, Zhen Xu Nanyang Technological University, Wenhao Wang Institute of Information Engineering, CAS, XiaoFeng Wang Nanyang Technological University, Danfeng Zhang Duke University, Zhi Li Huazhong University of Science and Technology, Haixu Tang Indiana University Bloomington, Zheli Liu Nankai University DOI Pre-print | ||
11:30 20mTalk | GadgetHunter: Region-Based Neuro-Symbolic Detection of Java Deserialization Vulnerabilities Research Papers Kaixuan Li Nanyang Technological University, Jian Zhang Beihang University, Chong Wang Nanyang Technological University, Sen Chen Nankai University, Zong Cao Imperial Global Singapore, Min Zhang East China Normal University, Yang Liu Nanyang Technological University Pre-print | ||
11:50 20mTalk | ReGA: Model-based Safeguard for LLMs via Representation-Guided Abstraction Research Papers Pre-print | ||
12:10 20mTalk | Two-Level Adaptation for Budget-Constrained Continuous Dynamic Dependence Analysis Research Papers Link to publication Pre-print | ||