Bias Ahead: Sensitive Prompts as Early Warnings for Fairness in Large Language Models
Large Language Models (LLMs) are being increasingly integrated into software systems, offering powerful capabilities but also raising concerns about fairness. Existing fairness benchmarks, however, focus on stereotype-specific associations, which limit their ability to anticipate risks in diverse, real-world contexts. In this paper, we propose sensitive prompts as a new abstraction for fairness evaluation: inputs that are not inherently biased but are more likely to elicit biased or inadequate responses due to the sensitivity of their content.
We construct and release SensY, a dataset of 12,801 prompts, categorized as sensitive and non-sensitive, spanning seven thematic domains, combining synthetic generation and real user inputs. Using this dataset, we query three open-source LLMs and manually analyze 4,500 responses to evaluate their adequacy in answering sensitive prompts. Results show that while models often provide factually correct answers, they frequently fail to acknowledge the ethical, relational, or contextual implications of sensitive inputs. In addition, we develop an automated classifier for predicting prompt sensitivity, achieving robust performance on sensitive prompts. Our findings demonstrate that prompt sensitivity can serve as an effective early-warning mechanism for fairness risks in LLMs. This perspective shifts fairness assessment from reactive mitigation toward preventive design, enabling developers to anticipate and manage bias before deployment.
Tue 17 MarDisplayed time zone: Athens change
14:00 - 15:30 | |||
14:00 15mTalk | Gender Bias in Generative AI-assisted Recruitment Processes Workshops & Tutorials Martina Ullasci Politecnico di Torino, Marco Rondina Politecnico di Torino, Riccardo Coppola Politecnico di Torino, Antonio Vetrò Politecnico di Torino | ||
14:15 25mTalk | Bias Ahead: Sensitive Prompts as Early Warnings for Fairness in Large Language Models Workshops & Tutorials Gianmario Voria University of Salerno, Martina De Lucia University of Salerno, Alessandra Raia University of Salerno, Andrea De Lucia University of Salerno, Gemma Catolino University of Salerno, Fabio Palomba University of Salerno | ||
14:40 25mTalk | Evaluation of Data Quality Disparity and Implications for Fair Machine Learning Workshops & Tutorials Mohit Sharma IIT Delhi, Pratik Mishra IBM Research, Sandeep Hans IBM India Research Lab, Abhijnan Chakraborty IIT Kharagpur, Vijay Arya IBM Research Pre-print | ||