Evaluation of Data Quality Disparity and Implications for Fair Machine Learning
The performance of machine learning (ML) models heavily depends on the quality of the data they are trained on. While prior work often treats data quality as uniform across a dataset, we investigate whether it varies across different population subgroups within a dataset and examine its implications, a phenomenon we refer to as Data Quality Disparity (DQD). Our analysis reveals that many real-world datasets inherently exhibit DQD, with underrepresented or marginalized groups frequently associated with lower-quality data compared to more privileged subgroups. Consequently, training ML models on raw data or applying standard preprocessing techniques without accounting for DQD can lead to uneven prediction performance across demographics. We introduce metrics to quantify DQD and assess its potential for discrimination, using both structured (tabular) and unstructured (image) datasets. Our work explores the connections between data quality and algorithmic fairness and underscores the need for developing fair data processing pipelines and learning algorithms that are DQD-aware.
Tue 17 MarDisplayed time zone: Athens change
14:00 - 15:30 | |||
14:00 15mTalk | Gender Bias in Generative AI-assisted Recruitment Processes Workshops & Tutorials Martina Ullasci Politecnico di Torino, Marco Rondina Politecnico di Torino, Riccardo Coppola Politecnico di Torino, Antonio Vetrò Politecnico di Torino | ||
14:15 25mTalk | Bias Ahead: Sensitive Prompts as Early Warnings for Fairness in Large Language Models Workshops & Tutorials Gianmario Voria University of Salerno, Martina De Lucia University of Salerno, Alessandra Raia University of Salerno, Andrea De Lucia University of Salerno, Gemma Catolino University of Salerno, Fabio Palomba University of Salerno | ||
14:40 25mTalk | Evaluation of Data Quality Disparity and Implications for Fair Machine Learning Workshops & Tutorials Mohit Sharma IIT Delhi, Pratik Mishra IBM Research, Sandeep Hans IBM India Research Lab, Abhijnan Chakraborty IIT Kharagpur, Vijay Arya IBM Research Pre-print | ||