A Gray Literature Study on Fairness Requirements in AI-enabled Software Engineering
Today, with the growing obsession with applying Artificial Intelligence (AI) to software across various contexts, much of the focus has been on the effectiveness of AI models, often measured through common metrics such as F1-score, while fairness receives relatively little attention. This paper presents a review of existing gray literature, examining fairness requirements in AI context, with a focus on how they are defined across various application domains, managed throughout the Software Development Life Cycle (SDLC), and the causes, as well as the corresponding consequences of their violation by AI models. Our gray literature investigation shows various definitions of fairness requirements in AI systems, commonly emphasizing non-discrimination and equal treatment across different demographic and social attributes. Fairness requirement management practices vary across the SDLC, particularly in data handling, model training, and monitoring. Fairness requirement violations are frequently linked, but not limited, to poor data quality, data representation issues, algorithmic bias, human judgment, and transparency gaps. The corresponding consequences include social harm, stereotype reinforcement, and data and privacy risks in AI-supported decisions. These findings emphasize the need for consistent frameworks and practices to integrate fairness into AI software and pay as much attention to it as is given to its effectiveness.