ICSE 2026
Sun 12 - Sat 18 April 2026 Rio de Janeiro, Brazil
Fri 17 Apr 2026 14:30 - 14:45 at Oceania IX - Testing and Analysis 18 Chair(s): Andy Zaidman

Language is a deep-rooted means of perpetration of stereotypes and discrimination. Large Language Models, now a pervasive technology in our everyday lives, can cause extensive harm when prone to generating toxic responses. The standard way to address this issue is to align the LLM, which, however, dampens the issue without constituting a definitive solution. Therefore, testing LLM even after alignment efforts remains crucial for detecting any residual deviations with respect to ethical standards. We present EvoTox, an automated testing framework for LLMs’ inclination to toxicity, providing a way to quantitatively assess how much LLMs can be pushed towards toxic responses even in the presence of alignment. The framework adopts an iterative evolution strategy that exploits the interplay between two LLMs, the System Under Test (SUT) and the Prompt Generator steering SUT responses toward higher toxicity. The toxicity level is assessed by an automated oracle based on an existing toxicity classifier. We conduct a quantitative and qualitative empirical evaluation using five state-of-the-art LLMs as evaluation subjects having increasing complexity (7–671B parameters). Our quantitative evaluation assesses the cost-effectiveness of four alternative versions of EvoTox against existing baseline methods, based on random search, curated datasets of toxic prompts, and adversarial attacks. Our qualitative assessment engages human evaluators to rate the fluency of the generated prompts and the perceived toxicity of the responses collected during the testing sessions. Results indicate that the effectiveness, in terms of detected toxicity level, is significantly higher than the selected baseline methods (effect size up to 1.0 against random search and up to 0.99 against adversarial attacks). Furthermore, EvoTox yields a limited cost overhead (from 22% to 35% on average).

Fri 17 Apr

Displayed time zone: Brasilia, Distrito Federal, Brazil change

14:00 - 15:30
14:00
15m
Talk
Drivora: A Unified and Extensible Infrastructure for Search-based Autonomous Driving TestingVirtual Attendance
Demonstrations
Mingfei Cheng Singapore Management University, Lionel Briand University of Ottawa, Canada; Lero centre, University of Limerick, Ireland, Yuan Zhou Zhejiang Sci-Tech University
14:15
15m
Talk
CITYWALK: Enhancing LLM-Based C++ Unit Test Generation via Project-Dependency Awareness and Language-Specific Knowledge
Journal-first Papers
Yuwei Zhang Institute of Software Chinese Academy of Sciences, Qingyuan Lu Institute of Software Chinese Academy of Sciences, Kai Liu Shanghai Stock Exchange Technology Co., Ltd., Wensheng Dou Institute of Software Chinese Academy of Sciences, Jiaxin Zhu Institute of Software at Chinese Academy of Sciences, Li Qian Shanghai Stock Exchange Technology Co., Ltd., Chunxi Zhang Shanghai Stock Exchange Technology Co., Ltd., Zheng Lin Shanghai Stock Exchange Technology Co., Ltd., Jun Wei Institute of Software at Chinese Academy of Sciences; University of Chinese Academy of Sciences
14:30
15m
Talk
How Toxic Can You Get? Search-Based Toxicity Testing for Large Language Models
Journal-first Papers
Simone Corbo Politecnico di Milano, Luca Bancale Politecnico di Milano, Valeria De Gennaro Politecnico di Milano, Livia Lestingi DEIB, Politecnico di Milano, Vincenzo Scotti Karlsruhe Institute of Technology, Matteo Camilli Politecnico di Milano
14:45
15m
Talk
Using Cooperative Co-evolutionary Search to Generate Metamorphic Test Cases for Autonomous Driving Systems
Journal-first Papers
Hossein Yousefizadeh University of Ottawa, Shenghui Gu University of Ottawa, Lionel Briand University of Ottawa, Canada; Lero centre, University of Limerick, Ireland, Ali Nasr Waterloo Research Center of Huawei
15:00
15m
Talk
Atomicity Violation Detection for Interrupt-Driven Programs via Incrementally Exploring Concurrent PathsVirtual Attendance
New Ideas and Emerging Results (NIER)
Yuanzhe Liu Xidian University, Bin Yu Xidian University, Ruixue Li Xidian University, Cheng Wen Xidian University, Xu Lu Xidian University, Chu Chen Qufu Normal University, Cong Tian Xidian University
15:15
15m
Talk
EVATest: Domain-Oriented Android GUI Testing based on Reward-Guided Retrieval-Augmented Generation
New Ideas and Emerging Results (NIER)
Bhavana Kondeti The University of Texas at San Antonio, Guanqun Yang Stevens Institute of Technology, USA, Yui Takashima The University of Texas at San Antonio, XUEQING Liu Stevens Institute of Technology, Xiaoyin Wang University of Texas at San Antonio