When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations
Retrieval-Augmented Generation (RAG)-based LLM systems rely on external document corpora that can evolve and change over time. However, current evaluation methodologies (e.g., RAGAS) assess correctness against static snapshots, failing to detect faults when routine updates, factual changes, or noise alter the underlying data. We introduce a metamorphic testing framework that evaluates the consistency of RAG systems under corpus evolution. We formalise a fault taxonomy and 11 mutation operators that systematically perturb the system at both the pre-chunk (retrieval index) and post-chunk (retrieved context) levels. An empirical evaluation across five datasets and over 28k mutants reveals metamorphic violation rates of 4.9-10.2%. In a meta-evaluation against ground truth, our metamorphic oracle achieves F1 scores of 0.927-1.000, while the best RAGAS metric reaches only 0.570. Finally, we provide actionable insights into mitigating these faults through retrieval re-configuration, generator upgrades, and LLM-based reranking.
Wed 14 OctDisplayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change
14:00 - 16:00 | |||
14:00 15mTalk | Benchmarking Contextual Understanding for In-Car Conversational Systems Journal First Philipp Habicht Humboldt Universität Berlin, Lev Sorokin BMW Group, Technical University of Munich, Abdullah Saydemir Technische Universität München, Ken Friedl BMW Group, Andrea Stocco Technical University of Munich, fortiss DOI | ||
14:15 15mTalk | Understanding Bugs in Modern Agentic Frameworks: A Study of Symptoms, Root Causes, and Triggering Conditions Research Papers | ||
14:30 15mTalk | DiffGAN: A Test Generation Approach for Differential Testing of Deep Neural Networks for Image Analysis Journal First Zohreh Aghababaeyan University of Ottawa, Canada, Manel Abdellatif École de Technologie Supérieure, Lionel Briand University of Ottawa, Canada; Lero centre, University of Limerick, Ireland, Ramesh S DOI | ||
14:45 15mTalk | Search-Based Testing of Vision Language Models for In-Car Scene Understanding Industry Showcase Lev Sorokin BMW Group, Technical University of Munich, Chen Yang TU Munich, Ken Friedl BMW Group, Andrea Stocco Technical University of Munich, fortiss Pre-print | ||
15:00 15mTalk | When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations Research Papers Jinhan Kim Università della Svizzera italiana, Samuele Pasini Università della Svizzera italiana, Paolo Tonella USI Lugano Pre-print | ||
15:15 15mTalk | Automated Assertion Generation and Regression Testing for Machine Learning Notebooks Research Papers Yingao (Elaine) Yao Cornell University, Vedant Nimje Veermata Jijabai Technological Institute, Varun Viswanath Dwarkadas J Sanghvi College of Engineering, Saikat Dutta Cornell University | ||
15:30 15mTalk | False-Positive Bug Reports in Deep Learning Compilers: Stages, Root Causes, and Mitigation Journal First Lili Huang College of Intelligence and Computing, Tianjin University, Qingchao Shen Tianjin University, Dong Wang Tianjin University, Yunping Wu Tianjin University, Meng Wang University of Bristol, Junjie Chen Tianjin University DOI | ||
15:45 15mTalk | CIPIHunter: Detecting Configuration-Induced Prediction Instability in Deep Learning Frameworks Research Papers Yanzhou Mu UNIST (Ulsan National Institute of Science and Technology), Korea, Shuo Meng Nantong University, Mijung Kim UNIST, Xiang Chen Nantong University, Chunrong Fang Nanjing University, Zhenyu Chen Nanjing University, Juan Zhai University of Massachusetts at Amherst | ||