ICSE 2026
Sun 12 - Sat 18 April 2026 Rio de Janeiro, Brazil
Thu 16 Apr 2026 15:00 - 15:15 at Oceania II - Testing and Analysis 12 Chair(s): Sam Malek

Deep Learning (DL) libraries (e.g., Pytorch) are popular in the development of AI applications. These libraries are complex and contain bugs. Researchers have proposed various bug-finding techniques for such libraries. Yet, there is much room for improvement. A key challenge in testing DL libraries is the lack of API specifications. Prior testing approaches often inaccurately model the input specifications of DL APIs, resulting in missing valid inputs that could reveal bugs or false alarms from invalid inputs. To address this challenge, we develop Centaur - the first neurosymbolic technique to test DL library APIs using dynamically learned input constraints. Centaur leverages the key idea that formal API constraints can be learned from a small number of seed inputs, and that the learned constraints can be solved using SMT solvers to generate valid and diverse test inputs for the API. We develop a novel grammar that represents first-order logic formulae over API parameters and expresses tensor-related properties (e.g., shape, data type, etc.) as well as relational properties between parameters. We use the grammar to guide a Large Language Model (LLM) to enumerate syntactically correct candidate rules, which are then validated using valid inputs. Further, we develop a custom refinement strategy to prune the set of learned rules to eliminate spurious or redundant rules. The learned constraints are then used to systematically generate valid and diverse inputs for the API by SMT solving (such as Z3) and a specialized sampling technique. We evaluate Centaur for testing PyTorch and TensorFlow. Our results show that Centaur generates constraints more accurately compared to prior approaches, namely DocTer and ACETest. In terms of coverage, Centaur covers 203, 149, and 9608 more branches than TitanFuzz, ACETest and Pathfinder, respectively. Using Centaur, we also detect 23 new bugs in PyTorch and TensorFlow, out of which 11 are already confirmed.

Thu 16 Apr

Displayed time zone: Brasilia, Distrito Federal, Brazil change

14:00 - 15:30
Testing and Analysis 12Research Track at Oceania II
Chair(s): Sam Malek University of California at Irvine
14:00
15m
Talk
Generator Solving for Symbolic Execution
Research Track
Siwei Wei State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences, and University of Chinese Academy of Sciences Beijing, China, Yan Cai Institute of Software at Chinese Academy of Sciences
14:15
15m
Talk
How Good are Input Grammar Miners? An Empirical Study
Research Track
Leon Bettscheider CISPA Helmholtz Center for Information Security, Andreas Zeller CISPA Helmholtz Center for Information Security
File Attached
14:30
15m
Talk
LSPRAG: LSP-Guided RAG for Language-Agnostic Real-Time Unit Test Generation
Research Track
Gwihwan Go Tsinghua University, Quan Zhang East China Normal University, Chijin Zhou East China Normal University, Zhao Wei Tencent, Yu Jiang Tsinghua University
14:45
15m
Talk
Breaking Single-Tester Limits: Multi-Agent LLMs for Multi-User Feature Testing
Research Track
Sidong Feng Monash University, Changhao Du Jilin University, huaxiao liu Jilin University, Qingnan Wang Jilin University, Zhengwei Lv ByteDance, Mengfei Wang ByteDance, Chunyang Chen TU Munich
15:00
15m
Talk
Testing Deep Learning Libraries via Neurosymbolic Constraint Learning
Research Track
M M Abid Naziri North Carolina State University, Shinhae Kim Cornell University, Feiran Qin North Carolina State University, Saikat Dutta Cornell University, Marcelo d'Amorim North Carolina State University
15:15
15m
Talk
MioHint: LLM-Assisted Request Mutation for Whitebox REST API TestingVirtual Attendance
Research Track
Jia Li The Chinese University of Hong Kong, Jiacheng Shen Duke Kunshan University, Yuxin Su Sun Yat-sen University, Michael Lyu The Chinese University of Hong Kong