FSE 2026
Sun 5 - Thu 9 July 2026 Montreal, Canada
Wed 8 Jul 2026 15:00 - 15:20 at MB 2.210 - LLM for SE 5 Chair(s): Banani Roy

Array programming (AP) frameworks (e.g., NumPy and Octave) are widely adopted in scientific computing. Critical defects can jeopardize the entire ecosystem. The stability of API designs enables differential testing on various implementations (e.g., two versions). However, two primary obstacles remain. First, current test generation cannot effectively generate valid inputs, as the APIs (e.g., matrix multiplication) have type constraints and semantic requirements. Second, unit testing approaches test APIs independently, but they share a core N-dimensional array structure (ndarray) as inputs. Modifying one API may alter the ndarray’s properties, breaking the correctness of others. We propose a differential testing tool for array programming, called ArrayDiff. We first collect semantic requirements from NumPy’s APIs and leverage LLMs to transfer NumPy’s requirements to other frameworks. Then, we propose an input-requirement-aware API call generator (IRA-ACG). Based on IRA-ACG, ArrayDiff employs search algorithms to evolve tests while ensuring valid inputs. ArrayDiff can generate valid and complex API call sequences to detect potential differences. We evaluate ArrayDiff and its ablation versions on five AP pairs. They detect 47 valid-input differences and 39 invalid ones, with 23 confirmed as bugs or document issues. IRA-ACG boosts the detection of valid-input differences, which constitute most confirmed bugs. Comparing ArrayDiff with TitanFuzz (LLM-based fuzzer) and Ghostwriter (unit tester) confirms the benefits of IRA-ACG and sequence-level testing.

Wed 8 Jul

Displayed time zone: Eastern Time (US & Canada) change

14:00 - 15:30
LLM for SE 5Tool Demonstrations / Ideas, Visions and Reflections / Research Papers at MB 2.210
Chair(s): Banani Roy University of Saskatchewan
14:00
20m
Talk
Red Teaming LLMs via Linguistic-Aware Fuzzing
Research Papers
Shuai Yuan University of Electronic Science and Technology of China, Nian Luo University Of Electronic Science And Technology Of China, Jingling Sun University of Electronic Science and Technology of China, Yihao Huang National University of Singapore, Singapore, Chengyu Zhang Loughborough University
14:20
10m
Talk
MIMIC-Py: An Extensible Tool for Personality-Driven Automated Game Testing with Large Language Models
Tool Demonstrations
Yifei Chen McGill University, Sarra Habchi Cohere, Canada, Lili Wei McGill University
14:30
10m
Talk
Towards Automated Test Adaptation in Fork Ecosystems via Large Language Models
Ideas, Visions and Reflections
Mukelabai Mukelabai Ruhr University Bochum, Keanu-Wesley Schurkus Ruhr University Bochum, Yannic Noller Ruhr University Bochum, Thorsten Berger Ruhr University Bochum
14:40
20m
Talk
Boosting LLMs for Mutation Generation
Research Papers
Bo Wang Beijing Jiaotong University, Ming Deng Beijing Jiaotong University, Mingda Chen Beijing Jiaotong University, Chengran Yang Singapore Management University, Singapore, Youfang Lin Beijing Jiaotong University, Mark Harman Meta Platforms, Inc. and UCL, Mike Papadakis University of Luxembourg, Jie M. Zhang Mistral AI and King's College London
15:00
20m
Talk
LLM-Assisted Input-Requirement-Aware Differential Testing of Array Programming Frameworks
Research Papers
Zhichao Zhou School of Information Science and Technology, ShanghaiTech University, Jingzhu He ShanghaiTech University
Pre-print
15:20
10m
Talk
AISysRev - LLM-based Tool for Title-abstract Screening
Tool Demonstrations
Aleksi Huotala University of Helsinki, Miikka Kuutila LUT University, Olli-Pekka Turtio University of Helsinki, Simo Sipilä University of Helsinki, Mika Mäntylä University of Helsinki
Pre-print