Conducting systematic reviews is laborious. In the screening or study selection phase, the number of papers can be overwhelming. Recent research has demonstrated that large language models (LLMs) can perform title-abstract screening and support humans in the task. To this end, we developed AISysRev, an LLM-based screening tool implemented as a containerized web application. The tool accepts CSV files containing paper titles and abstracts. Users specify inclusion and exclusion criteria. Multiple different LLMs can be used, such as Gemini, Claude, Mistral or ChatGPT via OpenRouter. We also support locally hosted models and any model compatible with the OpenAI SDK. AISysRev implements both zero-shot and few-shot prompting, and also allows for manual screening through interfaces that display LLM results as guidance for human reviewers. LLM calls are parallelized, meaning screening speed is typically between 100 to 300 papers per minute, depending on the model and the host. To demonstrate the tool’s use in practice, we conducted a qualitative trial study with 137 papers using the tool. Our findings indicate that papers can be classified into four categories: Easy Includes, Easy Excludes, Boundary Includes, and Boundary Excludes. The Boundary cases, where LLMs are prone to errors, highlight the need for human intervention. While LLMs do not replace human judgment in systematic reviews, they can reduce the burden of assessing large volumes of scientific literature. Video: https://www.youtube.com/watch?v=HeblemlgnAQ Tool: https://github.com/EvoTestOps/AISysRev
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
14:00 - 15:30 | LLM for SE 5Tool Demonstrations / Ideas, Visions and Reflections / Research Papers at MB 2.210 Chair(s): Banani Roy University of Saskatchewan | ||
14:00 20mTalk | Red Teaming LLMs via Linguistic-Aware Fuzzing Research Papers Shuai Yuan University of Electronic Science and Technology of China, Nian Luo University Of Electronic Science And Technology Of China, Jingling Sun University of Electronic Science and Technology of China, Yihao Huang National University of Singapore, Singapore, Chengyu Zhang Loughborough University | ||
14:20 10mTalk | MIMIC-Py: An Extensible Tool for Personality-Driven Automated Game Testing with Large Language Models Tool Demonstrations | ||
14:30 10mTalk | Towards Automated Test Adaptation in Fork Ecosystems via Large Language Models Ideas, Visions and Reflections Mukelabai Mukelabai Ruhr University Bochum, Keanu-Wesley Schurkus Ruhr University Bochum, Yannic Noller Ruhr University Bochum, Thorsten Berger Ruhr University Bochum | ||
14:40 20mTalk | Boosting LLMs for Mutation Generation Research Papers Bo Wang Beijing Jiaotong University, Ming Deng Beijing Jiaotong University, Mingda Chen Beijing Jiaotong University, Chengran Yang Singapore Management University, Singapore, Youfang Lin Beijing Jiaotong University, Mark Harman Meta Platforms, Inc. and UCL, Mike Papadakis University of Luxembourg, Jie M. Zhang Mistral AI and King's College London | ||
15:00 20mTalk | LLM-Assisted Input-Requirement-Aware Differential Testing of Array Programming Frameworks Research Papers Zhichao Zhou School of Information Science and Technology, ShanghaiTech University, Jingzhu He ShanghaiTech University Pre-print | ||
15:20 10mTalk | AISysRev - LLM-based Tool for Title-abstract Screening Tool Demonstrations Aleksi Huotala University of Helsinki, Miikka Kuutila LUT University, Olli-Pekka Turtio University of Helsinki, Simo Sipilä University of Helsinki, Mika Mäntylä University of Helsinki Pre-print | ||