Validation is a central activity when developing formal specifications. Similarly to coding, a possible validation technique is to define upfront test cases or scenarios that a future specification should satisfy or not. Unfortunately, specifying such test cases is burdensome and error prone, which could cause users to skip this validation task. This paper reports the results of an empirical evaluation of using pre-trained large language models (LLMs) to automate the generation of test cases from natural language requirements. In particular, we focus on test cases for structural requirements of simple domain models formalized in the Alloy specification language. Our evaluation focuses on the state-of-the-art GPT-5 model, but results from other closed- and open-source LLMs are also reported. The results show that, in this context, GPT-5 is already quite effective at generating positive (and negative) test cases that are syntactically correct and that satisfy (or not) the given requirement, and that can detect many wrong specifications written by humans.
| Slides (Talk.pdf) | 1.50MiB |
Thu 21 MayDisplayed time zone: Osaka, Sapporo, Tokyo change
16:05 - 17:55 | Session 4: LLMs Formal MethodsResearch Track at 2F Auditorium Chair(s): Carlo A. Furia Università della Svizzera italiana (USI) | ||
16:05 25mTalk | Can LLM Aid in Solving Constraints with Inductive Definitions? Research Track Weizhi Feng Institute of Software, Chinese Academy of Sciences, Shidong Shen Institute of Software, Chinese Academy of Sciences, Jiaxiang Liu Institute of Software, Chinese Academy of Sciences, Taolue Chen Birkbeck, University of London, Fu Song Institute of Software at Chinese Academy of Sciences; University of Chinese Academy of Sciences; Nanjing Institute of Software Technology, Zhilin Wu Institute of Software at Chinese Academy of Sciences; University of Chinese Academy of Sciences | ||
16:30 25mTalk | Towards Language Model Guided TLA+ Proof Automation Research Track | ||
16:55 25mResearch paper | Validating Formal Specifications with LLM-generated Test Cases Research Track Pre-print File Attached | ||
17:20 25mTalk | ModelWisdom: An Integrated Toolkit for TLA+ Model Visualization, Digest and Repair Research Track Zhiyong Chen Nanjing University, Jialun Cao Hong Kong University of Science and Technology, Chang Xu Nanjing University, Shing-Chi Cheung Department of Computer Science and Engineering, The HongKong University of Science and Technology, Hong Kong, China | ||