An empirical study on LLM-based classification of requirements-related provisions in food-safety regulations
As Industry 4.0 transforms the food industry, the role of software in achieving compliance with food-safety regulations is becoming increasingly critical. Food-safety regulations, like those in many legal domains, have largely been articulated in a technology-independent manner to ensure their longevity and broad applicability. However, this approach leaves a gap between the regulations and the modern systems and software increasingly used to implement them. In this article, we pursue two main goals. First, we conduct a Grounded Theory study of food-safety regulations and develop a conceptual characterization of food-safety concepts that closely relate to systems and software requirements. Second, we examine the effectiveness of two families of large language models (LLMs) – BERT and GPT – in automatically classifying legal provisions based on requirements-related food-safety concepts. Our results show that: (a) when fine-tuned, the accuracy differences between the best-performing models in the BERT and GPT families are relatively small. Nevertheless, the most powerful model in our experiments, GPT-4o, still achieves the highest accuracy, with an average Precision of 89% and an average Recall of 87%; (b) few-shot learning with GPT-4o increases Recall to 97% but decreases Precision to 65%, suggesting a trade-off between fine-tuning and few-shot learning; (c) despite our training examples being drawn exclusively from Canadian regulations, LLM-based classification performs consistently well on test provisions from the US, indicating a degree of generalizability across regulatory jurisdictions; and (d) for our classification task, LLMs significantly outperform simpler baselines constructed using long short-term memory (LSTM) networks and automatic keyword extraction.
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
10:30 - 12:30 | Requirement and SpecificationResearch Papers / Journal-First Paper at MB 2.430 Chair(s): Neil Ernst University of Victoria | ||
10:30 20mTalk | Automated Repair of Requirements for Cyber-Physical Systems in Simulink Requirements Tables Research Papers Aren Babikian University of Toronto, Alessio Di Sandro University of Toronto, Federico Formica McMaster University, Claudio Menghi University of Bergamo; McMaster University, Marsha Chechik University of Toronto Pre-print | ||
10:50 20mTalk | An empirical study on LLM-based classification of requirements-related provisions in food-safety regulations Journal-First Paper Shabnam Hassani University of Ottawa, Mehrdad Sabetzadeh University of Ottawa, Daniel Amyot University of Ottawa Pre-print | ||
11:10 20mTalk | Speculate: Generating REST API Specifications Using LLMs Research Papers Krishanu Singh IIT Delhi, Kushagra Karar IIT Delhi, Abhilash Jindal IIT Delhi, India, Guowei Yang University of Queensland | ||
11:30 20mTalk | Requirements Coverage-Guided Minimization for Natural Language Test Cases Journal-First Paper RONGQI PAN University of Ottawa, Feifei Niu Graz University of Technology, Lionel Briand University of Ottawa, Canada; Lero centre, University of Limerick, Ireland, Hanyang Hu Company A | ||
11:50 20mTalk | SpecWeaver: End-to-End HTTP API Specification Inference Across Multi-Layer Routing in Production Web Services Research Papers Wenbo Hu Institute of Information Engineering at Chinese Academy of Sciences, Jie Lu SKLP, Institute of Computing Technology, Chinese Academy of Sciences, Jingting Chen Institute of Information Engineering, Chinese Academy of Sciences, Feng Li Key Laboratory of Network Assessment Technology, Institute of Information Engineering, Chinese Academy of Sciences, China; School of CyberSpace Security at University of Chinese Academy of Sciences, China, Chenghang Shi SKLP, Institute of Computing Technology, CAS, Xiaonan Shi Institute of Information Engineering, Chinese Academy of Sciences and School of Cyber Security, University of Chinese Academy of Sciences, Jinchen Wang Institute of Information Engineering, Chinese Academy of Sciences and School of Cyber Security, University of Chinese Academy of Sciences, Wei Huo Institute of Information Engineering at Chinese Academy of Sciences | ||
12:10 20mTalk | CASCADE: Detecting Inconsistencies between Code and Documentation with Automatic Test Generation Research Papers Tobias Kiecker Humboldt-Universität zu Berlin, Jan Arne Sparka Humboldt-Universität zu Berlin, Martin Reuter Humboldt-Universität zu Berlin, Albert Ziegler XBow, Lars Grunske Humboldt-Universität zu Berlin Pre-print | ||