An Empirical Evaluation of Generative AI in Security Requirements Engineering and Threat Modeling
The manual generation of software development artifacts in large organizations—particularly security requirements and threat models—demands substantial effort and is prone to inconsistencies and coverage gaps. While recent advances in generative AI show promise for supporting Requirements Engineering, their adoption in security-critical and regulated environments remains limited due to concerns related to trust, data privacy, and domain specificity.
This paper presents an empirical evaluation of an LLM-based AI assistant designed to support security requirements engineering and threat modeling under controlled and auditable conditions. The proposed approach adopts a human–AI hybrid workflow, in which AI-generated artifacts are systematically reviewed and validated by specialists to preserve contextual accuracy and regulatory compliance. Using business documents as input, the assistant generates candidate security requirements and STRIDE-based threat models.
The study compares AI-only, manual, and hybrid workflows across 20 real-world projects conducted in a Brazilian public organization. Results show an average reduction of 18.3% in artifact generation time and a 13.6% increase in STRIDE threat coverage, while maintaining 73% semantic precision. Furthermore, the hybrid human–AI approach consistently outperformed the fully manual process in terms of completeness and overall quality.
These findings provide empirical evidence that generative AI can effectively support security requirements engineering when embedded within human-centered workflows and organizational governance structures, offering practical insights for adoption in regulated software development contexts.