Semantic-Driven Feature Identification In Software Variants Using Transformers
Feature identification in clone-and-own systems remains largely manual and time-consuming, as features are inherently semantic yet existing automated approaches rely on syntactic analysis, failing to capture semantically equivalent features across independently developed variants. Large Language Models (LLMs) offer a promising alternative through their semantic code understanding, but their context window limitation makes direct application to large software systems infeasible. In this paper, we present Semantic-driven Feature Identification using Transformers (SeFIT), a hierarchical decomposition-aggregation pipeline that overcomes this limitation by recursively partitioning systems into processable components, identifying features in each, and unifying them via transformer-based embeddings and HDBSCAN clustering. SeFIT requires no prior system knowledge, no manual effort, and scales independently of codebase size. We evaluate SeFIT on two established subject systems, ApoGames and ArgoUML, and establish a first performance baseline for fully automated semantic feature identification. Our results demonstrate that SeFIT is the first approach to overcome the context window problem and shows the practical potential of LLMs for variability mining in large-scale software systems, outperforming a syntax-based approach.