LLM4MTLs: Automated Generation and Empirical Evaluation of Model Transformation Languages
Model transformation languages (MTLs) are domain-specific languages used to transform models conforming to a given metamodel into other models, including textual models such as source code. Developing correct model transformations in these languages is challenging and requires both language-specific and domain knowledge, creating a need for automated assistance and thus motivating the use of large language models (LLMs) for MTL code generation. However, due to the limited availability of training data and executable examples, LLM-generated MTL code is often not syntactically valid or semantically usable out of the box.
This paper presents LLM4MTLs, an automated workflow for improving the reliability of LLM-generated MTL code through systematic prompt construction that allows flexible selection of prompting strategy, together with an evaluation suite and an empirical evaluation. The workflow systematically explores prompt constructions combining few-shot prompting, grammar prompting, and helper methods inclusion, and evaluates them using both syntactic and semantic metrics. We construct an evaluation suite spanning four MTLs (ATL, ETL, QVTo, and the Reactions language) with executable reference scripts and manually written test suites, and evaluate across three LLMs. e find that few-shot prompting consistently improves syntactic quality across all four MTLs and yields notable gains in semantic correctness, although its ability to improve semantic correctness decreases for larger and more complex transformations, such as long ATL codes. Grammar prompting stabilizes code generation when combined with few-shot examples, but in isolation, it can be ineffective or even counterproductive for certain model–language combinations. Furthermore, including helper methods in the prompt as a complementary amplifier is beneficial. Finally, LLM Model choice influences syntactic correctness and similarity for certain MTLs, particularly ETL and QVTo, while its influence on semantic correction remains limited across all MTLs.
Wed 1 JulDisplayed time zone: Brussels, Copenhagen, Madrid, Paris change
13:30 - 15:00 | Language ModelsECMFA 2026 at Petri Chair(s): Oszkár Semeráth Budapest University of Technology and Economics | ||
13:30 30mTalk | EMF-Kaizen: an intelligent assistant for domain-specific modelling and meta-modelling ECMFA 2026 Lissette Almonte Universidad Autónoma de Madrid, Jefferson Ivan Rengifo Universidad Autónoma de Madrid, Esther Guerra Universidad Autónoma de Madrid, Juan de Lara Autonomous University of Madrid Link to publication Pre-print Media Attached | ||
14:00 30mTalk | LLM4MTLs: Automated Generation and Empirical Evaluation of Model Transformation Languages ECMFA 2026 Bowen Jiang Karlsruhe Institute of Technology, Nathan Hagel Karlsruhe Institute of Technology (KIT), Haowei Cheng Waseda University, Benedikt Jutz Karlsruhe Institute of Technology (KIT), Arne Lange Karlsruhe Institute of Technology (KIT), Weixing Zhang Karlsruhe Institute of Technology (KIT), Rahul Sharma Karlsruhe Institute of Technology, Ralf Reussner Karlsruhe Institute of Technology (KIT) and FZI - Research Center for Information Technology (FZI), Anne Koziolek Karlsruhe Institute of Technology | ||
14:30 30mTalk | LLM-Powered Multi-Agent Systems: Exploring Documentation-Driven Metamodeling ECMFA 2026 James Pontes Miranda CEA LIST, Ansgar Radermacher Université Paris-Saclay, CEA List, Palaiseau, Marcos Didonet Del Fabro CEA-List, Fabien Baligand Université Paris-Saclay, CEA List, Palaiseau, Julie Bonnail Université Paris-Saclay, CEA List, Palaiseau, Pascal Bannerot Université Paris-Saclay, CEA List, Palaiseau, Kunal Suri Université Paris-Saclay, CEA List, Palaiseau | ||