Effectiveness of symmetric metamorphic relations on validating the stability of code generation LLM (ASE 2025 - Journal-First Track)

Who

Chan Pak Yuen, Jacky Keung, Zhen Yang

Track

ASE 2025 Journal-First Track

This program is tentative and subject to change.

Time Zone

The program is currently displayed in (GMT+09:00) Seoul.

Use conference time zone: (GMT+09:00) SeoulSelect other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Mon 17 Nov 2025 15:20 - 15:30 at Vista - Code Generation 1

Abstract

Pre-trained large language models (LLMs) are increasingly used in software development for code generation to enhance productivity. Companies often prefer private LLMs over public ones to mitigate the risk of exposing corporate secrets. Validating the stability of the outputs from these LLMs is crucial, and our study proposes using symmetric Metamorphic Relations (MRs) from Metamorphic Testing (MT) for this purpose. Our study involved an empirical experiment with eight private LLMs, two public LLMs and two publicly available datasets. The test scenario simulated a software development environment where private LLMs were used to generate source codes for software enhancements and maintenance, while ensuring that the company’s software assets remained secure and unexposed. We defined seven symmetric MRs that are based on the principles of symmetry and semantic preservation. These MRs were used to generate “Follow-up” datasets from “Source” datasets for testing purposes. Our evaluation aimed to detect the occurrence of violations (inconsistent predictions) between “Source” and “Follow-up” datasets. We then assess the effectiveness of MRs in identifying correct and incorrect non-violated predictions from ground truths, as well as how the MRs influenced LLM performance by measuring the correctness of the generated codes. Results showed that one public and four private LLMs did not violate “Case transformation of prompts” MR. Furthermore, the effectiveness and performance results indicated that proposed MRs effectively explain the instability of LLM’s outputs through “Case transformation of prompts”, “Duplication of prompts”, and “Paraphrasing of prompts”. Additionally, the findings revealed that using a mix of LLM technologies and pre-training with a vast number of datasets can cause the MRs under study to be ineffective. Most LLMs, except copilot, demonstrated low semantic understanding and primarily relied on statistical patterns to interpret the prompts. This study demonstrated that the proposed MRs could serve as a validation tool with violation measurements for the stability of code generation LLMs’ outputs under the simulation, where no ground truth is available. The effectiveness and performance results indicated that proposed MRs are indeed effective tools for explaining the instability of LLM’s outputs. Moreover, the findings highlighted the necessity for enhancing LLMs’ semantic understanding of prompts to improve stability and suggested potential future research directions. These include exploring different MRs, enhancing semantic understanding, and applying symmetry to prompt engineering.

Chan Pak Yuen

Department of Computer Science, City University of Hong Kong, Kowloon, Hong Kong, China

Jacky Keung

City University of Hong Kong

Hong Kong SAR China

Zhen Yang

Shandong University

China