FSE 2026
Sun 5 - Thu 9 July 2026 Montreal, Canada
Sun 5 Jul 2026 16:10 - 16:20 at MB 3.445 - LLMTrust

With the rapid adoption of large language models (LLMs) in automated code refactoring, assessing and ensuring functional equivalence between LLM-generated refactoring and the original implementation becomes critical. While prior work typically relies on predefined test cases to evaluate correctness, in this work, we leverage differential fuzzing to check functional equivalence in LLM-generated code refactorings. Unlike test-based evaluation, a differential fuzzing-based equivalence checker needs no predefined test cases and can explore a much larger input space by executing and comparing thousands of automatically generated test inputs. In a large-scale evaluation of six LLMs (CodeLlama, Codestral, StarChat2, Qwen-2.5, Olmo-3, and GPT-4o) across three datasets and two refactoring types, we find that LLMs show a non-trivial tendency to alter program semantics, producing 19-35% functionally non-equivalent refactorings. Our experiments further demonstrate that about 21% of these non-equivalent refactorings remain undetected by the existing test suites of the three evaluated datasets. Collectively, the findings of this study imply that reliance on existing tests might overestimate functional equivalence in LLM-generated code refactorings, which remain prone to semantic divergence.

Sun 5 Jul

Displayed time zone: Eastern Time (US & Canada) change

16:00 - 18:00
LLMTrustLLMTrust at MB 3.445
16:00
10m
Talk
Latent Security Threats in Modern Software Code Logs
LLMTrust
Rrezarta Krasniqi University of North Carolina at Charlotte, Abanti Chakraborty Shruti University of North Carolina at Charlotte
16:10
10m
Talk
A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings
LLMTrust
Simantika Bhattacharjee Dristi University of Virginia, Matthew B Dwyer University of Virginia
16:20
8m
Talk
LLMSafeGuard: A Training-Free Framework for Safeguarding LLM Decoding via Context-Wise Similarity Validation
LLMTrust
Ximing Dong Centre for Software Excellence at Huawei Canada, Shaowei Wang University of Manitoba, Dayi Lin Centre for Software Excellence, Huawei Canada, Ahmed E. Hassan Queen’s University
16:35
50m
Panel
Panel and Breakout Discussion
LLMTrust
Abdelwahab Hamou-Lhadj Concordia University, Montreal, Canada, Jin L.C. Guo McGill University, Song Wang York University, Jinqiu Yang Concordia University
17:25
10m
Day closing
Closing Remarks
LLMTrust