JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)
We build a benchmark to evaluate large language models (LLMs) for source code migration tasks, specifically upgrading functions from Java 8 to Java 11. We first collected a dataset of function pairs from open-source repositories, but limitations in data quality led us to construct a refined dataset covering eight categories of deprecated APIs. Using this dataset, the Mistral Codestral model was evaluated with CodeBLEU and keyword-based metrics to measure lexical and semantic similarity as well as migration correctness. Results show that the LLM can handle trivial one-to-one API substitutions with moderate success, achieving identical migrations in 11.11% of the cases, but struggles with more complex migrations such as CORBA or JAX-WS. These findings suggest that LLMs can partially reduce developer effort by automating repetitive migration tasks but cannot yet replace humans for more complex cases. The benchmark and analysis provide a foundation for future work on expanding datasets, refining prompting strategies, and improving migration performance across different LLMs.
Tue 14 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
09:00 - 10:30 | Code Transformation & Modernization 1ReCode 2026 at Bora Bora II Keynote and paper on code transformation & modernization | ||
09:00 15mDay opening | Opening by chairs ReCode 2026 | ||
09:15 60mKeynote | The Road to True Software Modernization with Autonomous Agents ReCode 2026 Omer Tripp Amazon | ||
10:15 15mTalk | JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11) ReCode 2026 Nishil Amin University College London, Zhiwei Fei Nanjing University, Xiang Li University College London, Justyna Petke University College London, He Ye University College London (UCL) | ||