FSE 2026
Sun 5 - Thu 9 July 2026 Montreal, Canada
Tue 7 Jul 2026 11:50 - 12:10 at MB 9D - Benchmarking

With the increasing popularity of large language models (LLMs) and LLM-based agents, reliable and effective code evaluation metrics (CEMs) have become crucial for progress across several software engineering tasks. While popular benchmarks often provide test cases to assess the correctness of generated code, crafting and executing test cases is expensive. Reference-based CEMs provide a cheaper alternative by scoring a candidate program based on its functional similarity to a reference. Although prior research has focused on reporting the weak correlation between these CEMs and functional correctness, the causes are only assumed, and plausible solutions remain unexplored. In this work, we critically evaluate four state-of-the-art reference-based CEMs, revealing their strong bias towards surface-level features rather than code functionality. Despite this surface bias, current evaluation datasets for these CEMs rarely include code pairs that are surface-similar yet functionally dissimilar, or functionally similar yet surface-dissimilar. To mitigate this gap, we propose LoCaL ( Looks Can Lie), a CEM evaluation benchmark, with 3117 code pairs at both the method and program levels. Each pair is labeled with a functional similarity score and aims to target regions where CEMs are likely to perform poorly. The functional similarity scores are calculated through differential fuzzing, which eliminates the need for predefined test cases and, at the same time, improves the reliability of the scores by executing an order of magnitude more tests than prior work. We find that all four CEMs show significant performance degradation on LoCaL, compared to the baselines. Finally, based on our findings, we draw the implication that exposing CEMs to LoCaL-like data might facilitate the development of metrics that are robust to surface bias.

Tue 7 Jul

Displayed time zone: Eastern Time (US & Canada) change

11:00 - 12:30
11:00
20m
Talk
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
Journal-First Paper
Roham Koohestani JetBrains Research & Delft University of Technology, Philippe de Bekker Delft University of Technology, Begüm Koç Delft University of Technology, Mali Izadi Google & TU Delft
11:20
10m
Talk
Investigating Test Overfitting on SWE-bench
Ideas, Visions and Reflections
Toufique Ahmed IBM, Jatin Ganhotra IBM Research, Avraham Shinnar IBM Research, Martin Hirzel IBM Research
Pre-print
11:30
20m
Talk
CrypFormBench: Benchmarking Formal Analysis Capability of Large Language Models for Cryptographic schemes
Research Papers
Zhaoxuan Li Institute of Information Engineering, Chinese Academy of Sciences;School of Cyber Security, University of Chinese Academy of Sciences, Qionglu Zhang State Key Laboratory of Information Security, Institute of Information Engineering Chinese Academy of Sciences, Beijing, China, Hengyuan Liu State Key Laboratory of Information Security, Institute of Information Engineering Chinese Academy of Sciences, Beijing, China, Xiaoyan Gu State Key Laboratory of Information Security, Institute of Information Engineering Chinese Academy of Sciences, Beijing, China, Xianhui Lu State Key Laboratory of Information Security, Institute of Information Engineering Chinese Academy of Sciences, Beijing, China, Hongbo Liu State Key Laboratory of Information Security, Institute of Information Engineering Chinese Academy of Sciences, Beijing, China, Bingzheng Wang State Key Laboratory of Information Security, Institute of Information Engineering Chinese Academy of Sciences, Beijing, China, Haihui Fan State Key Laboratory of Information Security, Institute of Information Engineering Chinese Academy of Sciences, Beijing, China, Ziming Zhao Zhejiang University, Rui Zhang Taiyuan University of Science and Technology, Li Zhou Institute of Software, Chinese Academy of Sciences
DOI Pre-print
11:50
20m
Talk
LoCaL: Countering Surface Bias in Code Evaluation Metrics
Research Papers
Simantika Bhattacharjee Dristi University of Virginia, Matthew B Dwyer University of Virginia
DOI Pre-print
12:10
20m
Talk
VerilogASTBench: Benchmark Construction of Verilog AST Dataset with Dual-Stage AST Semantic Enhancement Framework
Research Papers
luping zhang Nanjing University of Posts and Telecommunications, Chao Chen Nanjing University of Posts and Telecommunications, Dapeng Yan Nanjing University of Posts and Telecommunications, Hui Xu Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China, Mingsheng Cao University of Electronic Science and Technology of China, Jingkuan Song University of Electronic Science and Technology of China, Zhikuang Cai Nanjing University of Posts and Telecommunications, Yufeng Guo Nanjing University of Posts and Telecommunications
Pre-print