FSE 2026
Sun 5 - Thu 9 July 2026 Montreal, Canada
Tue 7 Jul 2026 16:00 - 16:20 at MB 3.445 - Robustness Chair(s): Boqi Chen

Constructing and curating high-quality code datasets requires significant resources, making them valuable intellectual property. Unfortunately, these datasets currently face severe risks of unauthorized use. Although digital watermarking offers a post hoc mechanism for copyright authentication, existing methods are predominantly based on the co-occurrence pattern, which is not robust and is susceptible to watermark detection and removal attacks. In this paper, we propose PuzzleMark, a robust watermarking method for code datasets. To reduce the risk of watermark exposure, PuzzleMark introduces a carrier selection strategy that leverages code complexity to evaluate the suitability of code snippets as watermark carriers, and selects those with high suitability for watermarking. To enhance the robustness of the watermark, PuzzleMark proposes a novel concatenation pattern to replace the traditional co-occurrence pattern, and implements two watermarking strategies through variable name concatenation. PuzzleMark adaptively embeds watermarks based on the inherent characteristics of the code, making it more stealthy while maintaining design simplicity. For watermark verification, PuzzleMark employs an independent-samples $t$-test to verify suspicious models under a black-box setting. Experimental results demonstrate that PuzzleMark achieves a 100% verification success rate and a 0% false positive rate, with negligible impact on model performance. Both our human study and our evaluation using four state-of-the-art watermark detection methods show that PuzzleMark exhibits strong imperceptibility, with an average suspicious rate $\leq$ 0.24 and an average recall $\leq$ 30.41%, respectively. Furthermore, the consistent retention of verifiability under two attack scenarios further corroborates the robustness of PuzzleMark. As a practical digital watermarking method, PuzzleMark provides strong protection for the intellectual property of code datasets and offers new insights for future research.

Tue 7 Jul

Displayed time zone: Eastern Time (US & Canada) change

16:00 - 17:20
RobustnessIdeas, Visions and Reflections / Research Papers at MB 3.445
Chair(s): Boqi Chen University of Ottawa
16:00
20m
Talk
PuzzleMark: Implicit Jigsaw Learning for Robust Code Dataset Watermarking in Neural Code Completion Models
Research Papers
Haocheng Huang Soochow University, Yuchen Chen Nanjing University, Weisong Sun Nanyang Technological University, Peizhuo Lv Nanyang Technological University, Yuan Xiao Nanjing University, Chunrong Fang Nanjing University, Yang Liu Nanyang Technological University, Xiaofang Zhang Soochow University
Pre-print
16:20
20m
Talk
Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-preserving Transformations
Research Papers
Jiyong Uhm Sungkyunkwan University, Minseok Kim Sungkyunkwan University, Michalis Polychronakis Stony Brook University, Hyungjoon Koo Sungkyunkwan University
Pre-print
16:40
20m
Talk
Verifying Structural Robustness of Deep Neural Network
Research Papers
Hai Duong George Mason University, Thanh Le National Institute of Information and Communications Technology (NICT), Lam Nguyen CMC Applied Technology Institute (CATI), ThanhVu Nguyen George Mason University
Pre-print File Attached
17:00
10m
Talk
Towards Reliable Testing for Machine Unlearning
Ideas, Visions and Reflections
Anna Mazhar Cornell University, Sainyam Galhotra Cornell University
DOI Pre-print