High-quality data augmentation for code comment classification
Code comments serve a crucial role in software development for documenting functionality, clarifying design choices, and assisting with issue tracking. They capture developers’ insights about the surrounding source code, serving as an essential resource for both human comprehension and automated analysis. Nevertheless, since comments are in natural language, they present challenges for machine-based code understanding. To address this, recent studies have applied natural language processing (NLP) and deep learning techniques to classify comments according to developers’ intentions. However, existing datasets for this task suffer from size limitations and class imbalance, as they rely on manual annotations and may not accurately represent the distribution of comments in real-world codebases. To overcome this issue, we introduce new synthetic oversampling and augmentation techniques based on high-quality data generation to enhance the NLBSE’26 challenge datasets. Our Synthetic Quality Oversampling Technique and Augmentation Technique (Q-SYNTH) yield promising results, improving the base classifier by 2.56%. The source code is publicly available at https://github.com/ThomBors/NLBSE2026.
Sun 12 AprDisplayed time zone: Brasilia, Distrito Federal, Brazil change
16:00 - 18:30 | NLBSE ToolsNLBSE at Oceania VI Chair(s): Fabio Marcos De Abreu Santos Colorado State University, USA, Moritz Mock Free University of Bozen-Bolzano | ||
16:00 5mDay opening | NLBSE Tool Competition Opening NLBSE | ||
16:05 7mShort-paper | High-quality data augmentation for code comment classification NLBSE | ||
16:12 7mShort-paper | TraCC: Efficient Multi-Label Code Comment Classification Through Knowledge Distillation and Adaptive Thresholding NLBSE A: Pir Sami Ullah Shah FAST National University, A: Abdullah Ashfaq National University of Computer & emerging Sciences (FAST-NUCES), A: Ahmed Fasseh National University of Computer & emerging Sciences (FAST-NUCES), A: Dilawar Shah National University of Computer & emerging Sciences (FAST-NUCES) | ||
16:19 7mShort-paper | X-LoRA MME: Multi-Model Ensemble with Mixture of Experts for Code Comment Classification NLBSE | ||
16:26 7mShort-paper | Distilling Semantics: Efficient Multi-label Code Comment Classification NLBSE A: Muhammad Abdul Majeed National University of Computer & emerging Sciences (FAST-NUCES), A: Ahmed Bin Asim National University of Computer & emerging Sciences (FAST-NUCES) | ||
16:33 5mProduct announcement | NLBSE Tool Competition Awards and Closing NLBSE | ||
16:38 1h50mTutorial | NLBSE NVIDIA Tutorial: Building an LLM-based coding copilot NLBSE | ||
18:28 2mDay closing | NLBSE Closing NLBSE | ||