On the synchronization between Hugging Face pre-trained language models and their upstream GitHub repository
Pre-trained language models (PTLMs) have revolutionized the field of natural language processing (NLP), enabling significant advancements in tasks such as text generation and translation. Similar to software package management, which involves centralized version control and distributed consumption, PTLMs are trained using code and environment scripts hosted in an upstream repository (e.g., a GitHub (GH) repository), while the family of model variants trained from a given repository’s scripts are distributed using dedicated downstream distribution platforms like Hugging Face (HF). Despite these similarities, coordinating development activities between GH and HF presents several challenges to avoid misaligned release timelines, inconsistent versioning practices, and other obstacles to seamless reuse of PTLM variants. To understand how commit activities are coordinated between these two platforms, we conducted an in-depth mixed-method study of 325 PTLM families consisting of 904 HF PTLM variants. Our analysis reveals that GH contributors typically make changes related to specifying the version of the model, improving code quality, performance optimization, and dependency management within the training scripts, while HF contributors make changes related to improving model descriptions, data set handling, and setup required for model inference. Furthermore, to understand the synchronization aspects of commit activities between GH and HF, we examined three dimensions of these activities—lag (delay), type of synchronization, and intensity—which together yielded eight distinct synchronization patterns. The prevalence of partially synchronized patterns, such as \emph{Disperse synchronization} and \emph{Sparse synchronization}, reveals structural disconnects in current cross-platform release practices. These patterns often result in isolated changes—where improvements or fixes made on one platform are never replicated on the other—and in some cases, indicate an abandonment of one repository in favor of the other. Such fragmentation risks exposing end users to incomplete, outdated, or behaviorally inconsistent models. Hence, recognizing these synchronization patterns is critical for improving oversight and traceability in PTLM release workflows.
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
14:00 - 15:30 | LLM for SE 4Journal-First Paper / Ideas, Visions and Reflections / Research Papers at MB 1.210 Chair(s): Mohamad Kassab Boston University | ||
14:00 20mTalk | Detecting Code-Comment Inconsistencies in Smart Contracts by Combining LLM and Program Analysis Research Papers Jiashuo Zhang Peking University, China, Jiachi Chen Zhejiang University, Ting Zhang Peking University, Yue Li Peking University, Daoyuan Wu Lingnan University, Yanlin Wang Sun Yat-sen University, Jianbo Gao Peking University, Ting Chen University of Electronic Science and Technology of China, Zhong Chen | ||
14:20 20mTalk | Unfulfilled Promises: LLM-Based Detection of OS Compatibility Issues in Infrastructure as Code Research Papers Georgios-Petros Drosos ETH Zurich, Georgios Alexopoulos University of Athens, Thodoris Sotiropoulos ETH Zurich, Dimitris Mitropoulos University of Athens, Zhendong Su ETH Zurich | ||
14:40 10mTalk | DePro: Understanding the Role of LLMs in Debugging Competitive Programming Code Ideas, Visions and Reflections Nabiha Parvez Military Institute of Science And Technology, Tanvin Pallab Military Institute of Science And Technology, Mia Mohammad Imran Missouri University of Science and Technology, Tarannum Shaila Zaman University of Maryland Baltimore County | ||
14:50 20mTalk | A Large-Scale Empirical Evaluation of LLMs for Automated Self-Admitted Technical Debt Repayment Journal-First Paper Mohammad Sadegh Sheikhaei Queen's University, Yuan Tian Queen's University, Kingston, Ontario, Shaowei Wang University of Manitoba, Bowen Xu North Carolina State University | ||
15:10 20mTalk | On the synchronization between Hugging Face pre-trained language models and their upstream GitHub repository Journal-First Paper Adekunle Ajibode Queen's University, Abdul Ali Bangash Lahore University of Management Sciences, Oussama Ben Sghaier Queen's University, Bram Adams Queen's University, Ahmed E. Hassan Queen’s University | ||