EfficientUICoder: A Bidirectional Token Compression Framework for Efficient MLLM-based UI Code Generation
Multimodal Large Language Models (MLLMs) have demonstrated exceptional performance in UI2Code tasks (i.e., generating code from UI mockups), significantly enhancing website development efficiency. However, UI2Code tasks incur substantially higher computational overhead compared to traditional code generation tasks. This overhead is primarily driven by the large number of input image tokens required to represent complex visual designs and the extensive volume of output code tokens needed to describe complete webpage structures. In this paper, we conduct a comprehensive preliminary study on popular MLLMs for UI2Code tasks, identifying significant redundancies in both image and code tokens. We observe that these redundancies not only exacerbate computational complexity but also hinder the model’s ability to focus on key UI elements, leading to excessively lengthy and often invalid HTML files.
To address these challenges, we propose EfficientUICoder, a bidirectional compression framework designed for efficient UI code generation. First, we introduce an Element and Layout-aware Token Compression method, which preserves essential UI element and layout information by detecting element regions and constructing a UI element tree for efficient representation. Second, we design a Region-aware Token Refinement strategy that refines selected tokens by leveraging attention scores to evaluate semantic importance, discarding lowattention tokens from selected region while integrating high-attention tokens from unselected regions. Third, we develop an Adaptive Duplicate Token Suppression mechanism, which dynamically modulates token probabilities during decoding by tracking HTML/CSS code structure frequencies and applying exponential penalty strategies to minimize repetitive generation. Extensive experiments demonstrate that EfficientUICoder achieves a 55%-60% compression ratio without compromising the quality of the generated webpages, effectively reducing output code redundancy. In terms of efficiency, EfficientUICoder achieves superior improvements, reducing computational cost by up to 44.9%, generated tokens by up to 41.4%, prefill time by up to 46.6%, and inference time by up to 48.8% on 34B-level MLLMs.
Wed 8 JulDisplayed time zone: Eastern Time (US & Canada) change
16:00 - 17:00 | |||
16:00 20mTalk | Look Before You Leap: Context-Sensitive GUI Grounding for Boosting Automated Extended Reality (XR) Testing Research Papers Shuqing Li The Chinese University of Hong Kong, Binchang Li Harbin Institute of Technology, Yepang Liu Southern University of Science and Technology, Cuiyun Gao Harbin Institute of Technology, Shenzhen, Jianping Zhang The Chinese University of Hong Kong, Shing-Chi Cheung Hong Kong University of Science and Technology, Michael Lyu The Chinese University of Hong Kong | ||
16:20 20mTalk | EfficientUICoder: A Bidirectional Token Compression Framework for Efficient MLLM-based UI Code Generation Research Papers Jingyu Xiao The Chinese University of Hong Kong, Zhongyi Zhang Huazhong University of Science and Technology, China, Yuxuan Wan The Chinese University of Hong Kong, Yintong Huo Singapore Management University, Singapore, Yang Liu Nanyang Technological University, Michael Lyu The Chinese University of Hong Kong Pre-print | ||
16:40 20mTalk | From Task to Tutorial: An Automated GUI Framework for Excel Tutorial Document and Video Creation Industry Papers Yuhang Xie Peking University, Jian Mu Nanjing University, Ma Xiaojun Microsoft, Chaoyun Zhang Microsoft, Lu Wang Microsoft Research, Mengyu Zhou Microsoft, Mugeng Liu Peking University, Si Qin Microsoft Research, Qingwei Lin Microsoft, Saravan Rajmohan Microsoft, Shi Han Microsoft Research, Dongmei Zhang Microsoft | ||