Beyond Model Optimization: Practical Energy-Efficient LLM Inference through Context-Aware Input Reduction
Large language models (LLMs) are increasingly power-hungry, and their rapid adoption has led to a surge in AI datacenter energy consumption, raising environmental concerns. However, most existing work on reducing LLM power use focuses on model-level or training-time optimizations, which are inaccessible to practitioners using closed-source “frontier” models such as GPT-5 or Gemini. To address this limitation, we propose two lightweight, input-level chat history reduction strategy—TF-IDF filtering and semantic search via sentence embeddings—that prune conversational context before inference without changing the model itself. We evaluate these strategies using a novel, production-scale, inference test framework built with vLLM and CodeCarbon for precise energy measurement, simulating real-world datacenter inference workloads. Experi- mental evaluation with Mistral-Small-24B-Instruct demonstrates that our methods achieve up to 66% energy savings per request for chat conversations with more than 19 messages, with only a modest (4–8%) drop in accuracy, and up to 93% savings for conversations longer than 200 messages. These results show that content-aware context pruning is a practical and highly effective approach to reducing inference energy consumption, enabling more sustainable and energy-efficient LLM deployment.
Tue 17 MarDisplayed time zone: Athens change
14:00 - 15:30 | |||
14:00 15mTalk | PPTAMη: Energy Aware CI/CD Pipeline for Container Based Applications Workshops & Tutorials Alessandro Aneggi Free University of Bozen-Bolzano, Andrea Janes Free University of Bozen-Bolzano, Xiaozhou Li Free University of Bozen-Bolzano | ||
14:15 25mTalk | GREENN: Granular evaluation of Energy Efficiency in Neural Networks Workshops & Tutorials Elena Ballesteros-Morallón University of Castilla-La Mancha, Felix García University of Castilla-La Mancha, Maria Gutierrez University of Castilla-La Mancha, Mª Angeles Moraga University of Castilla-La Mancha, Coral Calero Universidad de Castilla La Mancha | ||
14:40 25mTalk | Orchestrating AI-Driven Code Refactoring Based on Energy Measurements in CI Pipelines Workshops & Tutorials Carlos Pulido Hernández University of Castilla-La Mancha, Mª Angeles Moraga University of Castilla-La Mancha, Felix García University of Castilla-La Mancha | ||
15:05 25mTalk | Beyond Model Optimization: Practical Energy-Efficient LLM Inference through Context-Aware Input Reduction Workshops & Tutorials Kalle Pronk Fontys University of Applied Sciences, Qin Zhao Fontys University of Applied Science, Siebren Kazemier Q42 | ||