Efficient container allocation is central to the scalability and sustainability of microservice-based cloud systems. Traditional schedulers, based on static heuristics, struggle with dynamic workloads and often cause resource overprovisioning and energy waste. This paper formulates container placement as a reinforcement-learning problem and introduces a Proximal Policy Optimization (PPO) agent that learns adaptive allocation and power-management policies to balance performance, utilization, and energy efficiency. Implemented as a Kubernetes-native scheduler, the proposed approach enables sustainability-aware orchestration without altering core components. Experimental results show that the PPO-based agent outperforms classical heuristics in consolidation and scalability, achieving higher utilization with fewer active nodes. These findings demonstrate the potential of reinforcement learning to drive greener, more cost-effective cloud infrastructures.
Dennis Junger HTW Berlin, University of Applied Sciences, Environmental Informatics Unit, Volker Wohlgemuth HTW Berlin, University of Applied Sciences, Environmental Informatics Unit