Reduce, Reuse, Recycle: Is this Mantra under attack with Proliferation of Machine Learning?Keynote
Memory systems remain the dominant bottleneck in modern computing, yet traditional principles of “reduce, reuse, recycle” are increasingly strained by the explosive growth of machine learning workloads. While decades of computer architecture research emphasized locality-aware execution, memory hierarchy optimization, and efficient resource reclamation, contemporary AI systems often prioritize throughput and model scale at the expense of memory efficiency. This talk examines the classical philosophy of memory management through the lens of machine learning acceleration. It examines whether current ML frameworks, accelerators, and runtime systems still exploit temporal and spatial reuse effectively, minimize unnecessary data movement, and recycle memory resources efficiently under increasingly data-centric and streaming workloads. By analyzing inference and training behaviors across GPUs, CPUs, and FPGA-based accelerators, one can identify emerging inefficiencies caused by oversized models, redundant tensor movement, allocator fragmentation, and synchronization overheads. However, there are several opportunities enabled by near-memory computing, heterogeneous memory systems, sparsity-aware execution, and compiler/runtime co-design to restore efficiency-centric design principles. Future scalable AI platforms must move beyond raw computational capability and re-embrace memory-conscious computing, where intelligent reuse, reduction of movement, and efficient recycling of resources become first-class architectural objectives once again.
Tue 16 JunDisplayed time zone: Mountain Time (US & Canada) change
13:40 - 14:40 | |||
13:40 60mKeynote | Reduce, Reuse, Recycle: Is this Mantra under attack with Proliferation of Machine Learning?Keynote ISMM 2026 Lizy John University of Texas, Austin | ||