{"slug": "beyond-utilization-energy-conscious-gpu-sharing-for-inference-serving", "title": "Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving", "summary": "Researchers Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, and Neeraja J. Yadwadkar presented EnerTune, an inference serving system that reduces energy consumption by 1.4-2.3× and power draw by 1.3-2.6× over state-of-the-art baselines while meeting performance SLOs, at SOSP '26 on September 28, 2026. EnerTune uses analytical models to estimate per-model performance and power and the power draw of colocated models on shared GPUs, feeding an energy-aware bin-packing algorithm that jointly determines model placement and configuration. The work targets GPU multiplexing systems that optimize solely for utilization and can counterintuitively increase energy consumption.", "body_md": "Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving\nAuthors:\n \nPrasoon Sinha\n, \nDimitrios Liakopoulos\n, \nNathan Lemma\n, \nNeeraja J. Yadwadkar\nPublished:\n \nIn SOSP '26: ACM SIGOPS 32nd Symposium on Operating Systems Principles. September 28, 2026.\nAbstract:\n \nGPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can counterintuitively increase energy consumption. Designing policies that treat power and energy as first-order metrics requires understanding how deployment decisions—GPU allocation size, operating frequency, and batch size—affect energy, latency, and throughput. These relationships are complex, leading existing approaches to rely on extensive profiling. At scale, profiling becomes prohibitively expensive: each model can be deployed under hundreds of configurations, and profiling itself incurs significant energy cost, necessitating accurate yet energy-conscious methods. Further, such systems must model the power draw of colocated models on shared GPUs and adapt to dynamic workload fluctuations. We present EnerTune, an inference serving system that reduces energy consumption while meeting performance SLOs. EnerTune introduces analytical models to estimate per-model performance and power, and the power draw of colocated models on shared GPUs, and uses them in an energy-aware bin-packing algorithm to jointly determine model placement and configuration. EnerTune meets performance SLOs while reducing energy consumption by 1.4-2.3× and power draw by 1.3-2.6× over state-of-the-art baselines.\nBibTeX:\n \n\n```\n@inproceedings{sinha2026,\n  title = {{Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving}},\n  author = {Prasoon Sinha and Dimitrios Liakopoulos and Nathan Lemma and Neeraja J. Yadwadkar},\n  booktitle = {SOSP '26: ACM SIGOPS 32nd Symposium on Operating Systems Principles},\n  year = 2026,\n  doi = {10.1145/3830418.3843913},\n}\n```", "url": "https://wpnews.pro/news/beyond-utilization-energy-conscious-gpu-sharing-for-inference-serving", "canonical_source": "https://al.radbox.org/doi/10.1145/3830418.3843913", "published_at": "2026-09-30 04:41:22+00:00", "updated_at": "2026-09-30 04:48:12.351696+00:00", "lang": "en", "topics": ["ai-infrastructure", "mlops", "machine-learning", "ai-research"], "entities": ["Prasoon Sinha", "Dimitrios Liakopoulos", "Nathan Lemma", "Neeraja J. Yadwadkar", "EnerTune", "SOSP '26", "ACM SIGOPS"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/beyond-utilization-energy-conscious-gpu-sharing-for-inference-serving", "markdown": "https://wpnews.pro/news/beyond-utilization-energy-conscious-gpu-sharing-for-inference-serving.md", "text": "https://wpnews.pro/news/beyond-utilization-energy-conscious-gpu-sharing-for-inference-serving.txt", "jsonld": "https://wpnews.pro/news/beyond-utilization-energy-conscious-gpu-sharing-for-inference-serving.jsonld"}}