cd /news/ai-infrastructure/beyond-utilization-energy-conscious-… · home › topics › ai-infrastructure › article
[ARTICLE · art-142282] src=al.radbox.org ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving

Researchers Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, and Neeraja J. Yadwadkar presented EnerTune, an inference serving system that reduces energy consumption by 1.4-2.3× and power draw by 1.3-2.6× over state-of-the-art baselines while meeting performance SLOs, at SOSP '26 on September 28, 2026. EnerTune uses analytical models to estimate per-model performance and power and the power draw of colocated models on shared GPUs, feeding an energy-aware bin-packing algorithm that jointly determines model placement and configuration. The work targets GPU multiplexing systems that optimize solely for utilization and can counterintuitively increase energy consumption.

by read1 min views1 publishedSep 30, 2026

Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving Authors:

Prasoon Sinha , Dimitrios Liakopoulos , Nathan Lemma , Neeraja J. Yadwadkar Published:

In SOSP '26: ACM SIGOPS 32nd Symposium on Operating Systems Principles. September 28, 2026. Abstract:

GPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can counterintuitively increase energy consumption. Designing policies that treat power and energy as first-order metrics requires understanding how deployment decisions—GPU allocation size, operating frequency, and batch size—affect energy, latency, and throughput. These relationships are complex, leading existing approaches to rely on extensive profiling. At scale, profiling becomes prohibitively expensive: each model can be deployed under hundreds of configurations, and profiling itself incurs significant energy cost, necessitating accurate yet energy-conscious methods. Further, such systems must model the power draw of colocated models on shared GPUs and adapt to dynamic workload fluctuations. We present EnerTune, an inference serving system that reduces energy consumption while meeting performance SLOs. EnerTune introduces analytical models to estimate per-model performance and power, and the power draw of colocated models on shared GPUs, and uses them in an energy-aware bin-packing algorithm to jointly determine model placement and configuration. EnerTune meets performance SLOs while reducing energy consumption by 1.4-2.3× and power draw by 1.3-2.6× over state-of-the-art baselines. BibTeX:

@inproceedings{sinha2026,
  title = {{Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving}},
  author = {Prasoon Sinha and Dimitrios Liakopoulos and Nathan Lemma and Neeraja J. Yadwadkar},
  booktitle = {SOSP '26: ACM SIGOPS 32nd Symposium on Operating Systems Principles},
  year = 2026,
  doi = {10.1145/3830418.3843913},
}
── more in #ai-infrastructure 4 stories · sorted by recency
── more on @prasoon sinha 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/beyond-utilization-e…] indexed:0 read:1min 2026-09-30 · —