{"type": "article", "title": "How we made one of our largest inference workloads 4.7× more GPU-efficient", "publisher": "Web Pulse", "url": "https://wpnews.pro/news/how-we-made-one-of-our-largest-inference-workloads-4-7x-more-gpu-efficient", "original_source": "https://decagon.ai/blog/gpu-efficient-inference-serving-stack", "published": "2026-09-17T16:12:58+00:00", "accessed": "2026-09-17", "id": "how-we-made-one-of-our-largest-inference-workloads-4-7x-more-gpu-efficient"}