{"slug": "a-22000-gpu-bill-with-a-very-quiet-night-shift", "title": "A $22,000 GPU bill with a very quiet night shift", "summary": "A consultant working with an eight-engineer ML startup cut its $22,000 monthly AWS GPU inference bill by $4,800 by aligning cluster capacity with actual traffic patterns. The startup's inference requests were concentrated between 9am and 11pm US Eastern, but GPU instances ran at full capacity around the clock, so the team added a schedule that scales the cluster down to a minimal warm state at 11pm and back up at 8am. The fix took two days to implement and test, and required no architecture or model changes.", "body_md": "The one I think about most is the account where the infrastructure was correct and the cost was still wrong\n\nML startup. Eight engineers. Production inference API with real traffic\n\nTheir AWS bill was $22,000 a month. Mostly GPU compute for inference\n\nI went in expecting to find over-provisioned instances. I found the opposite. The instance types were appropriate for the workload. Utilization was reasonable. Nothing obviously wasteful\n\nWhat I found instead was the scheduling\n\nTheir inference traffic followed a completely predictable pattern. 90 percent of requests came between 9am and 11pm US Eastern time. The overnight period was nearly silent. A few health checks. Background jobs. No real user traffic\n\nThe GPU instances ran 24 hours a day\n\nAt full capacity. Full price. Through eight hours of near-zero utilization every single night\n\nWe implemented a scale-down schedule. At 11pm Eastern, scale the inference cluster down to a minimal warm state - enough to handle the background jobs and respond to the health checks. At 8am Eastern, scale back up before the traffic arrives\n\nThe change took two days to implement and test safely\n\nMonthly savings: $4,800\n\nNot from changing the architecture. Not from changing the model. Not from changing anything about how the product worked\n\nFrom matching the cost to the actual demand pattern instead of running at peak capacity all the time because that felt safer\n\nThe infrastructure was fine\n\nThe schedule was wrong\n\nSometimes that's all it is", "url": "https://wpnews.pro/news/a-22000-gpu-bill-with-a-very-quiet-night-shift", "canonical_source": "https://dev.to/vlad_z_16b6320e21f32bee0d/a-22000-gpu-bill-with-a-very-quiet-night-shift-5c81", "published_at": "2026-10-06 21:12:33+00:00", "updated_at": "2026-10-06 21:18:13.054470+00:00", "lang": "en", "topics": ["ai-infrastructure", "mlops"], "entities": ["AWS"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/a-22000-gpu-bill-with-a-very-quiet-night-shift", "markdown": "https://wpnews.pro/news/a-22000-gpu-bill-with-a-very-quiet-night-shift.md", "text": "https://wpnews.pro/news/a-22000-gpu-bill-with-a-very-quiet-night-shift.txt", "jsonld": "https://wpnews.pro/news/a-22000-gpu-bill-with-a-very-quiet-night-shift.jsonld"}}