A $22,000 GPU bill with a very quiet night shift A consultant working with an eight-engineer ML startup cut its $22,000 monthly AWS GPU inference bill by $4,800 by aligning cluster capacity with actual traffic patterns. The startup's inference requests were concentrated between 9am and 11pm US Eastern, but GPU instances ran at full capacity around the clock, so the team added a schedule that scales the cluster down to a minimal warm state at 11pm and back up at 8am. The fix took two days to implement and test, and required no architecture or model changes. The one I think about most is the account where the infrastructure was correct and the cost was still wrong ML startup. Eight engineers. Production inference API with real traffic Their AWS bill was $22,000 a month. Mostly GPU compute for inference I went in expecting to find over-provisioned instances. I found the opposite. The instance types were appropriate for the workload. Utilization was reasonable. Nothing obviously wasteful What I found instead was the scheduling Their inference traffic followed a completely predictable pattern. 90 percent of requests came between 9am and 11pm US Eastern time. The overnight period was nearly silent. A few health checks. Background jobs. No real user traffic The GPU instances ran 24 hours a day At full capacity. Full price. Through eight hours of near-zero utilization every single night We implemented a scale-down schedule. At 11pm Eastern, scale the inference cluster down to a minimal warm state - enough to handle the background jobs and respond to the health checks. At 8am Eastern, scale back up before the traffic arrives The change took two days to implement and test safely Monthly savings: $4,800 Not from changing the architecture. Not from changing the model. Not from changing anything about how the product worked From matching the cost to the actual demand pattern instead of running at peak capacity all the time because that felt safer The infrastructure was fine The schedule was wrong Sometimes that's all it is