cd /news/ai-infrastructure/show-hn-self-adjusting-vllm-at-produโ€ฆ ยท home โ€บ topics โ€บ ai-infrastructure โ€บ article
[ARTICLE ยท art-130569] src=twitter.com โ†— pub= topic=ai-infrastructure verified=true sentiment=โ†‘ positive

Show HN: Self-adjusting vLLM at production scale

Rivvr launched an autopilot product that automatically tunes vLLM inference deployments, claiming up to 2x higher tokens per second and 40-70% cuts in AWS bills. The company said the tool load-tests and tunes vLLM kernels, adjusts cluster topology to hold SLO targets, and handles incident recovery, rollouts, and rollbacks, aimed at scenarios where hundreds of callers hit an AI voice agent at once. Rivvr stated its team sets the latency and throughput requirements while Rivvr runs the infrastructure 24/7.

read1 min views6 publishedSep 15, 2026
Show HN: Self-adjusting vLLM at production scale
Image: source

๐—ฅ๐˜‚๐—ป๐—ป๐—ถ๐—ป๐—ด

@vllm_project๐—ถ๐˜€ ๐—ฒ๐—ฎ๐˜€๐˜†. ๐—ž๐—ฒ๐—ฒ๐—ฝ๐—ถ๐—ป๐—ด ๐—น๐—ฎ๐˜๐—ฒ๐—ป๐—ฐ๐˜† ๐—ฝ๐—ฟ๐—ฒ๐—ฑ๐—ถ๐—ฐ๐˜๐—ฎ๐—ฏ๐—น๐—ฒ ๐˜„๐—ต๐—ฒ๐—ป ๐Ÿฏ๐Ÿฌ๐Ÿฌ ๐—ฐ๐˜‚๐˜€๐˜๐—ผ๐—บ๐—ฒ๐—ฟ๐˜€ ๐—ฎ๐—ฟ๐—ฟ๐—ถ๐˜ƒ๐—ฒ ๐—ฎ๐˜ ๐—ผ๐—ป๐—ฐ๐—ฒ ๐—ถ๐˜€ ๐˜๐—ต๐—ฒ ๐—ต๐—ฎ๐—ฟ๐—ฑ ๐—ฝ๐—ฎ๐—ฟ๐˜. Imagine a delayed flight. Hundreds of passengers call the airlineโ€™s AI voice agent simultaneously. Requests queue, TTFT climbs, callers wait in silence. An engineer gets paged. Now your team is tuning configurations, rerunning load tests, and adding spare capacity for the next spike. Over-provisioning buys headroom but it also leaves you paying for that headroom between bursts. We built@rivvraito automate this operational work. You set the modelโ€™s latency and throughput targets. Rivvrโ€™s autopilot: โ€ข Load-tests and tunes vLLM kernels, shipping ๐˜‚๐—ฝ ๐˜๐—ผ ๐Ÿฎ๐˜… ๐—ต๐—ถ๐—ด๐—ต๐—ฒ๐—ฟ ๐—ง๐—ฃ๐—ฆ โ€ข Monitors metrics, adjusts cluster topology on the fly to ๐—ฎ๐˜‚๐˜๐—ผ๐—บ๐—ฎ๐˜๐—ถ๐—ฐ๐—ฎ๐—น๐—น๐˜† ๐—ธ๐—ฒ๐—ฒ๐—ฝ ๐—ฆ๐—Ÿ๐—ข ๐˜๐—ฎ๐—ฟ๐—ด๐—ฒ๐˜๐˜€ โ€ข Optimizes compute costs, ๐—ฐ๐˜‚๐˜๐˜๐—ถ๐—ป๐—ด ๐Ÿฐ๐Ÿฌ-๐Ÿณ๐Ÿฌ% ๐—ถ๐—ป ๐—”๐—ช๐—ฆ ๐—ฏ๐—ถ๐—น๐—น โ€ข Handles incident recovery, rollouts, and rollbacks For example, it can increase memory headroom to lower TTFT, migrate to lower-cost Spot instances, or switch VM sizes when AWS runs out of capacity. And these are just a fraction of Autopilot's capabilities. ๐—ฌ๐—ผ๐˜‚๐—ฟ ๐˜๐—ฒ๐—ฎ๐—บ ๐˜€๐—ฒ๐˜๐˜€ ๐˜๐—ต๐—ฒ ๐—ฟ๐—ฒ๐—พ๐˜‚๐—ถ๐—ฟ๐—ฒ๐—บ๐—ฒ๐—ป๐˜๐˜€. ๐—ฅ๐—ถ๐˜ƒ๐˜ƒ๐—ฟ ๐—ฟ๐˜‚๐—ป๐˜€ ๐˜๐—ต๐—ฒ ๐—ถ๐—ป๐—ณ๐—ฟ๐—ฎ๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ ๐Ÿฎ๐Ÿฐ/๐Ÿณ. Watch the demo below: 00:00

โ”€โ”€ more in #ai-infrastructure 4 stories ยท sorted by recency
โ”€โ”€ more on @rivvr 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain โ€” perfect for shipping the agent you just read about.

$git push zahid main
โ†’ Live at https://your-agent.zahid.host โœ“
Get free account โ†’ Pricing
from โ‚ฌ0/mo ยท no card required
LIVE [news/show-hn-self-adjustiโ€ฆ] indexed:0 read:1min 2026-09-15 ยท โ€”