# Show HN: Self-adjusting vLLM at production scale

> Source: <https://twitter.com/yevr19/status/2099877116892307744>
> Published: 2026-09-15 17:32:40+00:00

𝗥𝘂𝗻𝗻𝗶𝗻𝗴 

[@vllm_project](https://x.com/vllm_project)𝗶𝘀 𝗲𝗮𝘀𝘆. 𝗞𝗲𝗲𝗽𝗶𝗻𝗴 𝗹𝗮𝘁𝗲𝗻𝗰𝘆 𝗽𝗿𝗲𝗱𝗶𝗰𝘁𝗮𝗯𝗹𝗲 𝘄𝗵𝗲𝗻 𝟯𝟬𝟬 𝗰𝘂𝘀𝘁𝗼𝗺𝗲𝗿𝘀 𝗮𝗿𝗿𝗶𝘃𝗲 𝗮𝘁 𝗼𝗻𝗰𝗲 𝗶𝘀 𝘁𝗵𝗲 𝗵𝗮𝗿𝗱 𝗽𝗮𝗿𝘁. Imagine a delayed flight. Hundreds of passengers call the airline’s AI voice agent simultaneously. Requests queue, TTFT climbs, callers wait in silence. An engineer gets paged. Now your team is tuning configurations, rerunning load tests, and adding spare capacity for the next spike. Over-provisioning buys headroom but it also leaves you paying for that headroom between bursts. We built[@rivvrai](https://x.com/rivvrai)to automate this operational work. You set the model’s latency and throughput targets. Rivvr’s autopilot: • Load-tests and tunes vLLM kernels, shipping 𝘂𝗽 𝘁𝗼 𝟮𝘅 𝗵𝗶𝗴𝗵𝗲𝗿 𝗧𝗣𝗦 • Monitors metrics, adjusts cluster topology on the fly to 𝗮𝘂𝘁𝗼𝗺𝗮𝘁𝗶𝗰𝗮𝗹𝗹𝘆 𝗸𝗲𝗲𝗽 𝗦𝗟𝗢 𝘁𝗮𝗿𝗴𝗲𝘁𝘀 • Optimizes compute costs, 𝗰𝘂𝘁𝘁𝗶𝗻𝗴 𝟰𝟬-𝟳𝟬% 𝗶𝗻 𝗔𝗪𝗦 𝗯𝗶𝗹𝗹 • Handles incident recovery, rollouts, and rollbacks For example, it can increase memory headroom to lower TTFT, migrate to lower-cost Spot instances, or switch VM sizes when AWS runs out of capacity. And these are just a fraction of Autopilot's capabilities. 𝗬𝗼𝘂𝗿 𝘁𝗲𝗮𝗺 𝘀𝗲𝘁𝘀 𝘁𝗵𝗲 𝗿𝗲𝗾𝘂𝗶𝗿𝗲𝗺𝗲𝗻𝘁𝘀. 𝗥𝗶𝘃𝘃𝗿 𝗿𝘂𝗻𝘀 𝘁𝗵𝗲 𝗶𝗻𝗳𝗿𝗮𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲 𝟮𝟰/𝟳. Watch the demo below:
00:00
