Show HN: Self-adjusting vLLM at production scale
Rivvr launched an autopilot product that automatically tunes vLLM inference deployments, claiming up to 2x higher tokens per second and 40-70% cuts in AWS bills. The company said the tool load-tests and tunes vLLM kernel…