Online Learning for Cost-Efficient LLM Routing Ramp engineers developed a Thompson-sampling-based router that dynamically selects the most cost-effective large language model for each query, reducing inference costs by up to 40% while maintaining response quality. The system, detailed in a blog post on Ramp's engineering site, continuously learns from online feedback to balance exploration and exploitation across multiple LLM providers. Article URL: https://builders.ramp.com/post/thompson-sampling-model-routing Comments URL: https://news.ycombinator.com/item?id=49074574 Points: 1 Comments: 0