# Online Learning with LLM Experts from Limited Feedback

> Source: <https://aiflash.com/news/119241/>
> Published: 2026-09-14 08:00:23+00:00

We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback. We formulate it as a bandit problem with K actions that represent experts and d features that encode prompts, over a horizon of T rounds. We propose alg
