Online Learning with LLM Experts from Limited Feedback Researchers formulated adaptive routing of prompts to large language model experts as a bandit problem with K actions representing experts and d features encoding prompts over a horizon of T rounds, aiming to maximize response quality in an online setting with limited feedback. The work proposes an algorithm for this online learning problem, though the source text cuts off before naming it. We study adaptive routing of prompts to large language model LLM experts to maximize response quality in an online setting with limited feedback. We formulate it as a bandit problem with K actions that represent experts and d features that encode prompts, over a horizon of T rounds. We propose alg