04:00
2026-08-10
arxiv.org
machine-learning
Progressive Content Refinement with Decaying Reward Joint LinUCB
Researchers propose a new contextual bandit algorithm that explicitly models reward decay to improve iterative refinement in Large Language Models (LLMs), addressing over-exploitation from static promβ¦