# QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

> Source: <https://aiflash.com/news/128330/>
> Published: 2026-09-29 07:00:58+00:00

Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental ch
