QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents QwenGyre is an elastic reinforcement learning framework for training extreme-long-horizon LLM agents, where a single execution can span hours, hundreds of model-environment interactions, and nearly 1M tokens per rollout. The framework targets the two fundamental challenges of applying online reinforcement learning to such executions. Large language model LLM agents increasingly undertake extreme-long xlong horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning RL to such executions poses two fundamental ch