Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks Researchers propose a novel framework that enriches feedback by using environments as scaffolds to bootstrap self-evolving agents in long-horizon tasks, addressing reward sparsity in reinforcement learning for large language models. The approach aims to improve autonomous agent training beyond conventional supervised fine-tuning. Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning RL for long-horizon tasks is often hindered by severe reward sparsity. While conventional agent-side warming up via supervised fine-tuning SFT can a