T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks A Mixture-of-Experts model called T1, with 122B total parameters, has been trained with reinforcement learning to operate a real shell in a cloud sandbox for up to 300+ tool calls on long-horizon terminal tasks such as coding and scientific discovery. The model targets terminal agent reinforcement learning, a capability the source frames as increasingly important as agent usage shifts toward long-horizon work. Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool