Spade: Self-Play in Adaptive Synthetic Executable Environments
Researchers introduced SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play reinforcement learning framework in which a single large language model acts as both an Environment Designer and a Reaso…