DataFlex-RL: An Evaluation Platform for RLVR Data Policies Researchers introduced DataFlex-RL, an evaluation platform for comparing data policies in reinforcement learning with verifiable rewards (RLVR) under a common GRPO recipe. The platform targets how RLVR data policies decide which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. Data policies for reinforcement learning with verifiable rewards RLVR determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO rec