RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks? A new benchmark and dataset called RoboSPA evaluates Vision-Language-Action (VLA) models on robotic manipulation tasks with increased spatial and procedural complexity, moving beyond simple scenes and short-horizon tasks. The work highlights limitations in current VLA models' reasoning capabilities under more challenging conditions. Vision-Language-Action VLA models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural c