Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution, according to a study that characterizes the gap through end-to-end latency measurements of model inference and the robot execution chain. The researchers report that repeated Flow Matching denoising contributes substantially to inference cost, and propose stage-aware two-step flow denoising with system-level evaluation toward real-time VLAs. Vision-language-action VLA models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain. Repeated Flow Matching denoising contributes substantially to inference c