arXiv:2610.07025v1 Announce Type: new Abstract: Multi-person 3D pose estimation with WiFi channel state information (CSI) is challenging because reflections from different people overlap without directly identifying individual joints. Existing masked embedding objectives capture wireless relationships without explicit pose supervision, while structured decoders can retain coordinate errors. We propose WiSPER, a two-stage framework combining pose-aware predictive pretraining with conditional residual flow refinement. Pose-Aware Masked Embedding Learning (PAMEL) couples masked latent prediction with auxiliary pose-set supervision on the same CSI context, guiding the encoder toward joint localization from partial observations. Residual Flow refinement with Transformer (ReFT) generates a set of pose candidates to accommodate a variable number of people and refines each candidate through a conditional flow guided by its coarse coordinates and per-joint decoder features. Both stages use paired CSI and pose annotations during training, while inference requires only CSI. Experiments on the PiW3D dataset show that WiSPER achieves an overall mean per-joint position error of 63.72 mm, a 40.0% reduction relative to WiFi-JEPA. For experiments with two and three people, WiSPER reduces MPJPE by 42.1% and 38.1%, respectively. Pose-supervised pretraining configurations obtain lower errors than CSI-only JEPA, and enabling the trained residual refiner reduces overall MPJPE by 13.8-15.6% across the evaluated configurations.
WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI
A two-stage framework called WiSPER achieved an overall mean per-joint position error of 63.72 mm on the PiW3D dataset for multi-person 3D pose estimation from WiFi channel state information, a 40.0% reduction relative to WiFi-JEPA, according to the arXiv paper 2610.07025v1. WiSPER combines Pose-Aware Masked Embedding Learning (PAMEL), which pairs masked latent prediction with auxiliary pose-set supervision, and Residual Flow refinement with Transformer (ReFT), which refines pose candidates via a conditional flow; the method reduced MPJPE by 42.1% for two people and 38.1% for three people, and the trained residual refiner cut overall MPJPE by 13.8-15.6% across evaluated configurations. Training uses paired CSI and pose annotations, while inference requires only CSI.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.