cd /news/computer-vision/wisper-pose-supervised-predictive-an… · home › topics › computer-vision › article
[ARTICLE · art-146529] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI

A two-stage framework called WiSPER achieved an overall mean per-joint position error of 63.72 mm on the PiW3D dataset for multi-person 3D pose estimation from WiFi channel state information, a 40.0% reduction relative to WiFi-JEPA, according to the arXiv paper 2610.07025v1. WiSPER combines Pose-Aware Masked Embedding Learning (PAMEL), which pairs masked latent prediction with auxiliary pose-set supervision, and Residual Flow refinement with Transformer (ReFT), which refines pose candidates via a conditional flow; the method reduced MPJPE by 42.1% for two people and 38.1% for three people, and the trained residual refiner cut overall MPJPE by 13.8-15.6% across evaluated configurations. Training uses paired CSI and pose annotations, while inference requires only CSI.

by read1 min views2 publishedOct 7, 2026

arXiv:2610.07025v1 Announce Type: new Abstract: Multi-person 3D pose estimation with WiFi channel state information (CSI) is challenging because reflections from different people overlap without directly identifying individual joints. Existing masked embedding objectives capture wireless relationships without explicit pose supervision, while structured decoders can retain coordinate errors. We propose WiSPER, a two-stage framework combining pose-aware predictive pretraining with conditional residual flow refinement. Pose-Aware Masked Embedding Learning (PAMEL) couples masked latent prediction with auxiliary pose-set supervision on the same CSI context, guiding the encoder toward joint localization from partial observations. Residual Flow refinement with Transformer (ReFT) generates a set of pose candidates to accommodate a variable number of people and refines each candidate through a conditional flow guided by its coarse coordinates and per-joint decoder features. Both stages use paired CSI and pose annotations during training, while inference requires only CSI. Experiments on the PiW3D dataset show that WiSPER achieves an overall mean per-joint position error of 63.72 mm, a 40.0% reduction relative to WiFi-JEPA. For experiments with two and three people, WiSPER reduces MPJPE by 42.1% and 38.1%, respectively. Pose-supervised pretraining configurations obtain lower errors than CSI-only JEPA, and enabling the trained residual refiner reduces overall MPJPE by 13.8-15.6% across the evaluated configurations.

── more in #computer-vision 4 stories · sorted by recency
── more on @wisper 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/wisper-pose-supervis…] indexed:0 read:1min 2026-10-07 · —